Image processing method, electronic device, and computer-readable storage medium
By employing multi-dimensional detection and reference image adjustment technologies, the problem of monotonous visuals in product main images in e-commerce scenarios has been solved, thereby improving the visual expressiveness and conversion rate of images.
Patent Information
- Application Number
- CN202610458178.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-08
- Publication Date
- 2026-08-25
AI Technical Summary
In e-commerce scenarios, the visual presentation of product main images is monotonous and lacks appeal, resulting in limited traffic conversion. Existing technologies cannot effectively improve image processing effects.
By acquiring the visual information of the target object through multi-dimensional detection, a reference image is selected based on the category of the target object. The background area and object visual information of the reference image are combined to adjust the image to be processed and generate the target image, ensuring that the visual information of the target object remains unchanged and improving the overall visual expressiveness of the image.
It achieves the improvement of scene adaptability, visual harmony and style expression of image background without changing the visual information of the target object, thereby improving the overall visual appeal and conversion rate of the image.
Smart Images

Figure CN122636766A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology and the field of image processing, and more specifically, to an image processing method, an electronic device, and a computer-readable storage medium. Background Technology
[0002] Currently, in e-commerce scenarios, product main images directly impact user clicks and conversion rates. While products currently offer price advantages, their main images are generally visually monotonous and lack appeal, limiting traffic conversion. Related technologies largely rely on manual templates or generic generation models. The former suffers from rigid styles and cannot adapt to product category characteristics, while the latter is prone to altering the product image or obscuring key selling points, making it difficult to ensure content accuracy and visual harmony, resulting in poor image processing quality.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides an image processing method, an electronic device, and a computer-readable storage medium to at least solve the technical problem of poor image processing effect in related technologies.
[0005] According to one aspect of the embodiments of this application, an image processing method is provided, comprising: in response to receiving a to-be-processed image of a target object, performing multi-dimensional detection on the to-be-processed image to obtain object visual information of the target object, wherein the multi-dimensional detection is used to represent the process of detecting visual elements associated with the target object in the to-be-processed image in different dimensions; determining a reference image corresponding to the to-be-processed image based on the category of the target object, wherein the reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold; and adjusting the to-be-processed image based on the background region of the reference object in the reference image, the background region of the target object in the to-be-processed image, and the object visual information to obtain a target image of the target object.
[0006] According to one aspect of the embodiments of this application, an image processing method is provided, comprising: responding to an input command applied to an operation interface, displaying a target image of a target object on the operation interface; responding to a processing command applied to the operation interface, displaying a target image of the target object on the operation interface, wherein the target image is obtained based on the background region of a reference object in a reference image, the background region of the target object in the image to be processed, and object visual information, the reference image is determined based on the category of the target object, the object visual information is obtained by performing multi-dimensional detection on the image to be processed, the multi-dimensional detection being used to represent the process of detecting visual elements associated with the target object in the image to be processed in different dimensions, the reference image containing a reference object, and the similarity between the category of the reference object and the category of the target object being greater than a preset threshold.
[0007] According to another aspect of the embodiments of this application, a computing device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0008] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory storing an executable program; and a processor connected to the memory via a bus for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0009] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0010] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0011] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.
[0012] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0013] In this embodiment, in response to receiving a target object image to be processed, multi-dimensional detection is performed on the target object image to obtain object visual information. Multi-dimensional detection refers to the process of detecting visual elements associated with the target object in the target object image at different dimensions. Based on the category of the target object, a reference image corresponding to the target object image is determined. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold. Based on the background region of the reference object in the reference image, the background region of the target object in the target image, and the object visual information, the target image is adjusted to obtain the target image of the target object. By obtaining the visual information of the target object through multi-dimensional detection, combining it with the background features of the reference object in the reference image, and comparing the background semantic differences between the target image and the reference image, the deficiencies of the background of the target image in multiple dimensions such as scene adaptability, visual coordination, and style expressiveness are inferred. Then, based on these differences and the object visual information, a targeted improvement strategy is generated to adjust the background region, ensuring that the target object visual information remains unchanged, improving the overall visual expressiveness of the image, and thus solving the technical problem of poor image processing effect in related technologies.
[0014] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description
[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0016] Figure 1 This is a scene diagram illustrating an image processing procedure according to an embodiment of this application;
[0017] Figure 2 This is a flowchart of an image processing method according to an embodiment of this application;
[0018] Figure 3 This is a flowchart of an e-commerce product background improvement method according to an embodiment of this application;
[0019] Figure 4 This is a flowchart of a multi-dimensional intelligent analysis method according to an embodiment of this application;
[0020] Figure 5 This is a schematic diagram illustrating the construction of a competitor knowledge base according to an embodiment of this application;
[0021] Figure 6 This is a schematic diagram of an image processing structure according to an embodiment of this application;
[0022] Figure 7 This is a flowchart of an image processing method according to an embodiment of this application;
[0023] Figure 8 This is a schematic diagram of an image processing apparatus according to an embodiment of this application;
[0024] Figure 9 This is a schematic diagram of an image processing apparatus according to an embodiment of this application;
[0025] Figure 10 This is a structural block diagram of a computing device according to an embodiment of this application;
[0026] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0030] The technical solution provided in this application is mainly implemented using a deep learning model. Deep learning models can be widely applied in fields such as Natural Language Processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as to NLP tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios of this application include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0031] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0032] The Manufacturers to Consumer (M2C) business model refers to manufacturers bypassing traditional channels and selling goods directly to end users in order to reduce costs and improve efficiency.
[0033] Artificial Intelligence Generated Content (AIGC) refers to the process of automatically generating multimedia content such as images, text, and videos using artificial intelligence technology.
[0034] Optical Character Recognition (OCR) is a technology that uses image analysis techniques to automatically recognize and extract text information from images.
[0035] A bounding box (Bbox) is a rectangular area used in computer vision to define the position and extent of a target object in an image. It is often used to locate the main body of a product or text elements.
[0036] According to an embodiment of this application, an image processing method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0037] The technical solutions provided in this application can employ deep learning models with relatively large parameter scales, such as large models containing billions or even more model parameters. Here, "large model" is merely an example; this application does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in this application can be artificial intelligence-based language models (LM) or multimodal models (MM).
[0038] Considering the limited computing resources of mobile terminals, the methods described above in this application embodiment can be applied to, for example... Figure 1 The application scenarios shown are not limited to these. Figure 1 This is a scene diagram illustrating an image processing procedure according to an embodiment of this application, in such a case... Figure 1 In the application scenario shown, the deep learning model is deployed on server 10. Server 10 can connect to one or more client devices 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. These client devices 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users through a graphical user interface to invoke the deep learning model, thereby implementing the method provided in this embodiment.
[0039] In this embodiment, the system consisting of a client device and a server can perform the following steps: The client device acquires an image of the target object to be processed. The server performs multi-dimensional detection on the image to be processed to obtain the visual information of the target object. Based on the category of the target object, a reference image corresponding to the image to be processed is determined. Based on the background area of the reference object in the reference image, the background area of the target object in the image to be processed, and the visual information of the target object, the image to be processed is adjusted to obtain the target image of the target object.
[0040] It should be noted that with the rapid development of high-performance computing units, the methods provided in this application embodiment can also be applied to model-in-the-loop machines in other application scenarios. In one optional embodiment, the model-in-the-loop machine has multiple built-in models. Users can select one model to adjust as needed to obtain their own model. The high-performance computing unit built into the model-in-the-loop machine can then directly call the adjusted model to execute the methods provided in this application embodiment. In another optional embodiment, the deep learning model-in-the-loop machine has a pre-trained model built-in. Therefore, the high-performance computing unit built into the model-in-the-loop machine can directly call this model to execute the methods provided in this application embodiment.
[0041] Furthermore, when users need to train their own models, they can upload their own datasets via the client. These datasets are then sent to the server, allowing the server to adjust the pre-trained model using the dataset to obtain the user's customized model, which can then be deployed to the production environment. To facilitate users' model adjustment needs, the server provides complete adjustment tools, development frameworks, and processes, supporting multiple adjustment strategies. This allows the adjusted model to better adapt to different application domains and achieve a high degree of customization.
[0042] Under the aforementioned operating environment, this application provides the following: Figure 2 The image processing method shown. Figure 2 This is a flowchart of an image processing method according to an embodiment of this application. Figure 2 As shown, the method may include the following steps:
[0043] Step S202: In response to receiving the image to be processed of the target object, perform multi-dimensional detection on the image to be processed to obtain the object visual information of the target object.
[0044] Multidimensional detection refers to the process of detecting visual elements associated with the target object in the image to be processed in different dimensions.
[0045] The target object mentioned above is the physical object in the image to be processed that carries the main information. It has a clear shape, structure and visual identity, such as utensils and consumer goods. The boundary, color, attached text and icons of the target object in the image to be processed constitute the main content of information transmission. The original pixel content of the target object needs to be preserved.
[0046] The aforementioned image to be processed contains the original visual input of the target object and its current background. The background may lack consistency with typical visual expressions of the same type, such as color imbalance, scene chaos, or monotonous composition. As the starting point for processing, the value of this image lies in fully preserving the original semantics of the target object, while exposing the room for improvement in the background expression, which can be used for subsequent targeted adjustments based on the reference image.
[0047] The aforementioned object visual information is a set of structured features directly related to the target object, extracted through multi-dimensional visual analysis of the image to be processed. It covers the target object's boundary mask, spatial location, area ratio, surface color distribution, attached text content, icon labels, and relative layout relationships between elements. This object visual information is used to limit the constraints of image editing, ensuring that the integrity and recognizability of the main content are not affected during the background reconstruction process.
[0048] In one optional embodiment, upon receiving the image to be processed, an image segmentation algorithm is first invoked to generate a pixel-level mask for the target object to separate the subject from the background. Secondly, image recognition technology is used to identify embedded text content and promotional labels in the image to be processed, extracting their position and semantics. Then, a multimodal visual coding model is used to analyze the target object's color tendency, texture complexity, and compositional center of gravity, ultimately integrating these features into a unified object visual information vector. This object visual information does not include background semantics; it primarily focuses on the target object's own identifiability and information-carrying capacity, providing a constraint basis for subsequent background adjustments.
[0049] In another optional embodiment, after receiving the image to be processed, a parallel visual analysis process is initiated to extract features of the target object from four dimensions: geometric structure, semantic text, color texture, and spatial layout. Each dimension analysis module runs independently and outputs intermediate results, which are finally merged into unified object visual information. This object visual information does not involve background content and mainly represents the visual composition and semantic integrity of the target object itself.
[0050] For example, in an application scenario, such as an image of a rice cooker, the target object is the rice cooker itself. The image to be processed includes the background of the rice cooker placed on a dark table. Multi-dimensional detection will extract the outline mask of the rice cooker, the "3L capacity" text marked on its surface, the position of the brand logo, the distribution of the highlights of the metallic texture, and the spatial layout of the main body centered to the left, forming the visual information of the object, ensuring that the position, content and shape of these elements remain unchanged when the background is replaced later.
[0051] This application achieves precise constraints on the main content in image editing by systematically and structurally extracting the multidimensional visual attributes of the target object. This allows the background reconstruction process to be carried out without interfering with the core information, thus providing a quantifiable and comparable input basis for subsequent background style alignment and effectively improving the controllability and consistency of image processing.
[0052] Step S204: Based on the category of the target object, determine the reference image corresponding to the image to be processed.
[0053] The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold.
[0054] The aforementioned target object category is an identifier of the semantic classification to which the target object belongs. It is summarized based on the target object's functional attributes, morphological structure, and usage scenario, such as electric kettle, air fryer, and insulated lunch box. This category is used to establish semantic associations between the target object and similar visual expression patterns, and serves as the basis for retrieving and matching reference images.
[0055] The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold. This indicates that the reference object and the target object belong to the same or similar category in terms of physical form, function, or usage scenario. This ensures that the background style presented by the reference image has semantic relevance to the visual expression requirements of the target object, providing a reference basis that conforms to the category characteristics for subsequent background analysis.
[0056] The reference image mentioned above is another image that can be of the same category as the target object and contains the reference object. The background area of the reference image is extracted as a style guide source, reflecting the more typical visual expression pattern in the category. The selection of the reference image is based on category consistency. The main content of the reference image is not edited, but is only used to provide a reference for background scene, color preference and spatial composition, as a benchmark for background semantic alignment.
[0057] The aforementioned reference objects are entities in the reference image that belong to the same category as the target object. The existence of reference objects ensures that the background semantics of the reference image are reasonably related to the functional scenario of the target object. The reference objects themselves do not participate in image editing, but only serve as semantic anchors for background style analysis, so that the extraction of background patterns has intra-class consistency and avoids style mismatch due to subject differences.
[0058] The target image described above, while preserving the target object and its visual information, mainly generates an output image by semantic alignment and visual style transfer of the background area. The background content of the target image is similar in style to the reference image, but the main body does not undergo any deformation or content addition or deletion, thus improving the consistency of visual expression without compromising the original information carrying capacity.
[0059] In one optional embodiment, after obtaining the category identifier of the target object, all images containing the object of that category are retrieved from a pre-built visual sample library. The images are sorted and filtered according to the visual consistency, structural typicality and semantic coverage of the background area in the sample images. Finally, one or more reference images are selected as style reference sources. The selection criteria are based on category matching degree first, without depending on the image source, acquisition time or shooting conditions.
[0060] In one optional embodiment, a category alignment mechanism can be established between the target object and the background style. Unlike methods that rely solely on image content similarity for retrieval or manually preset templates, this application uses category constraints to ensure that the background semantics of the reference image always matches the functional scenario of the target object. For example, the background of an electric kettle should be associated with usage scenarios such as a kitchen, desktop, water vapor, and heat, rather than abstract color blocks or irrelevant indoor environments. The existence of a reference object provides a semantic basis for background style extraction, avoiding style mismatches caused by inconsistencies between the target object and the subject of the reference image. This process does not rely on subjective design but rather on unbiased selection based on the natural distribution patterns of images in that category within an objective sample library.
[0061] For example, the target object in the image to be processed can be a rice cooker, categorized as a rice cooker. Based on this category, all images containing rice cookers can be retrieved from a visual sample library, and images showing typical usage scenarios such as a light-colored kitchen countertop, cooked rice, steaming, and natural light can be selected as reference images. The reference object in this reference image is another rice cooker, whose shape, brand, and size may differ from the target object, but the category is the same. The everyday cooking scene presented in its background becomes the semantic basis for reconstructing the background of the image to be processed. If the target object is a thermos, the reference image will be selected from images showing the thermos placed on a desk, next to a coffee cup, or against a book background, ensuring that the background style is associated with office use and portable drinking scenarios.
[0062] By using the target object category as the sole retrieval criterion and combining it with reference objects of the same category in the reference image, semantic alignment of the background style is achieved. This ensures that the background expression of the reference image has category relevance and scene rationality, thereby providing a natural reference that conforms to human visual cognition for subsequent background replacement and avoiding the sense of incongruity caused by the background style being out of context.
[0063] Step S206: Based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the visual information of the object, adjust the image to be processed to obtain the target image of the target object.
[0064] The aforementioned background area is the remaining portion of the reference image after removing the pixels occupied by the target object or reference object, used to carry environmental semantics, spatial atmosphere, and visual contrast functions; the background area can be presented as a solid color, texture, artificial scene, or natural environment, and its visual attributes include hue, light and shadow, compositional hierarchy, and atmospheric tendency; in the embodiments of this application, the background area is a part that can be replaced by content and transferred by style, and its modification needs to be limited to the non-subject area defined by the boundary of the target object.
[0065] The background area of the reference object in the above reference image refers to the environmental part remaining after removing the pixels of the reference object. This background area carries the typical functional scene visual expression of this type of object, such as light, texture, spatial composition and color tendency. The content of the background area does not include the reference object itself, and mainly serves as a source of style guidance.
[0066] The background region of the target object in the image to be processed refers to the environmental part remaining after removing the pixels of the target object from the image to be processed. The current form of the background region may not match the typical visual expression associated with the target object category, but the spatial range of the background region corresponds to the boundary outline of the target object, constituting the original background region that needs to be replaced.
[0067] In one optional embodiment, the background area of the reference object in the reference image can be used as the style source, the background area of the target object in the image to be processed can be used as the replacement area, and the visual information of the object can be used as the unchangeable editing boundary. The content of the style source area can be migrated to the replacement area through image compositing. The migration process can be based on the outline of the target object defined by the visual information of the object to ensure that the pixels of the target object are not modified, covered or deformed. The main process is to replace the content and align the style of the background area, and finally output the target image.
[0068] This application achieves semantic transfer of the background region while maintaining the integrity of the target object. The background region of a reference object can be used as a style sample, the boundary mask of the target object as an editing mask, and the background region of the target object as a spatial carrier; these three elements together constitute the input conditions for image adjustment. This process does not generate new subjects or modify text or logos. It primarily uses pixel-level region replacement and tone blending to transition the background of the image to be processed from its original state to a state consistent with the expression of similar objects, achieving a natural convergence of visual expression.
[0069] For example, the target object in the image to be processed can be identified as a rice cooker. Its visual information includes the complete outline of the rice cooker, the "5L" lettering on its surface, the brand logo, and its centered position. The background area is a dark, matte tabletop. The reference object in the reference image is another rice cooker, whose background area is a light-colored wooden countertop with slight light reflection, and a bowl of rice placed on its left. Based on the outline mask in the object's visual information, the texture, lighting, and color structure of the background area in the reference image can be accurately mapped to the space occupied by the rice cooker's background in the image to be processed. The original pixels of the rice cooker, its lettering, and its logo are preserved, and only the tabletop area below it is replaced. This ensures that the rice cooker remains the main subject in the final target image, while the background presents a typical everyday usage scenario consistent with this type of product.
[0070] This application achieves a natural update of the background style of the image to be processed by referencing the background through semantic transfer without changing any visual content of the target object. This allows the target image to maintain information integrity while improving its consistency with typical visual expressions of the same type, thereby enhancing the rationality and coordination of visual communication.
[0071] Upon receiving the image of the target object to be processed, the system first performs multi-dimensional detection on the image, extracting visual information of the target object in multiple visual dimensions such as boundaries, text, icons, and spatial layout, ensuring that the main content is completely identified and unaffected by subsequent processing. Then, based on the category of the target object, reference images containing reference objects of the same category are selected from the visual samples to match the background style with the functional scene of the target object. Finally, using the background area of the reference object in the reference image as the style source and the background area of the target object in the image to be processed as the replacement space, combined with the main boundary defined by the object's visual information, the background content of the original image is transferred and fused to generate the target image. This allows the background to present a natural expression consistent with similar objects without changing any visual information of the target object, thus improving visual harmony.
[0072] For example, suppose the image to be processed is a picture of a rice cooker with a cluttered white tabletop as the background, dim lighting, and a lack of scene. First, the image undergoes multi-dimensional detection to identify the precise outline of the rice cooker, the "5L" marking on its surface, the position of the brand logo, and the spatial distribution of the main body centered and slightly lower, forming complete object visual information. Then, based on the category of "rice cooker," a reference image containing similar objects is selected from a pre-stored image set. This reference image features another rice cooker with a light-colored wooden countertop as its background, a bowl of rice on the right, and soft highlights from natural light at the top, creating a background area with a realistic, everyday usage scenario. Finally, using the rice cooker outline from the object visual information as a boundary mask, the system transfers the texture, lighting, and color structure of this background area from the reference image to the original tabletop area in the image to be processed. This ensures that the pixels of the rice cooker itself, the text, and the logo remain unchanged, only replacing the background below it. This results in the final target image where the rice cooker remains the main subject, but the background presents a typical usage environment consistent with this type of product, making the overall visual effect more natural and harmonious.
[0073] Through the above steps, in response to receiving the image of the target object to be processed, multi-dimensional detection is performed on the image to be processed to obtain the visual information of the target object. Multi-dimensional detection refers to the process of detecting visual elements associated with the target object in the image to be processed in different dimensions. Based on the category of the target object, a reference image corresponding to the image to be processed is determined. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold. Based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the visual information of the object, the image to be processed is adjusted to obtain the target image of the target object. By obtaining the visual information of the target object through multi-dimensional detection, and combining it with the background features of the reference object in the reference image, the semantic differences between the background of the image to be processed and the reference image are compared. This infers the deficiencies of the background of the image to be processed in multiple dimensions such as scene adaptability, visual coordination, and style expressiveness. Based on these differences and the object visual information, a targeted improvement strategy is generated to adjust the background region, ensuring that the visual information of the target object remains unchanged, thereby improving the overall visual expressiveness of the image and solving the technical problem of poor image processing effects in related technologies.
[0074] In the above embodiments of this application, multi-dimensional detection is performed on the image to be processed to obtain the object visual information of the target object, including: detecting the image to be processed to obtain the object region mask and visual description information of the target object, wherein the object region mask is used to represent the mask of the region where the target object is located in the image to be processed, and the visual description information is used to represent the description information of the target object; and determining the object visual information based on the object region mask and the visual description information.
[0075] The object region mask mentioned above is a two-dimensional matrix composed of pixel-level binary or soft masks. Each pixel indicates whether its corresponding position belongs to the coverage area of the target object. It is used to accurately delineate the spatial boundary of the target object in the image, ensuring that subsequent processing only affects the background area while the main body remains unchanged.
[0076] The aforementioned visual description information is a semantic textual or structured feature expression of the visual elements contained in the target object. The visual description information can be the text content, icons, brand information, promotional labels, color distribution and spatial position relationship attached to the image to be processed, reflecting the non-geometric visual semantics carried by the target object.
[0077] The aforementioned object visual information is a composite expression composed of object region mask and visual description information. The function of object visual information is to fully represent the spatial location and semantic content of the target object. As a core constraint in the image adjustment process, it ensures that the main content is not disturbed, covered or deformed in any processing.
[0078] In one optional embodiment, the pixel boundaries of the target object are first identified by an image segmentation model to generate an accurate object region mask. At the same time, text, icons and layout relationships associated with the target object in the image are extracted by an optical character recognition and multimodal semantic analysis model to form visual description information. Subsequently, the two are structurally aligned to construct unified object visual information, enabling it to have both spatial positioning and semantic recognition capabilities.
[0079] The object region mask provides geometric constraints, ensuring that editing operations are performed only in the background area; visual description information provides content semantics, identifying which elements cannot be changed, such as text, logos, or promotional labels. The object region mask and visual description information combine to form the object's visual information, serving as a protective boundary and content list for subsequent image adjustments, preventing the loss or misalignment of key information due to background replacement. This process does not use preset templates or rely on manual annotation; it automatically extracts information based on image content, achieving a complete understanding of the target object.
[0080] When detecting objects in the image to be processed, the first step is to use an image segmentation model to perform pixel-level recognition of the outline of the target object in the image, generating an accurate object region mask. This mask records the pixel affiliation of the target object in the image in the form of a binary matrix, ensuring that its boundaries are clear and unbroken. At the same time, the surface and surrounding areas of the target object are parsed using an optical character recognition and multimodal visual semantic analysis model, extracting the text content, icons, brand marks, promotional labels, and their relative spatial positions in the image to form structured visual description information. Subsequently, the object region mask and the visual description information are spatially aligned and semantically bound, mapping the position coordinates of the text and icons to the main area defined by the object region mask, thereby constructing complete and indivisible object visual information. This information carries both the geometric boundaries and semantic content of the target object, serving as a dual constraint for subsequent background adjustment, ensuring that the background area can be independently replaced and style transferred without changing any pixel and text distribution of the target object in the initial enhanced image.
[0081] For example, the image to be processed can be a picture of a thermos cup. The system generates a precise contour mask of the thermos cup using a segmentation model. Simultaneously, it identifies the word "stainless steel" on the cup through image recognition and detects the "limited-time free cup sleeve" icon attached to its left using a visual understanding model, determining that it is located in the upper right corner of the image. These elements together constitute visual descriptive information. Subsequently, the system binds the mask with the position coordinates of the aforementioned text and icon to form object visual information, clearly indicating that the thermos cup itself is not modified, and that "stainless steel" and "limited-time free cup sleeve" are completely preserved in their original positions.
[0082] By collaboratively constructing object region masks and visual description information, dual locking of the spatial boundaries and content elements of the target object is achieved, providing a precise and robust constraint basis for subsequent background replacement, and ensuring that the subject information remains intact, unchanged in position, and semantically clear during image processing.
[0083] In the above embodiments of this application, detecting the image to be processed to obtain the object region mask and visual description information of the target object includes: performing foreground segmentation on the image to be processed to obtain the object region mask, wherein the foreground segmentation is used to represent the extraction of pixels other than the background region in the image to be processed; and performing content recognition on the image to be processed to obtain visual description information.
[0084] The foreground segmentation described above is an image analysis method based on pixel-level classification. The role of foreground segmentation is to distinguish the non-background region occupied by the target object from the image to be processed. The output is a binary or probability mask with the same size as the image, which is used to accurately identify the position of each pixel of the target object, eliminate background interference, and form an independent subject region.
[0085] The aforementioned content recognition refers to the technical process of detecting and analyzing visual symbols and semantic elements in images. Through optical character recognition and visual semantic understanding models, text, icons, labels, brand logos and their distribution locations that are directly related to the target object are extracted from the image to form unstructured or structured semantic expressions.
[0086] The aforementioned visual description information is a set of textual or structured descriptions generated by content recognition, which includes the identifiable content visible on the target object and its spatial relationship in the image. It is used to define which elements must remain as is and cannot be covered or moved.
[0087] In one optional embodiment, an image segmentation model is first invoked to perform pixel-level classification on the input image to be processed, extracting regions that do not belong to the background as a continuous set of foreground pixels to form an accurate object region mask; subsequently, the system invokes a content recognition model to detect text, icons, and marking information on the surface of the target object and its adjacent regions within the entire range of the image to be processed, identify their semantic content and coordinate positions, and form visual description information.
[0088] By executing two parallel processing steps—foreground segmentation and content recognition—the target object can be deconstructed from both geometric and semantic dimensions. Foreground segmentation provides spatial boundaries indicating "where it cannot be moved," while content recognition provides a semantic list indicating "what cannot be changed." These two processes complement each other, forming a dual protection mechanism for subsequent image editing. This process does not rely on manual rules or preset templates; instead, it is automatically derived from image data by the model, possessing universal adaptability to arbitrary shapes and labeled content.
[0089] In one optional embodiment, when performing foreground segmentation on the image to be processed, a pixel-level classification model is used to determine the attribution of each pixel in the image. The non-background area covered by the target object is extracted into a continuous set of pixels, generating an object region mask with the same size as the original image. This object region mask identifies the precise contour boundary of the target object in binary form, ensuring that the background area and the subject area are separated at the pixel level. Simultaneously, when performing content recognition on the image to be processed, an optical character recognition and visual semantic analysis model is used to detect and semantically parse visual elements such as text, icons, and labels on the surface and surrounding areas of the target object, extracting its content text, graphic identifiers, and their spatial coordinates in the image, forming structured visual description information. The two processing paths are executed in parallel. The former provides spatial boundary constraints, while the latter provides a content semantic list, which together constitute a complete expression of the target object. This ensures that the area defined by the object region mask is not modified during the subsequent background replacement process, and that every piece of text and icon contained in the visual description information remains unchanged, thereby achieving independent processing of the subject content and the background area and avoiding the covering or offset of key information.
[0090] For example, the image to be processed can be a picture of a rice cooker, with the rice cooker itself located in the center of the image against a dark desktop background. The outline of the rice cooker can be precisely separated using a foreground segmentation model, generating an object region mask containing all its pixels. Simultaneously, a content recognition model detects the "5L capacity" text on the front of the rice cooker, the "Smart Reservation" icon on the side, and the "Stainless Steel Inner Pot" label on the bottom, recording their coordinate positions in the image to form visual descriptive information. This mask and descriptive information together constitute a complete representation of the target object, ensuring that the pixels of the rice cooker itself are not modified during subsequent background replacement, and that all text and icons remain in their original positions.
[0091] By combining the object region mask generated by foreground segmentation with the visual description information generated by content recognition, the spatial boundaries and content elements of the target object can be independently extracted and jointly constrained. This provides a precise and robust protection mechanism for subsequent background replacement without manual intervention, effectively preventing the main content from being deformed, occluded, or misaligned during image adjustment.
[0092] In the above embodiments of this application, the target image of the target object is adjusted based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the object visual information, including: comparing the background region of the reference object in the reference image and the background region of the target object in the image to be processed to obtain a background comparison result; and adjusting the image to be processed based on the background comparison result and the object visual information to obtain the target image.
[0093] The background area of the aforementioned reference object is the image area remaining after removing the main body of the reference object in the reference image. It includes the visual characteristics of the area, such as color distribution, texture structure, light and shadow form, and spatial layout, reflecting the typical presentation of this type of target object against this background.
[0094] The background area of the target object in the above-mentioned image to be processed is the background part retained after removing the main body of the target object in the image to be processed. Its visual characteristics may differ from the background in the reference image, such as color imbalance, monotonous scene, loose composition, etc.
[0095] The background comparison results mentioned above are quantitative or semantic differences generated by comparing the background regions of the reference object and the background regions of the target object in the image to be processed. These differences include comparison conclusions in dimensions such as tonal preference differences, scene type matching degree, spatial layout rationality, and visual attractiveness.
[0096] The aforementioned object visual information is a composite representation composed of object region mask and visual description information, used to clarify the spatial boundaries and content elements of the target object, ensuring that the main content is not disturbed during image adjustment.
[0097] In one optional embodiment, the background regions of the reference object in the reference image and the target object in the image to be processed are first extracted, and their color histograms, texture gradient distributions, semantic scene categories, and spatial composition patterns are extracted respectively. Then, the two are compared item by item in multiple visual dimensions to generate a background comparison result reflecting the degree of difference between the two. Based on this, combined with the constraints on the position and content elements of the subject in the object's visual information, the image editing model is driven to adjust only the background region in the target image, so that its visual features approach the reference background, while keeping the subject content unchanged, and the final target image is output.
[0098] By introducing visual patterns from external reference samples through a background contrast mechanism, background adjustments no longer rely on manual rules or generic templates, but are guided by existing visual paradigms among similar objects. The background contrast results provide adjustment direction, while the object's visual information provides adjustment boundaries. Together, they limit the scope of disturbance in the image editing process, allowing only semantically consistent and stylistically harmonious changes to the background without affecting the main structure, text labels, or layout relationships. This process does not preset fixed styles or enforce uniform templates; instead, it dynamically matches reasonable features of the reference background to achieve natural, controllable, and semantically consistent background rewriting.
[0099] For example, the image to be processed can be a picture of a rice cooker, with a solid gray tabletop as the background and no scene elements; the reference image is an image of a rice cooker of the same type, with a light-colored wooden countertop as the background, a bowl of rice placed on the right, and soft highlights formed by natural light at the top. The system extracts the background area of the target object in the image to be processed, identifying it as a single tone, without texture, and without composition; it extracts the background area of the reference object, identifying it as a typical kitchen scene, with warm colors, food accompaniment, and light and shadow layers; the comparison results of the generated backgrounds show that the background of the image to be processed is significantly lacking in scene richness, color layering, and atmosphere guidance; combining the outline mask of the rice cooker and the positions of text such as "5L capacity" and "stainless steel inner pot" in the object's visual information, the system only performs pixel-level redrawing of the background area, replacing the original gray tabletop with a wooden countertop and food accompaniment consistent with the reference image, while ensuring that the rice cooker itself and all text labels retain their original positions and shapes, thus generating the target image.
[0100] By using background contrast results to guide visual adjustments in the background area, the background style of the image to be processed is made to align with a more expressive reference background among similar objects. At the same time, the visual information of the object is used to ensure that the main content is not destroyed during the adjustment process, thereby improving the overall visual harmony and scene immersion of the image without changing the target object itself.
[0101] When comparing the background regions of the reference object in the reference image with the background regions of the target object in the image to be processed, we can extract their visual expressions in terms of color distribution, texture features, spatial composition, and scene semantics. By comparing each item, we can generate background comparison results that reflect the degree of difference, and identify the deficiencies of the background of the image to be processed in terms of style consistency, visual appeal, and scene matching. Based on this, we combine the subject boundary defined by the object region mask in the object's visual information with the content that does not need to be modified identified by the visual description information, and drive the image editing model to redraw the background region of the target image pixel by pixel, so that the background style is closer to the reference background. At the same time, we constrain the pixel position of the subject and the text labels to prevent them from shifting or being occluded. The final output is a target image with a background style that is closer to the same type of reference sample while preserving the integrity of the original subject.
[0102] In the above embodiments of this application, based on a reference image, background comparison is performed on the background region of the target object in the image to be processed to obtain a background comparison result, including: performing background recognition on the background region in the image to be processed to obtain first background information of the image to be processed, wherein the first background information is used to represent the visual state of the background region in the image to be processed; performing background recognition on the background region in the reference image to obtain second background information of the reference image, wherein the second background information is used to represent the visual state of the background region in the reference image; and performing background comparison on the image to be processed based on the first background information and the second background information to obtain a background comparison result.
[0103] The aforementioned background area is the set of pixels remaining after removing the main part of the target object from the image. Its visual state is composed of color distribution, texture structure, light and shadow form, spatial layout and semantic scene elements, which are used to carry the presentation environment of the target object.
[0104] The aforementioned background recognition is a technical process based on multimodal visual semantic analysis. It uses an image understanding model to perform semantic parsing on a specified region, extracting its structured visual expression in dimensions such as color tone, scene category, composition method, and atmosphere features, forming comparable descriptive information.
[0105] The aforementioned first background information is a structured visual description generated by background recognition of the background region in the image to be processed. It includes the background's tonal characteristics, scene type, layout pattern, and overall atmosphere, and is used to characterize the current state of the image background.
[0106] The aforementioned second background information is a structured visual description generated by background recognition of the background region in the reference image, reflecting the typical visual pattern of the reference background and serving as a reference style observed among similar objects.
[0107] The background comparison results mentioned above are generated by comparing the first background information and the second background information item by item in multiple visual dimensions. The difference judgment is used to quantify the relative relationship between the background of the image to be processed and the reference background in terms of style consistency, scene matching degree and visual expressiveness.
[0108] In one optional embodiment, the non-subject area of the target object in the image to be processed is first located as the background area. A visual semantic analysis model is invoked to identify its color tone, scene type, spatial composition, and atmosphere features to generate first background information. At the same time, the same processing is performed on the background area of the reference object in the reference image to generate second background information. Subsequently, the two are cross-matched in dimensions such as tone preference, scene category, composition complexity, and atmosphere appeal to output the background comparison results expressed in text or structured feature form, clarifying whether the background of the image to be processed deviates from the typical pattern of similar reference samples.
[0109] This application employs an independent and symmetrical background recognition mechanism to extract comparable background semantic expressions from both the image to be processed and the reference image, avoiding the limitations of direct pixel comparison and instead focusing on the alignment of higher-order visual semantics. The comparison process does not rely on preset templates but rather evaluates differences based on visual patterns derived from real samples, making the judgment results both category-adaptive and data-driven. This mechanism provides clear and interpretable directional guidance for subsequent background adjustments, ensuring that the adjustments are not random modifications but are guided by reasonable visual expressions already existing in similar objects.
[0110] For example, the image to be processed can be a display image of a rice cooker with a pure white flat background and no scene elements. The first background information is described as "light-colored, textureless background, no scene association, and no light and shadow hierarchy". The reference image is a display image of a rice cooker of the same type with a light wood-grain tabletop, a bowl of rice placed on the right, and soft side lighting at the top. The second background information is described as "warm-colored wood texture, a lifelike kitchen scene, food as a backdrop, and light and shadow transition". After comparing the two, the generated background comparison results indicate that the image to be processed is weaker than the reference background in terms of scene richness, color hierarchy, and atmosphere guidance, and lacks visual cues of the usage context.
[0111] By obtaining structured visual descriptions of the background regions in the image to be processed and the reference image through background recognition, and forming a quantifiable background difference judgment based on the semantic dimension comparison between the two, we can provide an objective guidance basis based on similar samples for subsequent image adjustment, and avoid the blindness and subjectivity of background modification.
[0112] When identifying the background region in the image to be processed, a visual semantic analysis model is used to extract the color tendency, scene category, spatial layout, and atmosphere features of the image to be processed, forming the first background information representing the current visual state of the background. The same processing is performed on the background region in the reference image to obtain the corresponding second background information, which reflects the typical background patterns observed in similar objects. Subsequently, the first background information and the second background information are compared in a structured manner in terms of color harmony, scene relevance, compositional richness, and visual appeal, generating a background comparison result that reflects the degree of difference between the two. This clarifies the direction of deviation of the background of the image to be processed from the reference sample in terms of style expression, thereby providing an objective basis based on visual semantics for subsequent adjustments and ensuring that background modification does not deviate from the reasonable expression rules of similar images.
[0113] In the above embodiments of this application, a background comparison is performed on the image to be processed based on the first background information and the second background information to obtain a background comparison result. This includes: performing a multi-dimensional evaluation on the image to be processed and the reference image based on the first background information and the second background information to obtain multiple evaluation results. The multi-dimensional evaluation includes at least one of the following: scene matching degree, tone coordination degree, visual attention degree, and element coordination degree of the background regions in the image to be processed and the reference image; and determining the background comparison result based on the multiple evaluation results.
[0114] The scene matching degree mentioned above refers to the degree of semantic consistency between the environment type presented by the background area of the image to be processed and the typical usage scenario reflected by the background area of the reference image. The scene matching degree reflects whether the background conforms to the common display scenarios of similar objects.
[0115] The aforementioned color harmony refers to the degree of visual similarity between the color composition (such as brightness, warmth, and saturation) of the background area of the image to be processed and the color preference of the background area of the reference image, in order to determine whether there is a color conflict or style discrepancy between the two.
[0116] The aforementioned visual attention refers to the ability of a background area in an image to guide the viewer's attention. It is assessed by characteristics such as texture complexity, light and shadow levels, and element distribution to determine whether it has the function of highlighting the subject and enhancing visual guidance.
[0117] The aforementioned element coordination refers to whether there are auxiliary elements in the background that have a functional or semantic connection with the main object (such as tableware, plants, lights, etc.). Their presence or absence, whether their position is reasonable, and whether their form is natural all affect the integrity and sense of life of the overall picture.
[0118] The background comparison results described above are structured judgments generated by combining multiple evaluation results, used to characterize the relative superiority or inferiority of the background of the image to be processed relative to the background of the reference image in terms of multi-dimensional visual performance.
[0119] In one optional embodiment, the first background information and the second background information can be used as inputs to perform independent scene semantic classification, color statistical analysis, attention heat modeling, and auxiliary element detection on the background regions of the two images respectively. The numerical or grade differences between the image to be processed and the reference image in four dimensions are calculated: scene matching degree, color tone coordination degree, visual attention degree, and element coordination degree. The evaluation of each dimension is based on preset semantic rules or statistical models, such as whether the scene type belongs to "living kitchen", whether the main color tone is warm light color, and whether there is food accompaniment. Finally, the evaluation results of the four dimensions are weighted and integrated according to preset weights to form a background comparison result that comprehensively reflects the background differences.
[0120] By breaking down background differences into multiple quantifiable and interpretable visual dimensions, comparison behavior no longer relies on vague intuition or single features, but is based on systematic semantic evaluation. Each dimension corresponds to a key factor in real images that affects user perception, ensuring that the evaluation results are close to human visual habits. This mechanism does not presuppose a single standard, but allows for dynamic selection of evaluation items based on category characteristics. For example, high-priced products pay more attention to element coordination, while fast-moving consumer goods emphasize visual attention, thereby enhancing the adaptability and rationality of the judgment.
[0121] For example, the image to be processed can be a display image of a rice cooker. The background of the image to be processed is a solid gray tabletop. The first background information is described as "no specific scene, neutral cool gray, no auxiliary elements, and no light and shadow guidance." The reference image is a display image of a rice cooker of the same type. Its background is a light wood grain tabletop, with a rice bowl on the right and soft side lighting at the top. The second background information is described as "a lifelike kitchen scene, warm light colors, food accompaniment, and light and shadow layers." The system evaluation results show that: the scene matching degree is low (the usage environment is not reflected), the color tone coordination degree is low (cool gray is not coordinated with the warm main object), the visual attention is low (no layer guidance), and the element coordination degree is zero (no related elements). Based on the four evaluation results, the background comparison results determine that the background of the image to be processed is weaker than the reference image in all four aspects and needs to be adjusted.
[0122] Through a multi-dimensional evaluation mechanism, background differences are decomposed into independently measurable visual attributes such as scene matching degree, color tone coordination degree, visual attention degree, and element coordination degree. Based on the comprehensive judgment of multiple indicators, background comparison results are generated, making the judgment of background quality semantically interpretable and category-appropriate, and avoiding adjustment deviations caused by misjudgment of a single feature.
[0123] Based on the first and second background information, the background regions of the image to be processed and the reference image can be independently evaluated from four dimensions: scene matching degree, tone coordination degree, visual attention degree, and element coordination degree. Scene matching degree is judged by comparing whether the environment type presented by the background belongs to the common usage scenario of the same object. Tone coordination degree is evaluated based on whether the color tendencies are similar in terms of warmth, coolness, brightness, and saturation. Visual attention degree is analyzed by the ability to guide attention through texture distribution and light and shadow layers. Element coordination degree detects whether there are auxiliary elements in the background related to the main function and the rationality of their layout. The evaluation results of the four dimensions are summarized and weighted according to preset rules to generate a background comparison result that reflects the relative difference in visual performance between the background of the image to be processed and the background of the reference image. This transforms the determination of background differences from subjective perception to multi-dimensional semantic alignment, ensuring that subsequent adjustments are based on real visual rules rather than random modifications.
[0124] In the above embodiments of this application, the background comparison result is used to indicate whether the background of the image to be processed needs to be adjusted. The adjustment of the image to be processed based on the background comparison result and the object visual information to obtain the target image of the target object includes: if the background comparison result indicates that the background of the image to be processed needs to be adjusted, generating a background adjustment strategy based on the second background information and the object visual information, and adjusting the image to be processed based on the background adjustment strategy to obtain the target image; if the background comparison result indicates that the background of the image to be processed does not need to be adjusted, determining the image to be processed as the target image.
[0125] The aforementioned background adjustment strategy is a structured editing instruction generated by the second background information and the object's visual information. It includes the scene type, tone tendency, spatial layout and auxiliary element configuration requirements of the target background, while clearly defining the constraints on the preservation of the main body area, which is used to guide the direction of content changes during the image editing process.
[0126] The aforementioned object visual information is a set of non-background constraints consisting of the main body region mask of the target object in the image to be processed, the position of the selling point text, the position of the brand logo, and the distribution of key visual elements. It is used to ensure that the integrity and recognizability of the main content are not destroyed during the image adjustment process.
[0127] When the background comparison result indicates that adjustment is required, the typical visual patterns contained in the second background information of the reference image can be semantically aligned with the immutable areas defined by the visual information of the object in the image to be processed. This generates a specific and executable background adjustment strategy, which includes the environmental type, color range, composition method, and whether auxiliary elements should be introduced in the target scene. Subsequently, the image editing model, based on this strategy and under the constraint of the object's visual information, mainly performs pixel-level redrawing of the background area to ensure that the main outline, text labels, and key details are not disturbed. If the background comparison result indicates that no adjustment is required, the original image to be processed is used as the final output without triggering any editing actions.
[0128] Differentiated processing is achieved through a conditional branching mechanism, avoiding indiscriminate editing of all images and improving processing efficiency and the rationality of results. When adjustments are required, the background adjustment strategy is guided by the semantic features of the reference background and bounded by the visual information of the object, achieving precise editing that is "directional and retains" so that background changes do not damage the information carrying capacity of the original image. This mechanism does not rely on manually preset templates, but dynamically generates customized editing instructions based on comparison results and object features, enhancing the generalization ability of images of different categories and original quality.
[0129] For example, the image to be processed can be a display image of an air fryer with a pure white background. The first background information shows "no scene, neutral tone, no auxiliary elements," and the background comparison results indicate that adjustment is needed. The second background information of the reference image is "light-colored kitchen countertop, food being cooked on the right, and natural light and shadow on the top." The object's visual information includes the outline of the air fryer, the words "15-minute quick cooking" on the front, and the location of the brand logo. Based on this, the system generates a background adjustment strategy: "Replace the background with a light wood grain tabletop, add an image of steaming food on the right, retain the original subject and text, and enhance the consistency of light and shadow." The image editing model then only redraws the background area to generate the target image. If the background of another image is already a similar everyday scene, the background comparison result indicates that no adjustment is needed, and the original image is directly output.
[0130] The system can determine whether to make adjustments based on the background comparison results. When adjustments are needed, it can combine the semantics of the reference background with the visual constraints of the object to generate a customized editing strategy. This ensures that background modifications are only made in non-subject areas and conform to visual rules, thereby achieving a reasonable evolution of the background style while maintaining the integrity of the subject and avoiding ineffective or destructive editing.
[0131] If the background comparison result indicates that background adjustment of the image to be processed is required, a background adjustment strategy containing the target background shape and editing constraints can be generated based on the scene type, tone tendency, spatial layout and auxiliary element features contained in the second background information of the reference image, combined with the main outline, selling point text position and brand logo range defined by the visual information of the object in the image to be processed. According to this strategy, only the background area is redrawn at the pixel level to ensure that the main content is preserved, thereby obtaining the target image. If the background comparison result indicates that no adjustment is required, the image to be processed can be directly output as the target image without triggering any editing operations. This mechanism avoids unnecessary processing through conditional judgment, ensuring that the editing behavior is always subject to the reference visual mode and subject protection constraints, thereby improving the targeting and information integrity of image processing.
[0132] In the above embodiments of this application, adjusting the image to be processed based on the background adjustment strategy to obtain the target image includes: using the object's visual information as an adjustment constraint and the background adjustment strategy as a guiding condition to adjust the background area in the image to be processed to obtain the target image. The adjustment constraint is used to keep the object's visual information in the image to be processed unchanged during the adjustment process, and the guiding condition is used to determine the adjustment direction of the background area during the adjustment process.
[0133] The aforementioned background adjustment strategy is a structured editing instruction derived from the background semantic features of the reference image. It includes the scene type, color distribution, spatial composition, and auxiliary element configuration requirements that the target background should present. It is used to clarify what visual form the background area should be changed to, but does not involve the intention to modify the main content.
[0134] In one optional embodiment, when performing image adjustments, the object's visual information is input into the image editing model as a spatial mask, limiting the editing operation to the background area outside the mask. At the same time, the background adjustment strategy is converted into a semantic guidance signal, which serves as a conditional input to the editing model, instructing that the original background be reconstructed into a new form that conforms to the semantic features of the reference background while preserving the subject. The image editing process is carried out simultaneously based on the dual mechanisms of region isolation and semantic guidance, ensuring that the morphological changes of the background strictly follow the strategy direction and that the subject content is constrained and protected.
[0135] By decoupling the editing behavior into two independent but collaborative control dimensions—constraint and guidance—background modifications become both directional and secure. Object visual information serves as a hard boundary, eliminating distortion problems such as subject deformation and text overlay found in traditional generation methods. Background adjustment strategies act as a soft guide, ensuring that background changes no longer rely on manual templates but dynamically match the visual patterns of high-performing samples. The combination of these two approaches achieves precise control over "changing the background without changing the subject," avoiding content distortion and style deviation.
[0136] For example, the image to be processed can be a rice cooker, whose visual information includes the main outline, the words "Smart Reservation" on the front, and the brand logo; the background of the reference image is a light wood grain countertop with steaming rice and natural side lighting; the system generates a background adjustment strategy of "replacing it with a warm-toned wood countertop, adding food and slight light and shadow layers on the right side"; the image editing model locks the rice cooker body and text area based on the mask in the object's visual information, and only applies the strategy guidance to the background area to complete the background redraw; in the final target image, the rice cooker and its text logo remain completely unchanged, and the background presents a visual form that is closer to a real-life scene.
[0137] By using the object's visual information as an adjustment constraint and the background adjustment strategy as a guiding condition, directional editing of the background area and mandatory protection of the main content are achieved. This ensures that image modifications only occur in non-critical areas, and that the direction of modification originates from the visual patterns of the reference sample. In this way, the visual suitability and expressive consistency of the background are improved while maintaining the integrity of the original information.
[0138] The aforementioned adjustment constraints refer to the set of non-background areas defined by the main body area mask of the target object in the image to be processed, the position of the selling point text, the distribution of the brand logo, and key visual elements. The purpose of the adjustment constraints is to define the unchangeable visual range during the image editing process, ensuring that the main content is not modified, obscured, or deformed.
[0139] The aforementioned guiding conditions refer to structured editing instructions generated based on the semantic features of the background of the reference image. These instructions include the target scene type, color tendency, spatial layout, and auxiliary element configuration requirements, and are used to indicate the visual form that the background area should be reconstructed to.
[0140] During image adjustment, the visual information of the object can be used as a boundary mask to restrict the editing operation to the background area. At the same time, the background adjustment strategy is used as a semantic guide input to the image editing model to drive the background pixels to be reconstructed into a shape that meets the guiding conditions without going out of bounds, thereby achieving precise processing of background updating only and subject preservation.
[0141] In the above embodiments of this application, adjusting the background region in the image to be processed to obtain the target image includes: adjusting the background region to obtain the adjusted region; and performing edge fusion of the object visual information and the adjusted region to obtain the target image.
[0142] The aforementioned adjustment area is a set of pixels generated after semantically guided image editing of the background area. Its shape and color conform to the scene type, tone tendency and spatial layout requirements set by the background adjustment strategy, but it has not yet formed visual coherence with the visual information of the object.
[0143] In one alternative embodiment, firstly, within the constraints defined by the object's visual information, the background area is semantically guided pixel redrawing to generate an adjustment area that conforms to the style of the reference background. Subsequently, pixel-level edge blending is performed at the boundary between the adjustment area and the object's visual information. Through hue gradient transition, light and shadow consistency calibration, and edge sharpness smoothing, a natural connection is formed between the adjustment area and the original subject, eliminating splicing marks and finally outputting a visually coherent target image.
[0144] After the background area content is replaced, this application introduces an edge blending step to make the visual information of the adjusted area and the object visually continuous at the pixel level, avoiding distortion of the subject's edge, color difference breakage, or light and shadow conflict caused by abrupt changes in the background. This process does not rely on manual intervention and automatically calculates the transition weight based on the pixel neighborhood relationship to ensure that the subject outline and background texture remain coordinated in terms of color gradient, light and dark direction, and edge blur.
[0145] By edge-blending the visual information of the adjusted area and the object, a natural connection between the subject and the new background at the pixel level is achieved after the background is modified, eliminating the sense of discontinuity and incoordination caused by the switching of areas, and improving the overall visual coherence and realism of the image.
[0146] The background area can be adjusted to obtain the adjustment area; the object's visual information and the adjustment area are then edge-blended to obtain the target image. First, the background area is redrawn pixel-level according to the background adjustment strategy, generating an adjustment area that conforms to the semantic features of the reference background. Then, at the boundary between the adjustment area and the object's visual information, based on the color gradient, brightness distribution, and edge sharpness features of the pixel neighborhood, a local gradient transition is performed. This gradually matches the texture and lighting of the adjustment area with the visual attributes of the subject's edges in the object's visual information, ensuring no abrupt color differences or contour breaks at the junction, ultimately outputting a visually coherent target image. This process, through natural pixel-level fusion, avoids stitching marks that appear after background replacement, improving the visual consistency between the subject and the new background.
[0147] In the above embodiments of this application, determining the reference image corresponding to the image to be processed based on the category of the target object includes: determining multiple initial reference objects based on the category of the target object; filtering duplicate reference objects among the multiple initial reference objects based on the object data of the multiple initial reference objects to obtain reference objects, and obtaining the reference image corresponding to the reference objects, wherein the similarity between the duplicate reference objects and other reference objects is greater than a preset threshold, and the other reference objects are reference objects other than the duplicate reference objects among the multiple initial reference objects.
[0148] The category of the target object mentioned above refers to the classification identifier of the object contained in the image to be processed. It is defined by attributes such as the physical form, function, or usage scenario of the product and is used to associate a set of images with similar visual expression patterns.
[0149] The initial reference objects mentioned above are multiple candidate objects that belong to the same category as the target object and are retrieved from the image database based on the category of the target object. Their object data includes image content, annotation information and performance indicators, which serve as the original sample source for background style analysis.
[0150] The aforementioned repeated reference object refers to an object whose object data is highly consistent with at least one other reference object in terms of visual structure, composition, or color distribution among multiple initial reference objects. The similarity between the two objects exceeds a preset value, indicating that there is redundancy in their expressed content.
[0151] The aforementioned reference objects are initial, non-redundant reference objects that have been selected and retained. The object data of the reference objects are representative and diverse, and are used to support the subsequent inductive analysis of background semantic features.
[0152] The reference image mentioned above is the original image file corresponding to the reference object. It contains complete visual content and is used to extract background semantic features as the basis for background adjustment strategies.
[0153] In one optional embodiment, firstly, based on the category of the target object, a batch of initial reference objects of the same type are retrieved from the image resource library to form an initial sample pool; then, visual features can be extracted from the object data of each initial reference object, and the structural similarity between objects can be calculated. When the similarity between an object and other objects exceeds a preset threshold, the initial reference object is determined to be a duplicate reference object and is removed; the remaining objects that are not removed are reference objects, and their corresponding reference images are further obtained for subsequent background analysis.
[0154] This application embodiment refines the reference object set through category-driven initial screening and redundancy elimination mechanism of similarity clustering, avoiding background feature bias or statistical deviation caused by duplicate samples. This process does not rely on manual annotation, but is based on automatic comparison of object data, ensuring that the selected reference objects have broad distribution and representativeness in visual expression, so that subsequent background semantic analysis is based on effective and non-redundant data.
[0155] For example, if the target object is a rice cooker, 500 initial reference object images can be retrieved based on its category. Among them, 120 images are duplicate images of the same brand, angle, and background composition. The pixel distribution and contour structure similarity of the object data can be calculated. If the similarity is found to be higher than a preset threshold, they are identified as duplicate reference objects and removed. The remaining 380 images cover real display images of different brands, different countertop materials, and different lighting conditions. These can be used as reference objects, and reference images can be extracted from them to generate background adjustment strategies.
[0156] Initial reference objects are determined based on the category of the target object, and non-repeating reference objects are selected by using a similarity threshold between object data. This improves the diversity and representativeness of the reference image set, avoids bias in background analysis due to sample redundancy, and enhances the accuracy and generalization ability of subsequent background semantic induction.
[0157] This application embodiment retrieves multiple initial reference objects from an image resource library based on the category of the target object, forming an initial sample set. Subsequently, visual features are extracted from the object data of each initial reference object, and the structural similarity between the initial reference object and the other initial reference objects is calculated. When the similarity of an initial reference object to other initial reference objects exceeds a preset threshold, it is identified as a duplicate reference object and removed from the set. The remaining initial reference objects that are not removed are the reference objects. Finally, the reference image corresponding to each reference object can be obtained as input data for subsequent background analysis. This process achieves initial sample coverage through category association, and then eliminates redundant samples with highly consistent visual expressions through automatic comparison of object data, ensuring that the reference objects have sufficient differences in composition, color, and scene layout, so that the reference images can reflect diverse visual expression patterns under the same category, avoiding the loss of representativeness in background semantic analysis due to sample convergence.
[0158] In the above embodiments of this application, obtaining a reference image corresponding to a reference object includes: obtaining multiple initial reference images corresponding to the reference object; and filtering the multiple initial reference images based on their image quality to obtain a reference image.
[0159] The aforementioned reference objects refer to representative samples of the same type of target objects that have been retained after category matching and redundancy removal. Their object data contains structured visual features and are used to associate with the original image set that can serve as a basis for background analysis.
[0160] The aforementioned initial reference image is the original image file directly corresponding to the reference object, obtained by searching in the image resource library. It may contain multiple versions formed by various shooting angles, lighting conditions, background environments, or differences in shooting equipment, and its content has not yet undergone quality screening.
[0161] The aforementioned reference images are a collection of images that conform to visual expression standards and are retained after image quality assessment and screening of the initial reference images. These images are used to support subsequent background semantic analysis, and their quality characteristics include sharpness, integrity of the subject composition, absence of obvious occlusion, and absence of excessive compression distortion. Further screening of the reference images can be performed based on selection criteria, such as image quality and the number of images.
[0162] In one optional embodiment, firstly, based on the unique identifier of each reference object, multiple initial reference images associated with it are retrieved from the image resource library to form an initial image set; then, each initial reference image can be evaluated for image quality, and the evaluation dimensions include whether the image resolution meets the minimum threshold, whether the subject is complete and uncropped, whether the background has large areas of blur or noise, and whether there are visual interferences such as watermarks or text coverings; the initial reference images that meet the quality standards are retained, and the remaining images are discarded, and the final image set is the reference image.
[0163] After the reference object is determined, the original image is further filtered through objective image quality assessment to avoid low-quality images introducing noise that interferes with background semantic analysis. This process does not rely on manual judgment, but is automatically judged based on pixel-level features and structural integrity indicators to ensure that the reference images involved in the analysis have consistency and reliability in visual expression, thereby improving the stability and generalization ability of subsequent background feature extraction.
[0164] For example, an electric kettle is used as a reference object. The system acquires thirty initial reference images of it. Among them, five images have missing lids due to low shooting angles, seven images have overexposed metal surfaces due to strong background reflections, three images contain platform watermarks, and two images are low-resolution thumbnails. The system filters them one by one according to preset quality standards and retains thirteen images with complete subjects, clear backgrounds, no obstructions, and qualified resolution as reference images for subsequent extraction of background semantic features of "light-colored countertop + wood texture + food matching".
[0165] Image quality is screened based on multiple initial reference images to ensure that the final reference images meet the analysis requirements in terms of sharpness, integrity and interference control, thereby reducing the risk of misjudgment of background features due to image defects and improving the accuracy and consistency of background semantic modeling.
[0166] This application embodiment obtains multiple initial reference images corresponding to the reference object, forming an initial image set. Subsequently, each initial reference image can be evaluated for image quality to determine whether it meets quality conditions such as resolution not lower than a preset value, complete and uncropped main outline, no obvious blur or noise in the background area, and no watermark or text coverage. Initial reference images that meet the standards are retained, while the remaining images are discarded, ultimately obtaining the reference image. This process achieves automatic filtering of the original image through objective image feature analysis, avoiding low-resolution, incomplete composition, or visually disturbed images from participating in subsequent background semantic analysis, ensuring that the reference image has basic usability and consistency in visual expression, and providing stable input for subsequent background feature extraction.
[0167] The system provided in this application mainly consists of the following four modules, which work together to achieve a closed loop from image understanding to automated improvement: the modules include multi-dimensional intelligent analysis of product main images; construction of a competitor knowledge base and comparison of competitor backgrounds; background problem analysis and improvement decisions; and background improvement processing based on image editing.
[0168] Figure 3 This is a flowchart of an e-commerce product background improvement method according to an embodiment of this application. The method includes the following steps:
[0169] Step S301: Multi-dimensional intelligent analysis of the product main image.
[0170] The system performs multi-dimensional intelligent analysis on the main images of e-commerce products to be improved. The multi-dimensional intelligent analysis includes three parallel processing sub-steps: subject detection, selling point detection, and background understanding.
[0171] Step S302: Construction of competitor knowledge base and comparison of competitor background.
[0172] Construct a competitor knowledge base for products in the same category as the product to be improved, and analyze and learn the background design features of the main images of competitor products based on the competitor knowledge base, extracting the common features and patterns of excellent background designs in the category.
[0173] Step S303: Background problem analysis and improvement decision.
[0174] The multi-dimensional analysis results of the product main image to be improved can be comprehensively compared with the background comparison results of competitors to determine whether the background of the product main image needs improvement. The background improvement decision-making process can involve inputting the semantic description information of the background of the product main image to be improved and the background design analysis report of competitors into a decision model. The decision model can be a rule-based judgment system, a machine learning classification model, or a large language model inference system. The decision model evaluates the gap between the background of the product main image to be improved and the excellent background of competitors from multiple dimensions. The evaluation dimensions can be scene matching degree, color tone coordination degree, visual appeal, product prominence, and overall coordination.
[0175] Among them, scene matching degree can be judged as whether the background scene of the main image of the product to be improved is consistent with or close to the mainstream scene of the category; color coordination degree can be judged as whether the background color of the main image of the product to be improved is consistent with the mainstream color preference of the category; visual appeal can be judged as whether the background of the main image of the product to be improved has sufficient visual appeal and scene appeal; product prominence can be judged as whether the background of the main image of the product to be improved can effectively set off and highlight the main product; and overall coordination can be judged as the degree of visual coordination between the main product, selling point text, and background.
[0176] Based on the above multi-dimensional evaluation results, the decision model outputs a binary judgment result. If the result is "yes," the background needs to be improved; if the result is "no," the background does not need to be improved. When the judgment result is "no," it means that the background of the current product main image already has a good visual effect, and no background improvement processing is needed. The original product main image is directly output as the final result. When the judgment result is "yes," it means that the background of the current product main image has room for improvement, and background image editing and improvement are performed.
[0177] Step S304: Background improvement processing based on image editing.
[0178] When it is determined that the background of a product's main image needs to be improved, based on the results of the comparison of competitor backgrounds and the multi-dimensional analysis information of the product's main image, the background of the product's main image is improved through image editing technology, generating a new product main image with an improved background.
[0179] When generating background improvement strategies, specific strategies can be generated based on competitor background comparison reports and analysis of the current product main image. Background improvement strategies can include a target background scene description, a target color scheme description, a target spatial layout description, and constraints on retaining the product's main subject. The target background scene description can determine the scene type the improved background should present based on the common characteristics of competitor backgrounds, such as "a light-colored kitchen countertop paired with a food display scene." The target color scheme description can determine the color scheme of the improved background, such as "soft warm colors, mainly light colors." The target spatial layout description can determine the spatial layout of the improved background, such as "the main product is centered slightly to the left, with cooked food displayed on the right." Constraints on retaining the main product specify the conditions under which the main product and its selling point text information must be fully retained during the background improvement process.
[0180] During image editing, background enhancement strategies can be implemented using image editing models to improve the background area of the product's main image. The main body mask can be used as an editing constraint input to ensure the product's main body remains unchanged during image editing. Background enhancement strategies can be transformed into image editing instructions or text prompts as guiding inputs to the image editing model. Diffusion-based image editing techniques or other image generation and editing techniques can be employed, using the original product main image as a basis, to regenerate or replace only the background area, generating a new background that conforms to the target background description. During image editing, constraints based on the main body mask ensure that key elements such as the product's main body, brand logo, and selling points are not modified or obscured; only the background area is edited and replaced. Edge blending processing between the main body and background is applied to the generated image, including tone transitions and consistent lighting.
[0181] Figure 4 This is a flowchart of a multi-dimensional intelligent analysis method according to an embodiment of this application. The method includes the following steps:
[0182] Step S401, subject detection.
[0183] The system performs subject detection on the input product image, identifying and extracting the main product region. Specifically, a deep learning object detection model or instance segmentation model is used to analyze and process the product image, detecting the position, outline, and boundary information of the main product in the image, and generating a subject region mask. The subject region mask is used to accurately separate the main product from the background, providing a foundation for subsequent background replacement and image editing.
[0184] When performing subject detection, a pre-trained semantic segmentation network or instance segmentation network can be used to perform foreground-background segmentation on the main image of the product to obtain a pixel-level subject region mask. The segmented subject mask is then post-processed, including morphological improvement, edge smoothing, and hole filling, to obtain a more accurate and smooth subject outline. Structured features such as the bounding box information, area ratio information, and position distribution information of the subject region are extracted for subsequent improvement decision-making.
[0185] The semantic segmentation network or instance segmentation network mentioned above can be a U-shaped network (U-Net), a mask region-based convolutional neural network (Mask R-CNN), or a segmentation anything model.
[0186] Step S402, Selling point detection.
[0187] The system performs feature detection on the input product main image, identifying key elements such as textual feature information, icons, and promotional labels. OCR and multimodal large modeling techniques can be used to detect and recognize text content, feature labels, brand logos, price information, and promotional icons in the product main image.
[0188] Step S403, background understanding.
[0189] The system performs background understanding analysis on the input product main image to extract the visual features and semantic information of the current background. Specifically, it uses a multimodal large language model (MLLM) or a visual understanding model to perform deep semantic analysis on the background area of the product main image to understand information such as the scene type, color style, spatial layout, and atmosphere of the current background.
[0190] In background understanding, visual understanding models can be used to perform semantic analysis on background areas, identifying scene elements (such as countertops, desktops, kitchens, living rooms, etc.), color tones (such as light, dark, warm, and cool tones), spatial composition features (such as simple and complex), and overall atmosphere and style in the background; generating structured semantic descriptions of the current background, including but not limited to descriptions of background scene categories, color characteristics, composition methods, and visual styles.
[0191] Figure 5 This is a schematic diagram illustrating the construction of a competitor knowledge base according to an embodiment of this application, such as... Figure 5 As shown, the method includes the following steps:
[0192] Step S501: Constructing a competitor knowledge base.
[0193] Collect main image data of competing products in the same category as the product to be improved, and build a competitor knowledge base. Based on the category information of the product to be improved (such as air fryers, rice cookers, etc.), retrieve competing products in the same category from the product database of the e-commerce platform, and prioritize high-quality competing products with high sales ranking, high click-through rate, and good conversion rate. Collect main image data of competing products, including multiple product display images from different angles and scenarios. Perform quality screening and classification on the collected competitor main images, remove low-quality and highly similar duplicate images, and build a structured competitor knowledge base. The data stored in the competitor knowledge base includes, but is not limited to, the original images of competitor main images, category tags of competitor products, sales indicators, user reviews, and other related information.
[0194] Step S502, competitor background feature analysis.
[0195] This analysis performs background feature analysis on the main images of competing products in the competitor knowledge base, extracting common features of excellent background designs within the category. It involves performing subject detection and background understanding processing on each main image of each competitor in the knowledge base, extracting the semantic features of each image's background. A multimodal large language model is used to summarize the common features of competitor backgrounds, generating an analysis report on competitor background design. This report includes summaries of preferences for mainstream background styles, color schemes, spatial layouts, and atmosphere. The report is presented in a structured text description format, such as "Competitors use light-colored kitchen countertops as a soft background to create a relaxed, everyday atmosphere, highlighting the product's usage and the food's appearance, resulting in a more lifelike and appetite-stimulating effect."
[0196] Figure 6 This is a schematic diagram of an image processing structure according to an embodiment of this application, such as... Figure 6 As shown, the system can perform subject detection and selling point detection on the product image to be processed, obtain the product mask image and product text description, perform background understanding on the product image to obtain the background information of the product image, construct a competitor knowledge base corresponding to competitor product images of the same category as the product image, perform background understanding on the competitor product images in the competitor knowledge base to obtain the background information of the competitor product images, and determine whether the background information of the product image needs to be adjusted based on the background information of the competitor product images and the background information of the product image. If the judgment result is no, no adjustment is needed; if the judgment result is yes, adjustment is needed and the image needs to be edited to obtain the target product image.
[0197] This application introduces a competitor learning mechanism, utilizing massive amounts of high-click-rate product image data from e-commerce platforms to dynamically capture and respond to consumers' real preferences for the main image background, thereby upgrading the background design paradigm from experience-driven to data-driven. During the improvement process, the main product and key selling points are kept unchanged to ensure that the core information carried by the image is not lost. Without increasing the visual information load, the semantic reconstruction and style transfer of the background enhance the product's sense of scene immersion and visual appeal.
[0198] This application can improve the background of the main image automatically, reduce the cost of manual design, and increase the click-through rate and conversion efficiency of tens of thousands of products on a large scale. By constructing a background improvement decision mechanism driven by competitor learning, it can intelligently extract better background styles based on high-click competitor images, and combine it with model editing technology that is consistent with the main body. Under the premise of preserving the main body of the product and key selling points, it can achieve semantic-level accurate replacement and natural integration of the background, thereby achieving a dual improvement in visual appeal and information accuracy.
[0199] According to an embodiment of this application, an image processing method is provided. Figure 7 This is a flowchart of an image processing method according to an embodiment of this application, such as... Figure 7 As shown, the method includes:
[0200] Step S702: In response to the input command applied to the operation interface, display the image of the target object to be processed on the operation interface.
[0201] When a user performs an input action on the interface, such as clicking to select an image or specifying an object, the system will recognize the action and display the image that needs to be processed on the interface as the input basis for subsequent processing.
[0202] Step S704: In response to the processing instructions applied to the operation interface, display the target image of the target object on the operation interface.
[0203] The target image is obtained based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the object's visual information. The reference image is determined based on the category of the target object. The object's visual information is obtained by performing multi-dimensional detection on the image to be processed. Multi-dimensional detection is used to represent the process of detecting visual elements associated with the target object in the image to be processed in different dimensions. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold.
[0204] When a user issues a processing command, such as wanting to change the background, enhance the image, or generate a more natural composite effect, the system does not simply manipulate the original image. Instead, it intelligently integrates multiple key pieces of information. First, it extracts background elements from reference images of the same category as the target object. Then, it extracts the original background surrounding the target object in the current image to be processed. Finally, it extracts the visual features of the target object itself through multi-dimensional visual analysis (such as color distribution, texture structure, light and shadow relationships, edge details, posture and shape). This information is then comprehensively calculated and integrated to generate a visually more realistic, harmonious, and natural target image. Finally, the composite result is presented to the user, achieving high-quality and intelligent image processing effects.
[0205] Through the above steps, in response to input commands applied to the operation interface, the image to be processed of the target object is displayed on the operation interface; in response to processing commands applied to the operation interface, the target image of the target object is displayed on the operation interface. The target image is obtained based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the object's visual information. The reference image is determined based on the category of the target object. The object's visual information is obtained through multi-dimensional detection of the image to be processed. Multi-dimensional detection represents the process of detecting visual elements associated with the target object in the image to be processed in different dimensions. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold. By obtaining the object's visual information through multi-dimensional detection, combined with the background features of the reference object in the reference image, the semantic differences between the background of the image to be processed and the reference image are compared. The deficiencies of the background of the image to be processed in multiple dimensions, such as scene adaptability, visual coordination, and style expressiveness, are inferred. Based on these differences and the object's visual information, a targeted improvement strategy is generated to adjust the background region, ensuring that the visual information of the target object remains unchanged, thereby improving the overall visual expressiveness of the image and solving the technical problem of poor image processing effects in related technologies.
[0206] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0208] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided. Figure 8 This is a schematic diagram of an image processing apparatus according to an embodiment of this application, such as... Figure 8 As shown, the device 800 includes: a detection module 802, a determination module 804, and an adjustment module 806.
[0209] The detection module is used to perform multi-dimensional detection on the image to be processed in response to receiving the image of the target object, thereby obtaining the visual information of the target object. The multi-dimensional detection refers to the process of detecting visual elements associated with the target object in the image to be processed in different dimensions. The determination module is used to determine the reference image corresponding to the image to be processed based on the category of the target object. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold. The adjustment module is used to adjust the image to be processed based on the background area of the reference object in the reference image, the background area of the target object in the image to be processed, and the visual information of the target object, thereby obtaining the target image of the target object.
[0210] It should be noted that the detection module 802, determination module 804, and adjustment module 806 correspond to steps S202 to S206 in the above embodiments. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also run as part of the device in the server 10 provided in the above embodiments.
[0211] In the above embodiments of this application, the detection module is used to detect the image to be processed and obtain the object region mask and visual description information of the target object. The object region mask is used to represent the mask of the region where the target object is located in the image to be processed, and the visual description information is used to represent the description information of the target object. Based on the object region mask and the visual description information, the visual information of the object is determined.
[0212] In the above embodiments of this application, the detection module is used to perform foreground segmentation on the image to be processed to obtain an object region mask, wherein the foreground segmentation is used to represent the extraction of pixels other than the background region in the image to be processed; and to perform content recognition on the image to be processed to obtain visual description information.
[0213] In the above embodiments of this application, the adjustment module is used to compare the background area of the reference object in the reference image with the background area of the target object in the image to be processed to obtain a background comparison result; and to adjust the image to be processed based on the background comparison result and the object visual information to obtain the target image.
[0214] In the above embodiments of this application, the adjustment module is used to perform background recognition on the background region in the image to be processed to obtain the first background information of the image to be processed, wherein the first background information is used to represent the visual state of the background region in the image to be processed; to perform background recognition on the background region in the reference image to obtain the second background information of the reference image, wherein the second background information is used to represent the visual state of the background region in the reference image; and to perform background comparison on the image to be processed based on the first background information and the second background information to obtain the background comparison result.
[0215] In the above embodiments of this application, the adjustment module is used to perform multi-dimensional evaluation of the image to be processed and the reference image based on the first background information and the second background information, and obtain multiple evaluation results. The multi-dimensional evaluation includes at least one of the following: scene matching degree, tone coordination degree, visual attention degree, and element coordination degree of the background region in the image to be processed and the reference image; and determines the background comparison result based on the multiple evaluation results.
[0216] In the above embodiments of this application, the adjustment module is used to generate a background adjustment strategy based on the second background information and object visual information when the background comparison result indicates that the background of the image to be processed needs to be adjusted, and to adjust the image to be processed based on the background adjustment strategy to obtain the target image; when the background comparison result indicates that the background of the image to be processed does not need to be adjusted, the image to be processed is determined to be the target image.
[0217] In the above embodiments of this application, the adjustment module is used to adjust the background area in the image to be processed using the object's visual information as the adjustment constraint and the background adjustment strategy as the guiding condition, so as to obtain the target image. The adjustment constraint is used to keep the object's visual information in the image to be processed unchanged during the adjustment process, and the guiding condition is used to determine the adjustment direction of the background area during the adjustment process.
[0218] In the above embodiments of this application, the adjustment module is used to adjust the background area to obtain the adjustment area; and to perform edge fusion of the object visual information and the adjustment area to obtain the target image.
[0219] In the above embodiments of this application, the determining module is used to determine multiple initial reference objects based on the category of the target object; based on the object data of the multiple initial reference objects, the duplicate reference objects among the multiple initial reference objects are filtered to obtain reference objects, and the reference images corresponding to the reference objects are obtained, wherein the similarity between the duplicate reference objects and other reference objects is greater than a preset threshold, and the other reference objects are reference objects other than the duplicate reference objects among the multiple initial reference objects.
[0220] In the above embodiments of this application, the determining module is used to obtain multiple initial reference images corresponding to the reference object; based on the image quality of the multiple initial reference images, the multiple initial reference images are filtered to obtain a reference image.
[0221] It should be noted that the preferred embodiments involved in the above embodiments of this application are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.
[0222] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided. Figure 9 This is a schematic diagram of an image processing apparatus according to an embodiment of this application, such as... Figure 9 As shown, the device 900 includes: a first display module 902 and a second display module 904.
[0223] The first display module is used to respond to input commands applied to the operation interface and display the image of the target object to be processed on the operation interface; the second display module is used to respond to processing commands applied to the operation interface and display the target image of the target object on the operation interface. The target image is obtained based on the background area of the reference object in the reference image, the background area of the target object in the image to be processed, and the object's visual information. The reference image is determined based on the category of the target object. The object's visual information is obtained by performing multi-dimensional detection on the image to be processed. Multi-dimensional detection is used to represent the process of detecting visual elements associated with the target object in the image to be processed in different dimensions. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold.
[0224] It should be noted that the first display module 902 and the second display module 904 mentioned above correspond to steps S702 to S704 in the above embodiments. The two modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware components or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the server 10 provided in the above embodiments.
[0225] Embodiments of this application may provide a computing device. Figure 10 This is a structural block diagram of a computing device according to an embodiment of this application. Figure 10 As shown, the computing device 100 may include one or more (only one is shown in the figure) processors 102, memory 104, memory controller, and peripheral interfaces.
[0226] The aforementioned computing device can be understood as an integrated smart terminal, including but not limited to servers, desktop computers, PCs (Personal Computers), all-in-one model machines, etc., and the computing device may have the model in the above embodiments of this application pre-installed.
[0227] Specifically, this computing device can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, this computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, this computing device also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other model types), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, this computing device can also create applications based on models, providing API calling capabilities, allowing models to be called into created applications through API interfaces, and providing application management tools to achieve application control.
[0228] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn and master AI technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for AI development, training, deployment, and application.
[0229] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0230] The processor can invoke an executable program stored in memory via a transmission device to execute any of the methods described in the above embodiments.
[0231] Embodiments of this application may provide an electronic device. Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 11 As shown, the electronic device may include: an input / output device 112; a memory 114; and a processor 116, wherein the processor 116 is connected to the input / output device 112 and the memory 114 via a bus 118.
[0232] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0233] The processor can invoke an executable program stored in memory via a transmission device to execute the method described in any of the above embodiments.
[0234] Those skilled in the art will understand that, Figure 11 The structure shown is for illustrative purposes only. The electronic device may also be a smartphone, tablet, PDA, mobile internet device (MID), PAD, or other terminal device. This diagram does not limit the structure of the aforementioned electronic device. For example, the electronic device may include more or fewer components (such as network interfaces, display devices, etc.) than shown in the diagram, or may have a different configuration than shown in the diagram.
[0235] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0236] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.
[0237] Optionally, in this embodiment, the storage medium may be located in a computing device or an electronic device.
[0238] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, which, when the executable program is running, controls the device where the computer-readable storage medium is located to execute the method described in any of the above embodiments.
[0239] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0240] Embodiments of this application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.
[0241] Embodiments of this application also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.
[0242] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0243] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0244] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0245] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0246] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0247] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An image processing method, characterized in that, include: In response to receiving a target object image to be processed, multi-dimensional detection is performed on the target object image to obtain object visual information, wherein the multi-dimensional detection is used to represent the process of detecting visual elements associated with the target object in the target object image in different dimensions; Based on the category of the target object, a reference image corresponding to the image to be processed is determined, wherein the reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold; Based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the visual information of the object, the image to be processed is adjusted to obtain the target image of the target object.
2. The method according to claim 1, characterized in that, Multi-dimensional detection is performed on the image to be processed to obtain the object visual information of the target object, including: The image to be processed is detected to obtain the object region mask and visual description information of the target object, wherein the object region mask is used to represent the mask of the region where the target object is located in the image to be processed, and the visual description information is used to represent the description information of the target object; Based on the object region mask and the visual description information, the visual information of the object is determined.
3. The method according to claim 2, characterized in that, The image to be processed is detected to obtain the object region mask and visual description information of the target object, including: Foreground segmentation is performed on the image to be processed to obtain the object region mask, wherein the foreground segmentation is used to represent the extraction of pixels other than the background region in the image to be processed; Content recognition is performed on the image to be processed to obtain the visual description information.
4. The method according to claim 1, characterized in that, Based on the background region of the reference object in the reference image, the background region of the target object in the image to be processed, and the object's visual information, the image to be processed is adjusted to obtain a target image of the target object, including: A background comparison is performed between the background region of the reference object in the reference image and the background region of the target object in the image to be processed to obtain a background comparison result; The image to be processed is adjusted based on the background contrast results and the object visual information to obtain the target image.
5. The method according to claim 4, characterized in that, Based on the reference image, a background comparison is performed on the background region of the target object in the image to be processed to obtain a background comparison result, including: Background recognition is performed on the background region in the image to be processed to obtain the first background information of the image to be processed, wherein the first background information is used to represent the visual state of the background region in the image to be processed; Background recognition is performed on the background region in the reference image to obtain the second background information of the reference image, wherein the second background information is used to represent the visual state of the background region in the reference image; Based on the first background information and the second background information, a background comparison is performed on the image to be processed to obtain the background comparison result.
6. The method according to claim 5, characterized in that, Based on the first background information and the second background information, a background comparison is performed on the image to be processed to obtain the background comparison result, including: Based on the first background information and the second background information, the image to be processed and the reference image are evaluated in multiple dimensions to obtain multiple evaluation results. The multi-dimensional evaluation includes at least one of the following: scene matching degree, tone coordination degree, visual attention degree, and element coordination degree of the background region in the image to be processed and the reference image. Based on the multiple evaluation results, the background comparison results are determined.
7. The method according to claim 4, characterized in that, The background contrast result is used to indicate whether background adjustment is needed for the image to be processed. Based on the background contrast result and the object visual information, the image to be processed is adjusted to obtain the target image of the target object, including: If the background comparison result indicates that the background of the image to be processed needs to be adjusted, a background adjustment strategy is generated based on the second background information and the object visual information, and the image to be processed is adjusted based on the background adjustment strategy to obtain the target image; If the background comparison result indicates that no background adjustment is needed for the image to be processed, then the image to be processed is determined to be the target image.
8. The method according to claim 7, characterized in that, The target image is obtained by adjusting the image to be processed based on the background adjustment strategy, including: Using the object's visual information as an adjustment constraint and the background adjustment strategy as a guiding condition, the background region in the image to be processed is adjusted to obtain the target image. The adjustment constraint is used to keep the object's visual information in the image to be processed unchanged during the adjustment process, and the guiding condition is used to determine the adjustment direction of the background region during the adjustment process.
9. The method according to claim 8, characterized in that, Adjusting the background region in the image to be processed to obtain the target image includes: The background area is adjusted to obtain the adjusted area; The target image is obtained by edge fusion of the visual information of the object and the adjustment area.
10. The method according to any one of claims 1-6, characterized in that, Based on the category of the target object, a reference image corresponding to the image to be processed is determined, including: Based on the category of the target object, multiple initial reference objects are determined; Based on the object data of the plurality of initial reference objects, duplicate reference objects among the plurality of initial reference objects are filtered to obtain the reference object, and the reference image corresponding to the reference object is obtained. The similarity between the duplicate reference object and other reference objects is greater than a preset threshold, and the other reference objects are reference objects other than the duplicate reference objects among the plurality of initial reference objects.
11. The method according to claim 10, characterized in that, Obtaining the reference image corresponding to the reference object includes: Obtain multiple initial reference images corresponding to the reference object; Based on the image quality of the plurality of initial reference images, the plurality of initial reference images are filtered to obtain the reference image.
12. An image processing method, characterized in that, include: In response to input commands applied to the operation interface, the image of the target object to be processed is displayed on the operation interface; In response to a processing instruction applied to the operation interface, a target image of the target object is displayed on the operation interface. The target image is obtained based on the background region of the reference object in a reference image, the background region of the target object in the image to be processed, and object visual information. The reference image is determined based on the category of the target object. The object visual information is obtained by performing multi-dimensional detection on the image to be processed. The multi-dimensional detection is used to represent the process of detecting visual elements associated with the target object in the image to be processed in different dimensions. The reference image contains a reference object, and the similarity between the category of the reference object and the category of the target object is greater than a preset threshold.
13. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor, connected to a memory via a bus, is used to run the program, wherein the program, when running, executes the method described in any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 12.
15. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 12.