Image processing method, computing device, electronic device and storage medium
By using user-driven region selection and descriptive text input for model integration into scenes, the method addresses the high cost and complexity of current model image generation, achieving realistic and coherent results.
Patent Information
- Application Number
- CN202510406488.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, when generating an image that integrates the target object into the scene, the cost is high, the user operation process is complicated, and the target object and the scene are not fusion effect.
By responding to the user's area selection instruction on the initial scene image, the mask matrix and the description text of the target object are obtained and inputted into the image generation model, and then the initial object scene image is generated and fused to ensure that the target object is naturally presented in the scene.
It reduces the cost of image generation and user operation complexity, improves the realism, immersion and consistency of background texture of the image, and achieves efficient integration of target objects and scenes.
Smart Images

Figure CN120318355A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of image generation and image processing. Specifically, it relates to an image processing method, a computing device, an electronic device, and a storage medium. Background Art
[0002] In the home goods industry, online display is a key link in attracting consumers and improving the shopping experience. E-commerce platforms increasingly rely on high-quality visual content to display products.
[0003] However, the current mainstream model image generation technology is usually based on on-site shooting and image editing methods. The on-site shooting method also involves complex processes such as post-production, resulting in a high cost of generating model images, a complex user operation process, and low flexibility and efficiency of image generation; and for the model images generated by related technologies, it is difficult for the model to naturally integrate into complex backgrounds. When dealing with natural poses of models in a home environment, such as sitting, lying, leaning, etc., it is difficult to perform precise generation control, resulting in single and unnatural actions, which affects the realism and attractiveness of the images. In summary, when generating images that integrate the target object into the scene in related technologies, the cost is high, the user operation process is complex, and the fusion effect of the target object and the scene is poor.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide an image processing method, a computing device, an electronic device, and a storage medium to at least solve the technical problems of high cost, complex user operation process, and poor fusion effect of the target object and the scene when generating images that integrate the target object into the scene in related technologies.
[0006] According to one aspect of the embodiments of this application, an image processing method is provided. The method includes: responding to a region selection instruction for a target region on an initial scene image, obtaining a mask matrix of the target region and a target description text of a target object, where the target description text is used to describe the target object; inputting the mask matrix, the target description text, and the initial scene image into an image generation model, and using the image generation model to generate an initial object scene image, where the initial object scene image contains the target object; fusing the initial scene image and the initial object scene image to obtain a target object scene image, where the target region in the target object scene image contains the target object.
[0007] According to another aspect of the embodiments of this application, a computing device is further provided, including: a memory storing an executable program; a processor for running the program, where when the program runs, it executes the methods in the various embodiments of this application.
[0008] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; a processor connected to the memory through a bus for running the program, wherein when the program runs, it executes the methods in various embodiments of the present application.
[0009] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in various embodiments of the present application.
[0010] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program that implements the methods in various embodiments of the present application when executed by a processor.
[0011] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a non-volatile computer-readable storage medium storing a computer program that implements the methods in various embodiments of the present application when executed by a processor.
[0012] According to another aspect of the embodiments of the present application, a computer program is further provided that implements the methods in various embodiments of the present application when executed by a processor.
[0013] In the embodiments of the present application, first, in response to a region selection instruction for a target region in an initial scene image, a mask matrix of the target region and a target description text of a target object are obtained, and the target description text is used to describe the characteristics of the target object; then, the mask matrix, the target description text, and the initial scene image are input into an image generation model together, and an initial object scene image is generated through the image generation model, and the described target object is included in the initial object scene image; finally, the initial scene image and the initial object scene image are fused to generate a final target object scene image, and the target object is accurately presented within the target region in the target object scene image. It is easy to notice that by responding to the user's region selection instruction for the target region on the initial scene image, obtaining the mask matrix of the region and the target description text of the target object, and then inputting the mask matrix, the target description text, and the initial scene image into the image generation model together, the user can achieve this by simply and intuitively selecting a region and inputting a description through an interface, which reduces the cost of image generation and the complexity of user operations; by means of the region selection instruction, the user is allowed to directly specify the editing region, and the target description text can guide the image generation model to generate an object image that meets the user's expectations, enabling the user to accurately control the generation process and improving the flexibility of image generation. The fusion of the initial scene image and the initial object scene image ensures the docking between the initial scene and the generated object, improves the realism, immersion, and coherence of the background texture of the image, and thus solves the technical problems in the related art that when generating an image that integrates a target object into a scene, the cost is relatively high, the user operation process is relatively complex, and the fusion effect of the target object and the scene is not good.
[0014] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. Brief Description of the Drawings
[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0016] Figure 1 is a schematic diagram of an application scenario of an image processing method according to an embodiment of the present application;
[0017] Figure 2 is a flowchart of an image processing method according to an embodiment of the present application;
[0018] Figure 3 is a schematic diagram of an optional user operation interface according to an embodiment of the present application;
[0019] Figure 4It is a schematic diagram of an optional generation effect display interface according to an embodiment of the present application;
[0020] Figure 5 It is a schematic diagram of an optional generation of a target object scene image according to an embodiment of the present application;
[0021] Figure 6 It is a schematic diagram of an optional generation of a target object scene image including a model structure according to an embodiment of the present application;
[0022] Figure 7 It is a schematic diagram of an optional post - processing process image effect display according to an embodiment of the present application;
[0023] Figure 8 It is a structural block diagram of a computing device according to an embodiment of the present application;
[0024] Figure 9 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0027] First, some nouns or terms that appear during the description of the embodiments of the present application are applicable to the following explanations:
[0028] A transducer can utilize the self - attention mechanism to capture global semantic information and enhance the coherence of the generated image.
[0029] An autoencoder can be a generative model composed of an encoder and a decoder, which realizes efficient feature expression by compressing data into a low-dimensional space and reconstructing it.
[0030] According to an embodiment of the present application, an image processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that used here.
[0031] The above image processing method provided by the embodiment of the present application can be applied to Figure 1 the application scenarios shown as follows, but not limited thereto. In the application scenarios shown as follows, Figure 1 the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include but are not limited to: smart phones, tablet computers, laptop computers, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. The client device 20 can interact with the user through a graphical user interface, thereby implementing the method provided by the embodiment of the present application.
[0032] In the embodiment of the present application, the system composed of the client device and the server can execute the following steps: The client device can interact with the server. The server can respond to a region selection instruction for a target region on the initial scene image, obtain a mask matrix of the target region and a target description text of the target object; input the mask matrix, the target description text, and the initial scene image into an image generation model, and use the image generation model to generate an initial object scene image; fuse the initial scene image and the initial object scene image to obtain a target object scene image.
[0033] It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided by the embodiment of the present application can also be applied to a model all-in-one machine. In an optional embodiment, multiple models are built in the model all-in-one machine, and the user can select and adjust one model according to needs to obtain his own model. Thus, the high-performance computing unit built in the model all-in-one machine can directly call the adjusted model to execute the above method provided by the embodiment of the present application. In another optional embodiment, a trained model is built in the model all-in-one machine. Thus, the high-performance computing unit built in the model all-in-one machine can directly call this model to execute the above method provided by the embodiment of the present application.
[0034] Further, when the user needs to train their own model, they can also upload their own dataset through the client. This dataset is sent from the client to the server, enabling the server to adjust the pre-trained model with this dataset to obtain the user's own model, which is then deployed to the production environment. To facilitate the user's model adjustment requirements, the server can provide complete adjustment tools, development frameworks, and processes, supporting multiple adjustment strategies, so that the adjusted model can better adapt to different field applications and achieve high customization.
[0035] Under the above operating environment, the present application provides an Figure 2 image processing method as shown. Figure 2 It is a flowchart of an image processing method according to an embodiment of the present application. As Figure 2 shown, it may specifically include the following steps:
[0036] Step S202, in response to a region selection instruction for a target region on the initial scene image, obtain a mask matrix of the target region and a target description text of the target object.
[0037] Among them, the target description text is used to describe the target object.
[0038] The above-mentioned initial scene image may refer to the background image to which the target object is to be added. For example, in the scenario of generating a home image containing a model figure, the initial scene image may be a home or product background image such as a living room or a bedroom; again, for example, in the scenario of generating an image of a person integrated with a tourist attraction, the initial scene image may be the background image of the tourist attraction. The initial scene image can be obtained by the user taking a photo, or can be selected by the user from a pre-set image library, or can also be downloaded by the user from the Internet. The content and acquisition method of the initial scene image can be determined according to actual needs and are not limited here.
[0039] The above-mentioned target region may be the region in the initial scene image where the target object is to be added. The target region can be the region boxed by the user in the initial scene image, or can be the default region in the initial scene image, such as the middle local region, etc., or can also be an intelligently matched region. For example, in the scenario of generating a home image containing a model figure, regions such as sofas and seats in the initial scene image can be intelligently matched as the target region. The target region can be determined according to actual needs and is not limited here.
[0040] The above-mentioned region selection instruction can be an instruction triggered when the user determines the target region, can be an instruction for the user to select a target region by framing on the initial scene image, or can be an instruction triggered when the user selects a default target region or intelligently selects a target region, or can be a region selection instruction automatically generated by the user inputting a prompt word in a voice manner, which can be generated according to the voice effect. For example, the user can generate a target object on the sofa in the initial scene image by voice input. It can also be a region selection instruction generated by the user inputting a prompt word in a text manner and analyzing and processing the prompt word. The specific determination method of the region selection instruction can be determined according to actual needs and is not limited here.
[0041] The above-mentioned mask matrix can refer to a matrix used to mark or select the target region in the initial scene image. The elements in the mask matrix can be binary values, such as 0 and 1, or can be other numerical values for marking specific regions, which are not limited here; the mask matrix can be a matrix with the same size as the initial scene image and can identify or select the target region of the region of interest in the initial scene image.
[0042] The above-mentioned target object can be an object to be added to the initial scene image. For example, it can be a model figure to be added, the user's selfie, or an animal to be added, etc. The target object can be an object that has not been actually obtained yet, but the user can describe it through natural language. For example, it can be a model image that has not been actually obtained yet, but the model image described by the user through natural language. The specific target object can be determined according to actual needs and is not limited here.
[0043] The above-mentioned target description text can be the language text used by the user to describe the expected target object to be added. For example, it can be a brief text description provided by the user, such as the description of "casual wear model"; it can also be the description obtained after expanding the brief text description provided by the user and following the description specification such as "human body attributes - clothing details - pose characteristics - scene adaptation". The specific target description text can be determined according to actual needs and is not limited here.
[0044] In an alternative embodiment, the system can respond to a region selection instruction made by the user, such as a box selection operation performed by the user on the initial scene image. It can capture the region selection instruction and, using computer vision techniques, such as a boundary detection algorithm based on deep learning, generate a mask matrix that matches the size of the initial scene image based on the user's box selection coordinates. The mask matrix can mark the target region as 1 and the remaining background regions as 0. It can further adopt a contour refinement mechanism to improve the mask boundary through edge enhancement and refinement algorithms, avoiding the problem of inconsistent image generation caused by slight deviations in the user's box selection and ensuring the precise fit of the mask boundary to the target region. The system can also integrate object recognition and segmentation technologies, which can automatically identify the object categories within the box selection region and further improve the mask matrix to ensure precise control of the target object during the generation process without affecting the background details outside the region. In the scenario of generating a home image containing a model figure, the brief text description provided by the user, such as the description "casual wear model", lacks details. The system can also set up a deep learning text generation model that can be pre-trained with a relatively large amount of text and image data related to home scenes and can expand the short description into a detailed text. After receiving the brief text description input by the user, the system can use the text generation model to analyze and semantically expand the brief text description, generating a long text description that includes the clothing style, specific poses, and details of the interaction with the scene of the model, generating a description that follows the description specification such as "human body attributes - clothing details - pose features - scene adaptation". This process not only enriches the input information but also ensures a high degree of coordination between the generated content and the scene, providing precise guidance for subsequent image generation.
[0045] In the above process, by implementing the user's intuitive response to the region selection instruction and automated text expansion, the operation complexity of the user during the image generation process can be significantly reduced, enhancing the user experience and enabling non-professional users to easily complete the customized generation of high-quality images. The mask matrix generated based on the contour refinement and object recognition technologies, as well as the detailed text description expanded by the text generation model, provide precise control signals for the subsequent image generation model, enabling the image generation model to accurately generate an image that blends with the scene background, avoiding the destruction of background information and enhancing the overall visual effect and authenticity of the generated image.
[0046] Step S204, input the mask matrix, the target description text, and the initial scene image into the image generation model, and use the image generation model to generate an initial object scene image.
[0047] Among them, the initial object scene image contains the target object.
[0048] The above-mentioned image generation model can be an image processing model capable of receiving three types of heterogeneous data, namely a mask matrix, a target description text, and an initial scene image, and outputting an initial object scene image. The image generation model can be specifically determined according to actual needs and is not limited here.
[0049] In an alternative embodiment, after obtaining the mask matrix of the target area and the target description text for the target object in the above preprocessing stage, the mask matrix, the target description text, and the initial scene image can then be input into the image generation model to generate an initial object scene image containing the target object, achieving the multi-modal fusion of conditional input. After the mask matrix, the target description text, and the initial scene image are input into the image generation model, these heterogeneous information can be effectively fused. For example, the input initial scene image can be encoded to obtain a scene graph hidden vector, the semantic feature vector can be extracted from the target description text using a text editor, and the mask matrix can be cross-modally aligned with the target description text through channel concatenation. Then, the obtained scene graph hidden vector, semantic feature vector, and cross-modal alignment result can be injected into the latent space together to obtain the initial object scene image. Here, the method of fusing these three types of heterogeneous data, namely the mask matrix, the target description text, and the initial scene image, to generate the initial object scene image can also be determined according to actual needs and is not limited here. During the generation process, the image generation model can encode the input initial scene image to extract the underlying features of the initial scene image. Then, in the decoding stage, the image generation model can utilize the semantic features of the mask matrix and the target description text to gradually denoise and naturally fuse the target object into the scene, realizing processes including the prediction of the target pose, the adjustment of the lighting consistency between the object and the environment, and the reconstruction of detailed textures, ensuring that the generated initial object scene image not only retains the original beauty of the scene but also precisely conforms to the user's description of the target object. A multi-modal attention mechanism can also be introduced to enable the image generation model to dynamically adjust the attention to different modal inputs during the generation process, ensuring that the generated result accurately reflects the user's intention.
[0050] In the above process, through the image generation model, three types of heterogeneous data, namely the mask matrix, the target description text, and the initial scene image, are fused to generate the initial object scene image, which can capture global semantic information, generate natural and smooth object forms, and in the case of complex textures and other scenes in the home environment, the generated initial object scene image is not only realistic and natural, but also highly integrated with the scene background, avoiding the problems of hard edges and background damage. Through the target description text, the system can generate highly customized images to meet the personalized needs of users for the target object, including clothing, posture, and specific ways of interacting with the scene, achieving image generation with high fidelity, naturalness, and customization capabilities. In the scenario of generating home images containing model figures, it can ensure that the generated images can meet the user's needs both in terms of content and form.
[0051] Step S206: Fuse the initial scene image and the initial object scene image to obtain the target object scene image.
[0052] Among them, the target area in the target object scene image contains the target object.
[0053] In an alternative embodiment, after the image generation model completes the preliminary generation of the target object, the system can perform a post-processing process on the image to fuse the initial object scene image with the initial scene image, and finally generate the target object scene image, including fine-tuning environmental factors such as lighting and shadows to ensure a high degree of harmony between the generated object and the scene. When fusing the initial scene image and the initial object scene image, the weighted summation method can be used. By using one or more masks as weights, the initial scene image and the initial object scene image are mixed together in a preset ratio to achieve the purpose of smooth transition and precise control. Or the transparency blending method can be used for fusion. Both the initial scene image and the initial object scene image have their own transparency coefficients, ranging from 0 to 1. When fusing, the transparency coefficient channels of the initial scene image and the initial object scene image are used together with the red, green, and blue channels to control the visibility and blending effect of the image layers. Transparency blending can achieve more complex stacking effects, such as fading in and out or partial transparent coverage. The selective fusion method can also be used to fuse the initial scene image and the initial object scene image. The selective fusion method allows the user to specify specific image regions or features for fusion before fusion, rather than the entire image. For example, if only the facial features of the model are wanted to be integrated into the scene, this part of the details can be selectively fused, and the rest can be left unchanged.
[0054] Optionally, the post - processing process can accurately identify and extract the changed areas in the initial object - scene image through a multi - level mask calculation process, including the target object itself and the shadows generated by the target object, so as to generate an accurate shadow mask including shadows. At the same time, the target object can also be segmented by using a self - trained saliency segmentation model or other methods to generate a human - body segmentation mask, so as to ensure the accuracy of subsequent image fusion and the integrity of the human - body area. Then, logical operations can be used to merge the human - body segmentation mask and the shadow mask to generate a final fusion mask. The fusion mask can contain the accurate boundary of the target object and also consider the shadow area generated by the interaction between the target object and the environment, providing an accurate control signal for image fusion. Subsequently, based on the fusion mask, the initial object - scene image and the initial scene image can be fused. The system can also adopt a detail - redrawing technique to perform refined processing on key areas such as the face, hands, and feet, further enhancing the vividness and visual effect of the image.
[0055] In the above process, the fusion of the target object and the initial scene image is realized, while maximizing the retention of the original information of the scene background, avoiding the deformation problem of background commodities, and ensuring the visual coherence and integrity of the generated image. The post - processing module can adjust and improve the shadow effect of the target object by accurately calculating the shadow mask, making it consistent with the lighting conditions in the scene, effectively enhancing the realism and immersion of the image, and improving the display effect of the target - object scene image in scenarios such as home furnishings.
[0056] For example, in the scenario where the above - mentioned image - processing method is applied to generate a home - furnishing image containing a model figure, when the user selects a target area on the home - furnishing scene image, that is, the initial scene image, through the interaction interface, a region - selection instruction can be generated. The system can respond to this region - selection instruction, and then obtain the mask matrix of the selected target area, dividing the entire initial scene image into a target area and a non - target area. The pixel values of the target area can be marked as 1, and the non - target area can be marked as 0. It can also be marked according to actual needs, which is not limited here. This process can be automatically completed by means of threshold segmentation or image - segmentation networks in computer vision. At the same time, the user can input a short description text through the interaction interface to describe the model figure to be generated, such as a description text like "a casual - dressed model sitting on the sofa". The system can use a fine - tuned deep - learning model to perform semantic expansion on the description text input by the user. The simple description text can be transformed into a long text containing more information about human body attributes, clothing details, pose features, scene adaptation, etc., to obtain a target description text. For example, following the description specification of "human body attributes - clothing details - pose features - scene adaptation", this can more specifically guide the image - generation model to generate a model image that meets the expectations.
[0057] Next, the obtained mask matrix, target description text, and initial scene image can be fed into the image generation model together. The image generation model can encode the initial scene image to obtain a scene graph hidden vector, encode the target description text to obtain a text feature vector, combine the mask matrix with the initial scene image and the text feature vector to form a conditional input. During the reverse denoising process, denoising can be gradually performed according to the conditional input until a model image matching the scene and the target description text is generated, that is, the initial object scene image. Finally, the initial scene image and the initial object scene image can be fused. The initial scene image and the initial object scene image can be fused according to methods such as the weighted summation formula. The fusion process can combine the generated initial object scene image into the initial scene image to ensure the naturalness and consistency of the human model and the scene.
[0058] In the embodiment of the present application, first, in response to a region selection instruction for a target region in the initial scene image, a mask matrix of the target region and a target description text of the target object are obtained. The target description text is used to describe the characteristics of the target object. Next, the mask matrix, the target description text, and the initial scene image are input into the image generation model together, and an initial object scene image containing the described target object is generated through the image generation model. Finally, the initial scene image and the initial object scene image are fused to generate a final target object scene image, in which the target object is accurately presented within the target region. It is easy to notice that by responding to the user's region selection instruction for the target region on the initial scene image, obtaining the mask matrix of the region and the target description text of the target object, and then inputting the mask matrix, the target description text, and the initial scene image into the image generation model together, the user can achieve this by simply and intuitively performing region selection and description input through the interface, reducing the cost of image generation and the complexity of user operations. With the help of the region selection instruction, the user can directly specify the editing region, and the target description text can guide the image generation model to generate an object image that meets the user's expectations, enabling the user to accurately control the generation process and improving the flexibility of image generation. The fusion of the initial scene image and the initial object scene image ensures the docking between the initial scene and the generated object, improving the realism, immersion, and coherence of the background texture of the image, and thus solving the technical problems in the related art that when generating an image integrating a target object into a scene, the cost is relatively high, the user operation process is relatively complex, and the fusion effect of the target object and the scene is not good.
[0059] In the above embodiments of the present application, fusing the initial scene image and the initial object scene image to obtain the target object scene image includes: obtaining a difference image between the initial scene image and the initial object scene image, where the difference image is used to reflect the difference between the initial scene image and the initial object scene image in the shadow area, and the shadow area is generated according to the relative position of the target object and the scene light; fusing the initial scene image and the initial object scene image based on the difference image to obtain the target object scene image.
[0060] The above difference image may refer to an image generated by comparing the initial scene image and the initial object scene image and the pixel value difference after preprocessing. The difference image can more prominently represent the differences between the initial scene image and the initial object scene image. The obtained difference image can highlight the difference degree between the two images of the initial scene image and the initial object scene image at the corresponding positions. Among them, the pixels with larger differences will have higher values, while the pixels with smaller differences or no differences can be close to zero.
[0061] The above shadow area may refer to the area of the illumination effect generated by the natural interaction of the target object in the initial scene image. For example, it may be the dark or semi-shadow part formed by a human model in a home scene image according to the physical illumination principle. When the human model is placed or generated into the scene, it can interact with the light source and objects in the environment, and some of these interactions are manifested as the shadow effects of the model and its surrounding environment.
[0062] In an alternative embodiment, after generating the initial object scene image, the consistency problem of lighting and shadows can be solved based on the above post-processing process, realizing the natural fusion of the initial scene image and the initial object scene image, and generating the final target object scene image. The process of obtaining the difference image between the initial scene image and the initial object scene image. Specifically, first, Gaussian blur processing can be performed on the initial scene image and the initial object scene image to reduce image details and highlight large light and shadow change regions, facilitating the accurate detection of subsequent shadow differences. The selection of Gaussian blur parameters can consider the image size and the amplitude of lighting changes to ensure the accuracy of the difference image. Subsequently, the pixel-level difference between the two blurred images of the initial scene image and the initial object scene image can be calculated to generate a difference image, which can reflect the changes between the initial scene image and the initial object scene image under lighting conditions, especially the inconsistency between the target object and the original light source in the scene, including the absence or incorrect direction of shadows. Specifically, for example, in the preprocessing stage, Gaussian blur processing may be performed on the initial scene image and the initial object scene image to remove small textures and noises in the image, retaining only large shapes and color blocks for subsequent difference calculation. The images after Gaussian blur processing can be called blurred images. Then, the system can calculate the pixel value differences between the two obtained blurred images, and the sum of absolute values or Euclidean distance can be used to measure the differences between pixels to obtain the difference image.
[0063] Next, based on the difference image, the system can generate a preliminary shadow mask through thresholding, that is, a region mask matrix for the shadow region. The shadow mask can mark the regions with significant lighting changes in the difference image, that is, the shadow areas affected by the relative positions of the target object and the light source. After obtaining the accurate shadow mask, the system can selectively fuse the initial scene image and the initial object scene image based on the difference image to correct or add the shadow of the target object, ensuring that the lighting conditions of the target object match the scene environment. The system can fuse only the target object and the shadow region of the target object according to the shadow mask, retaining the original lighting and texture information of the background region and avoiding unnecessary changes to the appearance of the background commodity.
[0064] In the above process, through the analysis of the difference image and the accurate identification of the shadow region, the system can automatically correct or add the shadow of the target object, ensuring that the lighting conditions of the target object are consistent with the scene environment, enhancing the realism and immersion of the image. It can perform fusion processing only on the shadow region, avoiding changes to the texture and light and shadow of the background commodity, maintaining the integrity of the background information, ensuring the visual coherence and naturalness of the generated image, and improving the overall quality of the image.
[0065] In the above embodiments of the present application, the initial scene image and the initial object scene image are fused based on the difference image to obtain the target object scene image, including: performing binarization processing on the difference image based on a preset threshold to obtain a region mask matrix of the shadow region; inputting the initial object scene image into a segmentation model, and using the segmentation model to segment the target object in the initial object scene image to obtain an object mask matrix of the target object; merging the region mask matrix and the object mask matrix to obtain a merged mask matrix; and fusing the initial scene image and the initial object scene image based on the merged mask matrix to obtain the target object scene image.
[0066] The above preset threshold can be set as an empirical value determined in the experiment to ensure the accurate extraction of the shadow region, and the specific value of the preset threshold is not limited here.
[0067] The above segmentation model may refer to a deep learning model for image segmentation tasks, which can accurately separate different objects or regions in an image and generate masks of these objects or regions. In the present application, the segmentation model can be used to accurately identify and extract the target object itself and the shadow region of the target object from the generated initial object scene image containing the target object, providing a basis for subsequent image fusion and improvement.
[0068] In an alternative embodiment, in order to accurately extract the shadow region, binarization processing can be performed on the difference image based on a preset threshold, and the pixel points with significant illumination changes are marked as the shadow region to generate a region mask matrix. The preset threshold can be adjusted through experiments to ensure that both the shadow boundary can be highlighted and the background changes are not overcaptured, thereby improving the accuracy of shadow recognition. The initial object scene image contains the target object and background information. In order to retain the details of the target object during the fusion process, a pre-trained segmentation model can be used to accurately segment the target object. The segmentation model separates the target object in the initial object scene image from other background elements and generates an object mask matrix. The object mask matrix can be marked as 1 in the target object region and 0 in the background region, providing accurate object boundary information for subsequent image fusion.
[0069] Next, after obtaining the region mask matrix and the object mask matrix, the two matrices, namely the region mask matrix and the object mask matrix, can be merged through logical operations to generate a merged mask matrix, ensuring that both the shadow region and the target object region are covered, providing a comprehensive control signal for the fusion process. Subsequently, the initial scene image and the initial object scene image can be fused based on the merged mask matrix, enabling the target object and its shadow to naturally blend into the scene while avoiding unnecessary changes to the background details.
[0070] In the above process, through differential image analysis and threshold binarization processing, the regional mask matrix of the shadow area can be accurately identified and generated, ensuring that the shadow of the target object is consistent with the lighting conditions in the scene, significantly enhancing the realism and immersion of the image. The combination of the object mask matrix and the shadow area mask achieves precise control of the target object and its shadow, avoiding unnecessary changes to the texture and light and shadow of the background commodities, maintaining the integrity of the background information, and improving the overall visual effect of the image.
[0071] In the above embodiment of the present application, the initial scene image and the initial object scene image are fused based on the combined mask matrix to obtain the target object scene image, including: performing a complement operation on the combined mask matrix to obtain the background retention matrix of the initial scene image, where the background retention matrix is used to retain the unedited area in the initial scene image; determining the first product of the background retention matrix and the initial scene image, and determining the second product of the combined mask matrix and the initial object scene image; obtaining the target object scene image based on the sum of the first product and the second product.
[0072] In an alternative embodiment, the background retention matrix can be generated by performing a complement operation on the combined mask matrix. The complement operation can convert the areas with a value of 1 in the combined mask matrix to 0, and the areas with a value of 0 to 1, thereby obtaining the background retention matrix. The background retention matrix marks the pixel positions of the unedited areas in the initial scene image, ensuring that these areas are completely retained during the fusion process and avoiding unnecessary impacts on the appearance and texture of the background commodities. After obtaining the background retention matrix and the combined mask matrix, the first product and the second product can be calculated. The first product matrix is the product of the background retention matrix and the initial scene image, representing the unedited background area retained in the initial scene image. The second product matrix is the product of the combined mask matrix and the initial object scene image, representing the updated information of the target object and the shadow area of the target object. In this way, the control effects of the combined mask matrix and the background retention matrix can be utilized to achieve precise selection and retention of image information, ensuring that the fusion process neither misses the details of the target object nor destroys the original information of the background area. Finally, the first product matrix and the second product matrix can be added to generate the target object scene image, which not only contains the updated target object information but also retains the background details in the scene, while ensuring the visual coherence and realism of the overall image.
[0073] In the above process, the background retention matrix generated through the filling operation and the accurate calculation of the product matrix can accurately retain the background details in the initial scene image, avoid texture damage, maintain the overall coordination of the scene and the consistency of product display. By using the image product and fusion technology guided by the combined mask matrix, it ensures that the generated target object and its shadow are naturally integrated into the scene, avoids problems such as hard edges or inconsistent lighting, significantly improves the realism and immersion of the image. Through matrix operations, selective retention and fusion of image information are achieved, avoiding complex image redrawing or manual correction processes, significantly reducing the time-consuming of image processing, improving production efficiency, and meeting the requirements for image update speed in scenarios such as home furnishing e-commerce.
[0074] In the above embodiments of the present application, based on the sum value of the first product and the second product, a target object scene image is obtained, including: obtaining an initial fusion image based on the sum value of the first product and the second product; performing enhancement processing on the region where the target object is located in the initial fusion image to obtain the target object scene image.
[0075] In an alternative embodiment, the first product and the second product can be subjected to a pixel-by-pixel addition operation to generate an initial fusion image. The initially fused image generated by the fusion may have deficiencies in the detail expression of key regions such as the face, hands, and feet of the target object. To further enhance the visual expressiveness of the target object, a detail redrawing technique can be used. By using a high-resolution redrawing model and focusing on the target object region, especially the parts with rich details such as the face and hands, fine processing is carried out. Through in-depth analysis and adjustment of the texture, light and shadow, and color of the target region, the clarity and naturalness of these regions are improved, ensuring that the target object is more vivid and realistic visually and more integrated with the scene.
[0076] In the above process, the application of the detail redrawing technique, especially the precise improvement of key regions such as the face and hands of the target object, significantly enhances the visual expressiveness of the object, increases the realism and attractiveness of the image, can improve the aesthetic degree of the target object, further enhances the interaction between the target object and the scene, and makes the image more immersive. Through the combination of precise mask matrix control and the detail redrawing technique, the natural integration of the target object and the home scene is achieved, while the detail expression of the target object is improved, overcoming the limitations of the prior art in background fidelity and object details, and providing a more efficient and attractive image display solution for scenarios such as home furnishing e-commerce.
[0077] In the above embodiments of the present application, obtaining the difference image between the initial scene image and the initial object scene image includes: respectively performing blurring processing on the initial scene image and the initial object scene image to obtain the first blurred image of the initial scene image and the second blurred image of the initial object scene image; determining the difference image between the first blurred image and the second blurred image.
[0078] In an alternative embodiment, the initial scene image and the initial object scene image can be blurred respectively in various ways such as Gaussian blur, median blur, adaptive blur, etc. Among them, median blur can slide a window and select the median of the pixel values within the window to replace the central pixel of the window, which can preserve the image edges. When dealing with complex texture backgrounds that appear in home scenes, median blur can better maintain the edge sharpness and avoid texture loss caused by the blurring process. Adaptive blur can dynamically adjust the blur intensity according to the local characteristics of the image, such as texture, illumination, color change, etc., and can intelligently process different regions in the image. For the dynamically changing light and shadow and materials in home scenes, it can provide a more natural blur effect and the ability to identify differences.
[0079] Optionally, the following further explains using the Gaussian blur processing method. The Gaussian blur algorithm can be applied to the initial scene image and the initial object scene image respectively to generate a first blurred image and a second blurred image. Gaussian blur performs weighted averaging on the image through a convolution kernel, which can effectively smooth the image, weaken the texture and details, and thus highlight the large-scale illumination changes in the image. This preprocessing step enables the system to focus on capturing the illumination differences between the target object and its shadow and the environment, without being interfered by complex texture details. When choosing the Gaussian blur parameters, the selection of the standard deviation of the Gaussian kernel can consider the image size and the amplitude of the illumination change. For high-resolution images, a larger standard deviation of the Gaussian kernel helps to capture more global illumination information; while in scenes with relatively local illumination changes, a smaller standard deviation of the Gaussian kernel can more accurately identify the difference regions. The optimal parameters determined through experiments can be used to balance the retention of details and the identification of illumination changes.
[0080] It is also possible to dynamically adjust the blur parameters, and the standard deviation of the Gaussian kernel can be automatically improved according to the complexity of the input image and the illumination conditions to meet the requirements for generating difference images in different scenarios. After generating the first blurred image and the second blurred image, the difference image can be determined by calculating the pixel difference between the first blurred image and the second blurred image. Specifically, the two images of the first blurred image and the second blurred image can be compared pixel by pixel to highlight the regions of illumination or shadow changes caused by the addition of the target object. The difference image can not only include the shadow part of the target object but also reflect the illumination inconsistency between the object and the environment, providing key information for subsequent fusion and improvement operations.
[0081] In the above process, blurring, as a preprocessing step for generating the difference image, effectively reduces the complexity of image calculation, speeds up the efficiency of generating the difference image, reduces the need for manual adjustment and intervention, significantly improves the speed of image processing, and meets the requirements for rapid image update and high-quality output in scenarios such as home e-commerce. The generation of the difference image enables the system to accurately identify the lighting difference between the added target object and the environment, provides guidance for subsequent fusion and detail improvement, can significantly enhance the authenticity of shadows and the consistency of lighting in the image, and enhances the overall naturalness and authenticity of the image.
[0082] In the above embodiments of the present application, the method further includes: obtaining a plurality of first sample object scene images and sample object mask matrices corresponding to the plurality of first sample object scene images, wherein different first sample object scene images include first sample objects in different poses; inputting the plurality of first sample object scene images into an initial segmentation model, and using the initial segmentation model to segment the first sample objects in the plurality of first sample object scene images to obtain a predicted object mask matrix of the first sample objects; adjusting the first model parameters of the initial segmentation model based on the sample object mask matrix and the predicted object mask matrix to obtain a segmentation model.
[0083] The above-mentioned plurality of first sample object scene images may refer to a data set collected and prepared in advance for training the segmentation model. For example, in the scenario of generating home images containing model figures, the plurality of first sample object scene images may be a large number of human body pose images in the home environment background.
[0084] In an alternative embodiment, a plurality of first sample object scene images and corresponding sample object mask matrices may be collected in advance. These first sample object scene images may represent images of the target object in different poses and different backgrounds. The sample mask matrix can accurately mark the position of the target object in the image, distinguish the target object from the background, and the first sample object scene images and the sample object mask matrix constitute the data set required for training the segmentation model. The following preprocessing process may also be performed on the first sample object scene images. To improve the generalization ability of the segmentation model, data augmentation operations such as random rotation, flipping, and scaling can be performed on the sample images to generate more variant training data, so that the segmentation model can maintain stable segmentation performance when facing target objects in different poses. At the same time, considering the uneven distribution of target objects in different poses in the data set, oversampling or undersampling techniques can be used to ensure that samples of various poses can be fully learned by the model, and to avoid overfitting or insufficient generalization ability of the segmentation model to certain poses.
[0085] Next, the pre - processed multiple first - sample object scene images can be input into the initial segmentation model. The initial segmentation model is used to segment the target objects in the multiple first - sample object scene images and output a predicted object mask matrix. The predicted object mask matrix can be compared with the sample object mask matrix to evaluate the segmentation accuracy of the model. According to the error feedback, the first model parameters of the initial segmentation model can be adjusted to improve the segmentation performance, so that the initial segmentation model can output an accurate object mask matrix for target objects in different poses, completing the training process.
[0086] In the above process, by introducing sample data containing target objects in various poses, as well as multiple rounds of training and parameter adjustment, the segmentation model can learn to handle a wider range of human poses, improving the generalization ability of the model. The comparison between the sample object mask matrix and the predicted object mask matrix enables the system to quantify the segmentation error of the model. Through parameter adjustment, the system can gradually improve the segmentation performance of the model, ensuring that the boundaries of the target objects are clearer and the separation from the background is more accurate.
[0087] In the above - mentioned embodiments of the present application, obtaining the mask matrix of the target area and the target description text of the target object includes: determining the position coordinates of the target area according to the area selection instruction and obtaining the initial description text of the target object; performing spatial coordinate mapping on the position coordinates to obtain the mask matrix; inputting the initial description text into the text generation model, and using the text generation model to expand the initial description text to obtain the target description text.
[0088] The above - mentioned text generation model can be a deep - learning model with language understanding and generation capabilities. For example, in the scenario of generating a home image containing a model figure, the text generation model can be a deep - learning model fine - tuned according to the specific requirements of the home scene, and can generate a text description that better meets the requirements of the home environment.
[0089] In an optional embodiment, the pre - processing process can process user input, accurately identify the target area, and expand the model description text to guide the subsequent image - to - image processing. The system can respond to the received area selection instruction. The area selection instruction can be that the user selects the area of the home scene to be edited through an intuitive interface. For example, this operation can include drawing a rectangular box on the initial scene image. The system can parse out the position coordinates of the area selected by the user. For example, it can include the coordinate points of the upper - left corner and the lower - right corner. The obtained coordinate points can be used to define the boundary of the target area, providing position guidance for the subsequent generation of the mask matrix. A simple and friendly user interface can also be designed so that the user can easily and accurately select the target area, reducing operation errors. The system can also have an intelligent expansion function, that is, it can automatically identify and adjust the mask boundary according to the area selected by the user to ensure that the complete body of the target object, including details such as clothing and accessories, is covered.
[0090] Next, the system can perform a spatial coordinate mapping on the position coordinates to construct a mask matrix that matches the target area. The mask matrix can be a two-dimensional array with the same size as the original image, where the pixel values within the target area are marked as 1 to indicate that editing is required, and the pixel values in the background area can be marked as 0 to indicate that the original appearance is to be retained. The generation of the mask matrix can ensure that the subsequent image generation model can perform precise editing within the specified area without affecting other parts of the scene. At the same time, the user can provide only a short initial description text, such as the target object of casual clothing. The system can input the initial description text into a text generation model. For example, a deep learning model fine-tuned with home scene data can be used to generate a more detailed, specific, and better scene-adapted target description text. Here, the text generation model can have a profound understanding of scenes such as home, and can expand the text description according to the scene characteristics to ensure that the generated target object matches the scene style.
[0091] In the above process, the user input can be simplified. Only the area needs to be selected and a short text description provided. The preprocessing module can automatically identify the target area and expand it to generate a detailed text, greatly simplifying the user operation process and improving the user experience. The generation of the mask matrix ensures that the model can accurately identify and edit only within the specified area, avoiding unnecessary changes to the texture, light and shadow, and appearance of the background items, and maintaining the overall aesthetic and authenticity of the scene. The text expansion technology not only enriches the description content but also ensures the scene adaptability of the description, making the generated target object not only have a natural pose but also the details such as clothing and accessories are coordinated with the scene style, enhancing the visual appeal of the image and the user experience.
[0092] In the above embodiments of the present application, the mask matrix, the target description text, and the initial scene image are input into an image generation model, and the initial object scene image is generated by using the image generation model, including: encoding the initial scene image to obtain a scene graph hidden vector; using a text encoder to extract features from the target description text to obtain a semantic feature vector; performing channel merging on the scene graph hidden vector, the semantic feature vector, and the mask matrix to obtain the initial object scene image.
[0093] The above-mentioned scene graph hidden vector refers to a high-dimensional representation form obtained by encoding the input initial scene image through an encoder. For example, in the scenario of generating a home image containing a model figure, it can be a high-dimensional representation form obtained by encoding the input home scene image through an autoencoder. The scene graph hidden vector condenses the main features and information of the image and can be used as the input of the scene condition in the image generation model.
[0094] In an alternative embodiment, the initial scene image can be encoded, which can be accomplished by a pre-trained encoder, such as an autoencoder. The encoder can transform the high-resolution scene image into a low-dimensional scene graph latent vector, which contains the key features of the scene image, such as color, texture, and object layout. The low-dimensional representation not only saves computational resources but also enables the image generation model to better capture the global information of the scene, providing a solid foundation for subsequent editing and image generation. Then, a text encoder can be used to extract features from the target description text to generate a semantic feature vector. The text encoder can understand the semantic information in the text description, such as the casual wear attribute in the casual wear target object, as well as details such as the gender and action of the target object, and encode this information into vector form for fusion with the features of the scene image. The semantic feature vector contains information such as the appearance, pose, and style of the target object.
[0095] Next, the scene graph latent vector, the semantic feature vector, and the mask matrix can be merged channel-wise, which can be done in the latent space to ensure the combination of scene, text, and editing region information. The image generation model uses these fused features to generate an initial object scene image that matches the target description text. The above process utilizes the global modeling ability of the image generation model to ensure a high degree of fusion and natural transition between the object and the background in the generated image.
[0096] In the above process, through multi-modal fusion, the image generation model can accurately generate the target object that matches the text description, while ensuring a high degree of fusion between the target object and the scene background, avoiding damage to the background texture during model image generation, and improving the realism and immersion of the image. The introduction of the text encoder enables the image generation process to be controlled according to the text description input by the user. The user can flexibly set the pose, clothing, and appearance of the target object through the text, meeting the diverse needs of target object display in the home furnishing e-commerce scenario, not only enhancing the visual effect of the image but also simplifying the user operation and realizing an efficient and high-quality image generation function.
[0097] In the above embodiments of the present application, the method further includes: obtaining a sample scene image, a second sample object scene image, and a sample description text of a second sample object, where the second sample object scene image contains the second sample object; inputting the sample scene image and the sample description text into an initial image generation model, and using the initial image generation model to generate a predicted object scene image; adjusting the second model parameters of the initial image generation model based on the second sample object scene image and the predicted object scene image to obtain an image generation model.
[0098] In an alternative embodiment, sample scene images and corresponding second sample object scene images can be collected in advance, where the second sample object scene images contain images of real target objects incorporated into a home scene in different poses and costumes. At the same time, sample description texts related to each second sample object scene image can be obtained, and these texts describe information such as the attributes, costumes, and poses of the target objects. Constructing such a dataset provides rich scene and object examples for the training of the image generation model, ensuring that the image generation model can learn diverse scene features and object descriptions. Then, the sample scene images and sample description texts can be input into the initial image generation model, and the initial image generation model generates a predicted object scene image by learning the correlation between the sample images and the text descriptions. The initial image generation model can encode the sample scene images, extract the features of the sample scene images, then fuse these features with the semantic features of the sample description texts, and finally decode to generate the predicted image. The predicted object scene image can be close to the second sample object scene image to reflect the learning effect of the model.
[0099] Next, the second sample object scene image can be compared with the predicted object scene image to identify the differences between the two, and these differences can be reflected in pixel-level errors in the image, such as unnatural poses, missing costume details, improper fusion with the background, etc. These differences can be utilized to adjust the second model parameters of the initial image generation model through the backpropagation algorithm to reduce the gap between the predicted image and the real image, enabling the initial image generation model to achieve an appropriate accuracy when generating the predicted image.
[0100] In the above process, by using diverse training data, the obtained image generation model can not only generate target objects that match the input text description but also ensure that they naturally blend into different scenes, enhancing the generalization ability and adaptability of the image generation model. Through the adjustment of model parameters, the generated image results are closer to the real scene, and details such as the appearance, costume, and pose of the target object are more realistic, improving the visual effect and user experience of the image. Through the text-image matching learning during the training process, the obtained image generation model can more deeply understand the text description and accurately convert the text description into image content, enhancing the matching degree and consistency between the scene image and the text description.
[0101] The technical solution proposed in this application will be described below in combination with an optional embodiment. This application proposes a method and system for quickly generating intelligent models for home scenes based on a transformer image-to-image model. The application scenarios of this application can at least include: merchants upload product scene pictures, such as living rooms and bedrooms, select the target area and input model descriptions, such as clothing, attributes, etc., and the system automatically generates intelligent model images integrated with the scene, replacing traditional shooting, so as to achieve a simpler interaction. After the user selects the area, a short text is input, such as "casual wear model", without complex parameter adjustment. Background fidelity, based on the image-to-image model, can keep the original scene lighting and texture unchanged, and only modify the target area. Quick generation, the generation time for a single image is less than 1 minute, with high efficiency. And highly customizable, giving users a high degree of freedom and control, and the pose and appearance of the model can be flexibly controlled through text to meet diverse display needs.
[0102] Figure 3 is a schematic diagram of an optional user operation interface according to an embodiment of the present application, as Figure 3 shown. In the user operation interface, the user can view the upload history in the scene picture part, upload the scene picture, and can bind the product by importing the products in the store; the user can input model attributes, such as the content of "female" shown in the figure; the model pose can be generated by a template or by text. The user can select the position where the expected model appears, and can choose the selection method of box selection or point selection. The user can supplement the model pose or select the model pose. There are prompt words shown in the figure, "Please supplement the model pose, such as [sitting on the sofa, hands on the sofa], 0 / 50, standing pose 1, standing pose 2, sitting pose 1, sitting pose 2"; the user can select different model clothes shown in the figure, and finally can click "Start Generation" to generate.
[0103] Figure 4 is a schematic diagram of an optional generation effect display interface according to an embodiment of the present application, as Figure 4 shown. The generation effect display interface can have a prompt word "AI human model; customize the exclusive virtual model for your product. Try the products below and start experiencing"; the user can view the generated target object scene image in the generation effect display interface, or can select different initial scene images; in the generation result part, the user can select "My Creation" or "Team Creation" to view or continue to edit the generated target object scene image.
[0104] In the current wave of digital transformation embraced by the home decoration industry, online display platforms have become an important element in attracting consumers' attention and improving the shopping journey. To stand out in the field of home visual promotion, many home brands and e-commerce platforms are actively exploring innovative paths, using highly attractive model images to vividly showcase the charm of products such as furniture and decorations. In view of the above challenges, this application proposes an intelligent model generation system designed for home e-commerce platforms, adopting an intelligent model generation method based on a transformer image-to-image model. The system allows users to select model poses, appearance features, clothing combinations, and specify scene images and model positions. Then, the system accurately generates matching intelligent model images, significantly enhancing the aesthetic expressiveness and authenticity of the output images, and providing a cost-effective model image production solution for merchants.
[0105] The technical solution proposed in this application has high operational convenience. Users only need to simply select an area and input text to easily describe the position and appearance preferences of the intelligent model, greatly simplifying the operation process. The scene is highly realistic. Without disturbing the product display information in the original scene image, the intelligent model can seamlessly integrate, maintaining the overall consistency and visual integrity of the scene and enhancing the user experience. It provides a good visual experience, can generate more natural human postures and background interaction effects, and better attracts users in the e-commerce scenario.
[0106] This application proposes an intelligent image generation method based on a transformer architecture. The specific implementation process includes a preprocessing module, an image-to-image core processing module, and a post-processing improvement module. These modules work together to achieve efficient image editing of "select and generate" in the home scene. The following elaborates on each technical link in detail.
[0107] Figure 5 It is a schematic diagram of an optional method for generating a scene image of a target object according to an embodiment of this application. As Figure 5 shown, the processing result of the mask generation module for the position coordinates, the processing result of the text expansion module for the initial description text, and the initial scene image can be input into the image generation model. The output result of the image generation model can obtain the target object scene image after passing through the illustrated mask detection module, image overlay fusion, and detail redrawing.
[0108] The preprocessing module includes an adaptive mask generation unit that receives the rectangular selection coordinates (x1, y1, x2, y2) input by the user and constructs a high-dimensional mask matrix M ∈ {0, 1} through spatial coordinate mapping H×W, where the parameter H can represent the height of the mask matrix, and the parameter W can represent the width of the mask matrix. The pixel values of the target area can be set to 1, and the background area to 0, which can greatly reduce the time-consuming of user interaction operations. The semantic expansion unit uses a deep learning model fine-tuned with home scene data to expand the semantics for the short text description provided by the user, such as "casual wear model". The training data can construct <short text, long text, scene graph> pairs, with 250,000 groups here, and the quantity is not limited; among them, the expanded text can follow the description specification of "human body attributes - clothing details - pose features - scene adaptation". This module greatly improves the compatibility of the generated text with the target environment by aligning the input text with the scene graph.
[0109] The function of the image-to-image generation model is to generate a model image that matches the given text description based on the masked part of the scene graph. The core network structure of this model is constructed based on the Transformer, aiming to build a multi-modal conditional coupling generation framework. Specifically, the Transformer-based image generation framework adopted in this application realizes the global modeling ability. This architecture can more effectively handle the limitations in complex scene generation and includes the following technical details: The multi-condition fusion mechanism. The input conditions include three types of heterogeneous data. First, the input scene graph is encoded by an autoencoder to obtain the scene graph latent vector. The semantic feature vector extracted by the contrastive language-image pre-training text encoder is used. The preprocessed mask matrix is cross-modally aligned with the previously processed features through channel concatenation and injected into the latent space together. The training method. In terms of the dataset, home scene graph data pairs can be used, which can include <scene graph, model image>, and the corresponding model description text. During the training process, the model image is used as the supervision information to improve the similarity between the generated image and the real model image, thereby enhancing the generation ability of the model.
[0110] Figure 6 is a schematic diagram of an optional generation of a target object scene image including a model structure according to an embodiment of the present application, such as Figure 6As shown, in an optional process of generating a scene image of a target object, the initial description text is subjected to feature extraction to obtain pooled features, which are then processed by a multi-layer perceptron. The guidance vector is processed through time steps, sinusoidal position encoding, and a multi-layer perceptron, and then can be input into a multi-modal diffusion block, a single diffusion block, and a modulation layer through the illustrated processing; the initial description text is processed by a text generation model to obtain a target description text, which can be input into parts such as the illustrated multi-modal diffusion block for processing; the latent vector can be input into parts such as the illustrated multi-modal diffusion block after being processed by a linear layer; the input identifier can be input into the multi-modal diffusion block and the single diffusion block after being subjected to rotational position embedding; there are multiple multi-modal diffusion blocks and multiple single diffusion blocks as illustrated, and the output result can obtain the illustrated latent vector through the modulation layer and the linear layer. The illustrated latent vector can be obtained through T times of iterative (×T) processing, and then through a shape transformation layer and an autoencoder decoder, a scene image of the target object can be obtained.
[0111] The above-mentioned initial description text (Prompt) can provide a text description for guiding the generation of an AI model, including features such as the appearance, clothing, and actions of the model. Feature extraction can be used to encode the text description and convert the initial description text into a feature vector that can be understood by the model; the pooled feature (Pooled) can extract key information for conditional generation; the multi-layer perceptron (MLP) can be used to process and transform the output features to meet the input requirements of the model. The guidance vector (Guidance) can integrate various conditional information to guide the decision-making in the image generation process; the time step (Timestep) can be the time identifier in the denoising process of the image generation model, used to control the order and degree of the denoising steps; the sinusoidal position encoding (Sinusoidal Encoding) can provide position information for the image generation model to help the model understand the relative relationships of different positions in the image during the generation process. The text generation model can be used to expand and refine the input text description; the latent vector (Latent) can be the image feature representation output by the image encoder, used for image generation in the latent space; the linear layer (Linear) can be used to transform the latent vector and the conditional vector to adapt them to the input format of the image generation model.
[0112] The input identifier (Ids) can represent the sequence identifier input into the model, including the conditional information and position encoding in the diffusion process, used for the model generation and denoising process. The above-mentioned rotational position embedding (RopE) can provide rotation-invariant position information to help maintain symmetry and direction perception when processing images, especially when processing human models. The image generation model can simultaneously process image, text, and position information to achieve image generation under multi-modal conditions, and can be responsible for performing denoising operations at each time step of the diffusion process to gradually reconstruct the image.
[0113] The Modulation layer can be used to adjust the weights and biases of the model, enabling the model to dynamically adjust the generation process according to the conditional vector and ensuring that the output image meets specific conditions. The Shape Transformation layer (n, 4c, h, w → n, c, 2h, 2w) can be responsible for adjusting the shape of the latent vector, converting from n samples, each with 4c channels and an h×w size, to n samples, each with c channels and a 2h×2w size, achieving upsampling of the image and enhancement of details. The autoencoder decoder can convert the generated latent vector back into a visualizable image form, reconstructing the image details and realizing the final image generation. The finally generated AI model image has an appearance, clothing, and actions that highly integrate with the input scene graph, meeting the display requirements of the home scene.
[0114] In the entire process from the input conditions to the final image generation described above, first, the input initial text description (Prompt) is transformed into pooled features through feature extraction. At the same time, the text generation model refines the text description to generate more detailed guidance information. The image input is converted into a latent vector through the autoencoder. Subsequently, the latent vector and conditional information are transformed through the Linear layer and are passed into the image generation model together with the Timestep and Sinusoidal Encoding. Then, multi-modal image generation can be performed based on the conditional information (Guidance) and input identifiers (Ids), where the Modulation layer helps the model dynamically adjust the generation strategy. The Shape Transformation layer performs upsampling operations during the generation process to ensure rich details. Finally, the VAE decoder converts the latent vector into the output image (Image), completing the entire image generation process. The above process makes full use of the advantages of the Transformer image-to-image model in processing complex image tasks, combining the encoding of images and texts by feature extraction and the refinement of text descriptions by the text generation model, so as to be able to generate high-quality AI model images that not only conform to the text description but also highly integrate with the scene graph, meeting the needs of the home e-commerce industry.
[0115] Post-processing module. In the post-processing stage, in order to effectively solve the shadow adaptation problem between the generated region and the background and ensure the integrity of the human body, this application introduces a set of fine multi-level mask calculation processes. This process is elaborated in detail in the following core steps: including shadow mask generation, this application adopts a simple and effective method to pre-extract the human mask containing shadows. In this link, the main purpose is to accurately identify and extract the changing regions in the generated image, including shadows, for subsequent merging with the human segmentation mask. The specific steps are as follows: Gaussian blur difference calculation. First, for the generated image (I gen ) and the original image (I src) Perform Gaussian blur processing separately. The standard deviation of the Gaussian kernel can be set to 5, that is, (σ = 5). Here, the standard deviation of the Gaussian kernel can also be determined according to actual needs and is not limited here; obtain the blurred image (I gen blur and I src blur ). This step aims to reduce the detailed noise in the image and facilitate subsequent difference calculation. Subsequently, calculate the L1 norm difference map (M diff ) between these two blurred images. This difference map reflects the difference between the generated image and the original image in the shadow area. Perform binarization processing. Based on the difference map (M diff ), use the thresholding method to generate a mask. Subsequently, delete the mask that misdetects the background area through erosion and dilation operations to ensure the integrity of the human body area. Finally, obtain the mask (M shadow ) containing the shadow area. The setting of the threshold τ is the empirical value determined in the experiment to ensure the accurate extraction of the shadow area. Through this step, a clear shadow area mask can be obtained, providing a basis for subsequent mask merging.
[0116] Generation of the human body segmentation mask. The generation of the human body segmentation mask is an important step to ensure the integrity of the human body in the generated image. This application uses a self-trained saliency segmentation model for human body mask segmentation. Model training and application: Use a training set containing human body annotation data in a home scene, covering various postures such as sitting, lying, and standing, to train the segmentation model. After training, use the model scene map as the input, and the model outputs a complete human body mask (M body ). This mask accurately reflects the contour and posture of the human body in the scene.
[0117] Mask merging and image fusion. After obtaining the shadow mask (M shadow ) and the human body segmentation mask (M body ), merge these two masks with the initial mask through logical operations to generate the final mask (M final ), and fuse the generated image and the original image accordingly. Logical operation and mask merging: First, according to the logical OR (∪) operation, merge the segmentation mask and the shadow mask. The merging strategy can be as follows:
[0118] M final = M shadow ∪M body ;
[0119] Through the above merging, the specified generation area can be retained, and at the same time, the human body contour and the shadow area are fused to ensure a natural transition of the generated image in terms of shadow and human body integrity.
[0120] Next, overlay fusion can be performed. After obtaining the final mask, the generated image (I gen ) is fused with the original image (I src ) using the overlay fusion technique. The fusion formula can be as follows:
[0121] I fused = I src ·(1 - M final ) + I gen ·M final ;
[0122] This step ensures the complete retention of background information, avoids texture damage, and makes the fused image more natural and harmonious in visual effect. Finally, detail redrawing can be performed. The fused image (I fused ) is redrawn in detail, especially enhancing the details of key areas such as the face and hands and feet. By applying the redrawing model, the generated human model area is finely processed to further improve the image fidelity and visual quality.
[0123] Figure 7 is a schematic diagram showing the image effect of an optional post - processing process according to an embodiment of the present application. As Figure 7 shown, it respectively shows the difference mask (diff - mask); erosion, with the number of iterations being 3 (erode, iter = 3); extended mask, with the growth parameter being 15 and the blur parameter being 10 (mask - expand), grow = 15, blur = 10); the initial object scene image; the image effect of the image after the overlay process and the target object scene image obtained after detail redrawing.
[0124] The technical solution proposed in this application has improved background fidelity, showing significant advantages in background processing. It can effectively retain the original background information, avoid unnecessary damage or changes to background items in the overlapping area with the generated model, and ensure the consistency of the appearance of background products. In addition, the application of the post-processing technology of overlaying back further ensures a high retention of background materials and significantly improves the overall quality of the image. The naturalness and richness of human poses are enhanced. This application also performs excellently in the generation of human models. The generated human body and shadows are not only realistic and natural, significantly improving the vividness of the image, but also the human actions are rich and diverse, with stronger adaptability to the scene. This characteristic enhances the usability of the generated pictures in e-commerce scenarios. The user-friendliness is improved. This application has made remarkable progress in terms of user-friendliness. By improving the operation steps, the originally cumbersome process is simplified to only two operations, greatly reducing the operation difficulty and time cost of users. At the same time, it avoids cumbersome operations such as users needing to provide reference models or complex pose control methods, enhancing the user experience and satisfaction. In addition, this application brings a more economical and efficient solution to the field of image generation and processing by reducing the cost and time consumption of generating a single image.
[0125] In preprocessing, this application introduces a technology that converts the user-selected coordinates into a rectangular binary mask, and combines it with a fine-tuned deep learning model to perform a detailed expansion of the input short text to match the input scene picture. This innovative measure not only greatly simplifies the user operation process and reduces the usage threshold, but also more accurately captures and meets the user's personalized image generation needs, generating a detailed text description that matches the scene. The user only needs to provide a concise text description and perform a basic selection action, and the system can automatically generate an image that meets the expectations. In postprocessing, by integrating mask overlay and detail redrawing technologies, the beauty of the output image in terms of facial and limb details can be improved, and the coherence between the generated area and the background area can be enhanced. Finally, the overlay of the generated image with the original input picture largely retains the material and appearance of the products in the background area.
[0126] This application proposes a new strategy for generating intelligent models in home scenes based on a deep image conversion transformer model. This strategy innovatively accepts the scene picture, selected area coordinates, and short text description provided by the user as inputs, and generates a model scene picture through target area editing. It effectively reduces the deformation problem of background products and improves the usability of the generated images. At the same time, compared with the method of generating models using rigid pose control, the models generated by this strategy have smoother and more natural movements, greatly enhancing the expressiveness and vividness of the images, and providing a new solution for generating intelligent models in home scenes.
[0127] Optionally, Figure 8 is a structural block diagram of a computing device according to an embodiment of the present application. As Figure 8As shown, the computing device A may include: one or more (only one is shown in the figure) processors 102, a memory 104, and a peripheral interface 106. Among them, the processor 102, the memory 104, and the peripheral interface 106 are interconnected via a bus 108.
[0128] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal A through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0129] The processor can call the information and application programs stored in the memory through a transmission device to execute the steps in the embodiments.
[0130] Embodiments of the present application can provide an electronic device. Figure 9 It is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 9 shown, the electronic device may include: an input / output device 902; a memory 904, and a processor 906. Among them, the processor 906 is connected to the input / output device 902 and the memory 904 via a bus 908.
[0131] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal A through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0132] The processor can call the executable program stored in the memory through the transmission device to execute the following method: in response to a region selection instruction for a target region on an initial scene image, obtain a mask matrix of the target region and a target description text of the target object, where the target description text is used to describe the target object; input the mask matrix, the target description text, and the initial scene image into an image generation model, and use the image generation model to generate an initial object scene image, where the initial object scene image contains the target object; fuse the initial scene image and the initial object scene image to obtain a target object scene image, where the target region in the target object scene image contains the target object; and execute the methods in various embodiments of the present application.
[0133] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.
[0134] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a processing unit to execute the methods in various embodiments of the present application.
[0136] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. This program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0137] An embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium may be used to store the program code executed by the method provided in the above embodiment.
[0138] Optionally, in this embodiment, the above storage medium may be located in a computing device.
[0139] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method described in any one of the above embodiments.
[0140] An embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and when the computer program is executed by a processor, it implements the method provided in the above embodiment.
[0141] The above computer program product may refer to a software program that has been written, tested, and released, and can run on a computer or other devices. The computer program product may include application programs, operating systems, tool software, etc., and is used to implement specific functions or solve specific problems.
[0142] An embodiment of the present application further provides a computer program product. Optionally, the above computer program product may include a non-volatile computer-readable storage medium, and the above non-volatile computer-readable storage medium may be used to store a computer program, and when the computer program is executed by a processor, it implements the method provided in the above embodiment.
[0143] The above non-volatile computer-readable storage medium may refer to a medium for storing data. The non-volatile computer-readable storage medium can retain data without loss when powered off and can be used to store data for long-term preservation, such as operating systems, application programs, and user files. The non-volatile storage medium may include hard disk drives, solid-state drives, optical discs, and flash storage devices, etc.
[0144] An embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the above computer program is executed by a processor, it implements the method provided in the above embodiment.
[0145] The above computer program may refer to a set of instructions used to tell a computer to perform specific tasks or operations. The computer program can be written by a programmer using a specific programming language and may include algorithms, data structures, logic, and control flows, etc. The computer program can be used for various purposes, including application software, operating systems, etc.
[0146] In the above embodiments of the present application, the descriptions of the various embodiments each have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0147] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0148] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0150] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical disks and other various media that can store program codes.
[0151] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An image processing method, characterized in that, including: responding to a region selection instruction for a target region on an initial scene image, obtaining a mask matrix of the target region and a target description text of a target object, where the target description text is used to describe the target object; inputting the mask matrix, the target description text, and the initial scene image into an image generation model, and using the image generation model to generate an initial object scene image, where the initial object scene image contains the target object; fusing the initial scene image and the initial object scene image to obtain a target object scene image, where the target object is included in the target region in the target object scene image.
2. The image processing method according to claim 1, wherein Fusing the initial scene image and the initial object scene image to obtain a target object scene image, including: obtaining a difference image between the initial scene image and the initial object scene image, where the difference image is used to reflect the difference between the initial scene image and the initial object scene image in a shadow region, and the shadow region is generated according to the relative position of the target object and scene light; fusing the initial scene image and the initial object scene image based on the difference image to obtain the target object scene image.
3. The image processing method according to claim 2, wherein Fusing the initial scene image and the initial object scene image based on the difference image to obtain the target object scene image, including: performing binarization processing on the difference image based on a preset threshold to obtain a region mask matrix of the shadow region; inputting the initial object scene image into a segmentation model, and using the segmentation model to segment the target object in the initial object scene image to obtain an object mask matrix of the target object; merging the region mask matrix and the object mask matrix to obtain a merged mask matrix; fusing the initial scene image and the initial object scene image based on the merged mask matrix to obtain the target object scene image.
4. The image processing method according to claim 3, wherein Fusing the initial scene image and the initial object scene image based on the merged mask matrix to obtain the target object scene image, including: performing a complement operation on the merged mask matrix to obtain a background retention matrix of the initial scene image, where the background retention matrix is used to retain an unedited region in the initial scene image; determining a first product of the background retention matrix and the initial scene image, and determining a second product of the merged mask matrix and the initial object scene image; obtaining the target object scene image based on the sum value of the first product and the second product.
5. The image processing method according to claim 4, wherein Obtaining the target object scene image based on the sum value of the first product and the second product, including: obtaining an initial fusion image based on the sum value of the first product and the second product; performing enhancement processing on a region where the target object is located in the initial fusion image to obtain the target object scene image.
6. The image processing method according to claim 2, characterized in that Obtaining a difference image between the initial scene image and the initial object scene image, including: Blur the initial scene image and the initial object scene image respectively to obtain a first blurred image of the initial scene image and a second blurred image of the initial object scene image; Determine the difference image between the first blurred image and the second blurred image.
7. The image processing method according to claim 3, wherein The method further includes: Obtain a plurality of first sample object scene images and sample object mask matrices corresponding to the plurality of first sample object scene images, where different first sample object scene images contain first sample objects in different poses; Input the plurality of first sample object scene images into an initial segmentation model, and use the initial segmentation model to segment the first sample objects in the plurality of first sample object scene images to obtain predicted object mask matrices of the first sample objects; Adjust first model parameters of the initial segmentation model based on the sample object mask matrices and the predicted object mask matrices to obtain the segmentation model.
8. The image processing method according to claim 1, characterized in that Obtain a mask matrix of the target region and a target description text of the target object, including: Determine the position coordinates of the target region according to the region selection instruction, and obtain the initial description text of the target object; Perform spatial coordinate mapping on the position coordinates to obtain the mask matrix; Input the initial description text into a text generation model, and use the text generation model to expand the initial description text to obtain the target description text.
9. The image processing method according to claim 1, wherein Input the mask matrix, the target description text, and the initial scene image into an image generation model, and use the image generation model to generate an initial object scene image, including: Encode the initial scene image to obtain a scene graph hidden vector; Use a text encoder to extract features from the target description text to obtain a semantic feature vector; Perform channel merging on the scene graph hidden vector, the semantic feature vector, and the mask matrix to obtain the initial object scene image.
10. The image processing method according to claim 9, wherein The method further includes: Obtain a sample scene image, a second sample object scene image, and a sample description text of a second sample object, where the second sample object scene image contains the second sample object; Input the sample scene image and the sample description text into an initial image generation model, and use the initial image generation model to generate a predicted object scene image; Adjust second model parameters of the initial image generation model based on the second sample object scene image and the predicted object scene image to obtain the image generation model.
11. A computing device, characterized in that, Includes: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the method according to any one of claims 1 to 10.
12. An electronic device, characterized in that, Includes: A memory storing an executable program; A processor connected to the memory through a bus for running the program, where when the program runs, it executes the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method described in any one of claims 1 to 10.
14. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the method described in any one of claims 1 to 10.
Citation Information
Cited By
Image processing system
CN121685570A
Image processing system
CN121685570B