Method, device, equipment and storage medium for generating display image of target object

By extracting the subject area image of the target object from the original image and generating candidate background images, and combining feature matching for fusion, the problems of low efficiency of product image generation and single style in the prior art are solved, and display image generation is achieved that is adapted to various categories.

CN117152302BActive Publication Date: 2025-08-12DOUYIN VISION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310967809.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2025-08-12
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

Existing product image generation technology relies on material libraries or image templates, making it difficult to batch generate multiple styles, and manual production takes time and cannot adapt to the display image needs of various categories.

Method used

By extracting the subject area image of the target object from the original image, generating candidate background images based on its description information, combining feature matching to generate display images, reducing dependence on material libraries and image templates.

Benefits of technology

It realizes batch generation of display images of multiple styles, adapts to display scenarios of various categories, improves generation efficiency and style richness, and reduces dependence on existing materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152302B_ABST
    Figure CN117152302B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of image processing technology and proposes a method, apparatus, device, and storage medium for generating a display image of a target object. The method comprises: obtaining an original image of the target object and extracting a main body area image of the target object from the original image; obtaining description information of the target object and, based on the description information, generating a candidate background image suitable for the target object; and fusing the main body area image with the candidate background image to generate a display image of the target object. By implementing the technical solution of the present disclosure, it is convenient to batch generate candidate background images of various styles, no longer relying on existing material libraries or image templates, and can adapt to various display image generation scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, apparatus, device, and storage medium for generating a display image of a target object. Background Art

[0002] With the widespread adoption and development of internet technology, online shopping has become a mainstream trend, making the generation of product images crucial. Currently, product images are primarily generated manually or by machines. Manual creation is time-consuming and relies on image resources, making it difficult to generate in bulk. While machine creation can generate images in bulk, it still relies on existing image libraries or manually defined image templates, limiting the variety of product images. Summary of the Invention

[0003] In view of this, the embodiments of the present disclosure provide a method, apparatus, device, and storage medium for generating a display image of a target object, so as to solve the problem that the generation of product images depends on a material library or image templates.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating a display image of a target object, comprising: obtaining an original image of the target object and extracting a main area image of the target object from the original image; obtaining description information of the target object and generating a candidate background image suitable for the target object based on the description information; and fusing the main area image with the candidate background image to generate a display image of the target object.

[0005] The method for generating a display image of a target object provided by the embodiment of the present disclosure extracts the main area image of the target object from the original image and generates a corresponding candidate background image based on its description information. This can automatically generate a candidate background image for the target object based on the characteristics of the target object itself, facilitate batch generation of candidate background images of various styles, and then fuse the main area image with the candidate background image to obtain a batch of display images. This no longer relies on existing material libraries or picture templates and can be adapted to generation scenarios of display images of various categories.

[0006] In a second aspect, an embodiment of the present disclosure provides a device for generating a display image of a target object, comprising: an image acquisition module for acquiring an original image of the target object and extracting a main area image of the target object from the original image; a description information acquisition module for acquiring description information of the target object and, based on the description information, generating a candidate background image suitable for the target object; and a fusion module for fusing the main area image with the candidate background image to generate a display image of the target object.

[0007] In a third aspect, an embodiment of the present disclosure provides an electronic device, which includes a memory and a processor, wherein the memory is used to store a computer program. When the computer program is executed by the processor, the method for generating a display image of a target object according to the first aspect or any corresponding embodiment thereof is implemented.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, characterized in that the computer-readable storage medium is used to store a computer program, which, when executed by a processor, implements the method for generating a display image of a target object of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The features and advantages of the various embodiments of the present disclosure will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present disclosure in any way. In the accompanying drawings:

[0010] Figure 1 is a flowchart of a method for generating a display image of a target object according to an embodiment of the present disclosure;

[0011] Figure 2 is a flowchart of a method for generating a display image of a target object according to an embodiment of the present disclosure;

[0012] Figure 3 is a flowchart of a method for generating a display image of a target object according to an embodiment of the present disclosure;

[0013] Figure 4 It is a structural block diagram of a device for generating a display image of a target object according to an embodiment of the present disclosure;

[0014] Figure 5 Schematic diagram of the hardware structure of the electronic device according to the embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0016] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0017] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0018] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0019] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0020] The existing methods for producing product images are mainly divided into the following categories:

[0021] Manual production, such as designers or operations staff manually creating product images that meet target needs based on the products and product sales scenarios. However, this method has the following drawbacks:

[0022] (1) It takes longer. Usually, it takes at least 20-30 minutes to create a product image, mainly due to the search and selection of production materials, layout, color adjustment, etc.

[0023] (2) It is difficult to generate in batches. Since it takes a long time and relies on the production of image materials, manually produced product images are often only suitable for generating scenarios with a small number of images and cannot support batch generation.

[0024] Machine production relies on existing mature intelligent layout or image generation model algorithms. The model can select materials suitable for the current product based on the material library, perform intelligent image layout, and mass-produce product images for manual selection or direct external delivery. However, this method has the following drawbacks:

[0025] (1) Relying on the existing material library, it is necessary to make intelligent selections from the material pool, such as selecting background elements, button elements, decorative elements, etc. corresponding to the product image, combining and superimposing them to generate the target product image;

[0026] (2) Relying on manually set image templates, such as image templates set by designers, to output the result images according to the corresponding layout, resulting in the richness of the generated styles being directly related to the number of manual templates.

[0027] Based on this, this technical solution can generate background images and related element images based on the characteristics of the target object itself, solve the dependence on the original material library and artificial image templates, and can be adapted to the generation scenarios of display images of various categories, facilitating the batch generation of candidate background images of various styles.

[0028] According to an embodiment of the present invention, an embodiment of a method for generating a display image of a target object is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0029] In this embodiment, a method for generating a display image of a target object is provided, which can be used in electronic devices such as computers, mobile phones, and laptop computers. Figure 1 FIG. 1 is a flow chart of a method for generating a display image of a target object according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0030] Step S101 : obtaining an original image of a target object, and extracting a main body area image of the target object from the original image.

[0031] Target objects represent various categories and models of products, such as refrigerators, air conditioners, kettles, and snacks. Original images are pre-captured or collected images of the target objects in various spatial environments. The main area image is the main body image of the target object.

[0032] The original image contains the target object body and its current spatial usage environment. For the original image, the position of the target object body can be identified, the target object body outline can be identified, and the subject area image can be extracted from the original image according to the subject outline.

[0033] Step S102: Acquire description information of the target object, and generate a candidate background image suitable for the target object based on the description information.

[0034] The descriptive information is textual information describing the characteristics of the target object. The candidate background image is an image of the spatial environment in which the target object resides. Specifically, because the descriptive information can reflect the characteristics of the target object, combining this descriptive information with a pre-trained background generation model can generate the corresponding candidate background image.

[0035] Taking a refrigerator as an example, its description information is "warm home environment". Then, combining this description information can generate multiple candidate background images suitable for the refrigerator, which can specifically include warm home scenes such as restaurant background images, kitchen background images, and living room background images.

[0036] Step S103: Fusing the subject area image with the candidate background image to generate a display image of the target object.

[0037] The subject area image is combined with the candidate background image for feature matching, and the subject area image and the candidate background image are effectively fused according to the feature matching results to highlight the subject area image in the candidate background image, so that the fusion between the candidate background image and the subject area image is more harmonious, and a display image for the target object is obtained.

[0038] The method for generating a display image of a target object provided in this embodiment extracts the main area image of the target object from the original image and generates a corresponding candidate background image based on its description information. This can automatically generate a candidate background image for the target object based on the characteristics of the target object itself, making it convenient to batch generate candidate background images of various styles. The main area image can then be merged with the candidate background image to obtain a batch of display images, thereby no longer relying on existing material libraries or picture templates and being adaptable to generation scenarios of display images of various categories.

[0039] In this embodiment, a method for generating a display image of a target object is provided, which can be used in electronic devices such as computers, mobile phones, and laptop computers. Figure 2 FIG. 1 is a flow chart of a method for generating a display image of a target object according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0040] Step S201: Obtain an original image of the target object and extract a main body area image of the target object from the original image. Detailed descriptions refer to the descriptions of the corresponding steps in the above embodiment, which will not be repeated here.

[0041] Step S202 : obtaining description information of the target object, and generating a candidate background image suitable for the target object based on the description information.

[0042] For obtaining the description information, the object identification of the target object may be used to obtain the information, or the attribute features of the target object in the original image may be identified and determined.

[0043] In some specific implementations, the above step S202 may include:

[0044] Step S2021: Obtain the object identifier of the target object, and query object information that matches the object identifier.

[0045] The object information includes at least one of image information, text information, audio information, and video information of the target object.

[0046] The object identifier uniquely represents the target object and can be represented by a code, such as the target object's corresponding barcode or product serial number. Object information describes the target object and is associated with the object identifier. Specifically, this object information includes one or more of image information, text information, audio information, and video information, such as the original image of the target object, the target object's details page, the target object's title and text description, and the target object's category information.

[0047] Step S2022: parse one or more description contents related to the target object from the object information, and generate description information of the target object based on the one or more description contents.

[0048] The description content is a description text of the target object. After determining the object information, the electronic device can identify the description content contained in the object information and merge the identified one or more description contents into the description information.

[0049] For example, for a detail page image, OCR can be used to identify the text within the detail page and use the identified text as the description. For the original image, a picture-speaking model can be used to extract the content contained in the original image and convert it into a textual description. For video information, the content contained in each frame can be extracted to obtain the description of the target object. For audio information, speech can be converted into text to obtain the corresponding description. Finally, all the descriptions are combined to generate the description of the target object.

[0050] In step S2023, based on the description information, a candidate background image suitable for the target object is generated. The description information is combined with a pre-trained background generation model to generate the corresponding candidate background image. The training method for the background generation model is described in further detail below and is not detailed here.

[0051] In some specific implementations, the above step S202 may further include:

[0052] Step S2024 : identifying attribute features of the target object in the original image of the target object, and generating description information of the target object based on the identified attribute features.

[0053] Attribute features represent characteristics of the target object, such as its category, color, and shape. When no corresponding object identifier exists for the target object, the main body of the target object in the original image can be identified to determine its attribute features. Subsequently, the content understanding model and the target object's attribute features are combined to generate a description of the target object.

[0054] In step S2025, based on the description information, a candidate background image suitable for the target object is generated. The description information is combined with a pre-trained background generation model to generate the corresponding candidate background image. The training method for the background generation model is described in further detail below and is not detailed here.

[0055] In some optional implementations, the description information of the target object is input into a trained background generation model to output a candidate background image of the target object. The training method of the background generation model is as follows:

[0056] Step a1: Acquire an image sample set of a target object, where the image sample set includes multiple image samples of the target object.

[0057] Step a2: for any target image sample in the image sample set, extract the background area image excluding the target object from the target image sample.

[0058] Step a3: generating a description sample for describing the target object in the target image sample, and processing the description sample using the background generation model to output a predicted background image corresponding to the description sample.

[0059] Step a4: comparing the predicted background image and the background area image, and generating error information based on the comparison result, so as to correct the background generation model through the error information.

[0060] An image sample set is a pre-collected set of original images of a target object. This set includes multiple original images from different usage scenarios, and these original images serve as image samples. Specifically, the image sample set can be captured using a camera or collected online. The method for collecting the image sample set is not specifically limited herein.

[0061] The image sample includes the target object itself and the spatial environment it is located in, and the spatial environment exists as the background of the target object. For any target image sample in the image sample set, image segmentation is performed to extract the target object and the background area image.

[0062] The target object in the target image sample is described to generate a description sample corresponding to the target object. The description sample is then input into the background generation model for prediction processing, and the predicted background image corresponding to the description sample is output.

[0063] The predicted background image is compared with the background region image to generate a comparison result. The comparison result determines the error information between the predicted background image and the background region image. The background generation model is iteratively trained based on this error information to achieve corrections to the background generation model, making the output of the background generation model more accurate.

[0064] Here, the background image generation model is trained by combining the background area image in the target image sample and the description text of the target object, thereby achieving batch generation of candidate background images of the target object based on the description information.

[0065] Step S203: fuse the subject area image with the candidate background image to generate a display image of the target object. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0066] The method for generating a display image of a target object provided in this embodiment uses the target object's object identifier to determine its associated object information. Parsing the object information yields descriptive information, enabling the subsequent batch generation of candidate background images based solely on the object identifier. Describing the descriptive information based on the target object's inherent attributes ensures that the descriptive information matches the target object, ensuring the accuracy of candidate background image generation.

[0067] In this embodiment, a method for generating a display image of a target object is provided, which can be used in electronic devices such as computers, mobile phones, and laptop computers. Figure 3 FIG. 1 is a flow chart of a method for generating a display image of a target object according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0068] Step S301: Obtain an original image of the target object and extract a main body area image of the target object from the original image. Detailed descriptions refer to the descriptions of the corresponding steps in the above embodiment, which will not be repeated here.

[0069] Step S302: Obtain description information of the target object, and generate a candidate background image suitable for the target object based on the description information. Detailed descriptions refer to the descriptions of the corresponding steps in the above embodiment, which will not be repeated here.

[0070] Step S303: Fusing the subject area image with the candidate background image to generate a display image of the target object.

[0071] Specifically, the above step S303 may include:

[0072] Step S3031: Identify one or more candidate regions for matching the target object in the candidate background image.

[0073] The candidate area is a spatial area suitable for the target object. Specifically, the candidate area can be a spatial area for placing the target object, such as a dining room, living room, kitchen, floor, table, wall, stove, etc. The candidate area matches the target object, such as a kettle and a table, a refrigerator and the floor, etc. In addition, the candidate area can also be a scene that matches the target object. For example, if the target object is a refrigerator, then the candidate area can represent the home environment. For another example, if the target object is a car fragrance, then the candidate area can represent the car environment. After acquiring the candidate background image, the environmental features contained in the candidate background image are identified to determine one or more candidate areas that can match the target object.

[0074] For example, when the target object is a refrigerator, a spatial area for matching the refrigerator can be identified from the candidate background image generated for the refrigerator. The spatial area can be the restaurant floor, the kitchen floor, or of course other areas suitable for placing the refrigerator. There is no specific limitation here.

[0075] Step S3032: Identify the attribute features of the target object in the main area image, and determine the actual matching area in the candidate area based on the attribute features.

[0076] The actual matching region is the area that matches the target object. Attribute features characterize the target object's type, color, shape, and other characteristics. Since the main region image represents the area where the main body of the target object resides in the original image, the target object's attribute features can be identified from the main region image. Then, combining the target object's attribute features, the actual matching region of the target object is selected from multiple candidate regions.

[0077] For example, by identifying the target object in the main area image, it is determined that the target object is a tea kettle, and the candidate areas identified from the candidate background image are the stove, table top, and floor. At this time, the shape and category of the kettle can be combined to determine from multiple candidate areas that the actual matching area that matches the kettle is the table top.

[0078] Step S3033: fuse the subject area image to the actual matching area in the candidate background image.

[0079] The subject area image matches the actual matching area, meaning the target object in the subject area image fits within the actual matching area. At this point, the subject area image and the actual matching area are fused to effectively fuse the subject area image with the candidate background image, ensuring a more harmonious integration of the target object and its surroundings.

[0080] In some optional implementations, before step S3032 , the method may further include: scaling the main body area image according to size information of the target reference object in the candidate background image to generate a scaled main body area image.

[0081] Accordingly, the above step S3033 may include: fusing the scaled subject area image to the actual matching area in the candidate background image.

[0082] To further ensure effective integration of the target object with its surroundings, the size ratio of the subject area image and the target reference object in the candidate background image must match. Here, the subject area image can be scaled based on the size information of the target reference object in the candidate background image to obtain a scaled subject area image. This ensures a good size match between the scaled subject area image and the target reference object in the candidate background image.

[0083] In the above embodiment, the subject area image is adaptively scaled in combination with the size of the target reference object to optimize the fusion of the subject area image and the candidate background image, ensuring that the fused display image is more accurate and beautiful.

[0084] As an optional implementation, after generating the display object, necessary descriptive text can be supplemented for the target object, and the supplemented descriptive text can be adaptively adjusted in combination with the display image and the size of the target object in the display image to integrate it into the display image and highlight the characteristics of the target object in the display image.

[0085] Step S304: determining the segmentation degree of the target object in the main body area image, and determining display effect information of the display image.

[0086] The segmentation degree indicates the completeness and accuracy of the extraction of the target object in the main area image. Specifically, the steps of determining the segmentation degree of the target object in the main area image include:

[0087] Step b1: Detect the edge contour of the target object in the main area image.

[0088] Step b2: Compare the edge contour with the standard contour of the target object to generate a contour comparison result.

[0089] Step b3: identifying redundant image information other than the target object in the main body area image, and determining the segmentation degree of the target object in the main body area image based on the redundant image information and the contour comparison result.

[0090] The standard contour is the actual contour of the target object. The main region image is an image containing the main body of the target object. The edges of the target object can be depicted in the main region image to determine the edge contour of the target object. At the same time, redundant image information other than the target object is removed from the main region image.

[0091] The standard outline of the target object is compared with the edge outline of the target object extracted from the main area image to determine whether the two are consistent and obtain the corresponding outline comparison result. The redundant image information and the outline comparison result are combined to determine the segmentation degree of the target object.

[0092] In the above embodiment, the target object's edge contour is compared with the standard contour, and the target object's segmentation degree is determined according to the contour comparison result and redundant image information to determine the target object's integrity and avoid introducing redundant information during segmentation.

[0093] The display effect information represents the effect of the target object in the display image, and is used to characterize the generated effect of the display image. Specifically, the display effect information of the display image is realized by completing the training of the quality assessment model. The training steps of the quality assessment model may include:

[0094] Step c1: obtaining a display image sample set, wherein the display samples in the display image sample set have marked standard conversion information, wherein the standard conversion information is used to generate display effect information of the display samples.

[0095] Step c2: for any target display sample in the display image sample set, process the target display sample using the quality assessment model to generate predicted conversion information corresponding to the target display sample.

[0096] Step c3: comparing the predicted conversion information with the standard conversion information, and generating error information based on the comparison result, so as to calibrate the quality assessment model through the error information.

[0097] Standard conversion information is used to generate display effect information for evaluating the display sample generation effect. Specifically, the standard conversion information may include the number of times the display sample has been viewed, the number of comments and positive reviews on the display sample, and the purchase volume of the target object corresponding to the display sample.

[0098] The display image sample set is a display sample set generated by fusing the target object and the candidate background image. The display samples in the display image sample set all have corresponding standard conversion information.

[0099] For any target display sample in the display image sample set, it is input into the quality assessment model for evaluation processing of conversion information, and predicted conversion information corresponding to the target display sample is output.

[0100] The predicted conversion information is compared with the standard conversion information to determine the comparison result. The error between the predicted conversion information and the standard conversion information is determined through the comparison result. The quality assessment model is iteratively trained based on this error information to achieve correction of the quality assessment model, making the conversion information output by the quality assessment model more accurate.

[0101] In the above implementation, the quality assessment model is trained to evaluate the display effect information of the display image, ensuring that the display image finally determined conforms to the subject characteristics of the target object and has a high degree of matching with the target object.

[0102] Step S305 : generating quality information of the display image according to the segmentation degree and the display effect information, so as to determine whether to retain the display image according to the quality information.

[0103] The quality information indicates the degree of fusion and matching between the target object and the candidate background image, as well as the integrity of the target object. The quality information of the display image can be determined by combining the segmentation degree and the display quality information. If the fusion and matching between the target object and the candidate background image are poor, and / or the integrity of the target object is poor, the display image can be determined to have poor quality and can be deleted. Otherwise, the display image is considered to have good quality and can be retained.

[0104] In some optional embodiments, if there are multiple display images generated for the target object, the above method may further include: sorting the multiple display images according to display effect information of each display image, and selecting the actual display image of the target object based on the sorting result.

[0105] Combined with the display effect information of each display image, each display image is scored to obtain a score for each display image. Based on the scores, the display images are sorted from high to low or from low to high to obtain a corresponding sorting result. Based on this sorting result, the display images with higher scores can be identified and designated as the actual display images.

[0106] Here, the display images are sorted in combination with the display effect information so as to select actual display images with a higher matching degree, so that the actual display images can accurately represent the target object features and usage scenarios.

[0107] The method for generating a display image of a target object provided in this embodiment combines candidate regions and attribute characteristics of the target object to determine the actual matching region for the target object. This method intelligently adjusts the fusion layout of the subject region image and candidate background images based on different backgrounds and the attribute characteristics of the target object, ensuring that the fused display image accurately reflects the applicable scenario of the target object. The quality of the generated display image is determined based on the segmentation degree of the target object in the subject region image and the display effect information of the display image, facilitating the selection of high-quality display images.

[0108] This embodiment also provides a device for generating a display image of a target object. This device is used to implement the above-mentioned embodiments and preferred embodiments, and the details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0109] This embodiment provides a device for generating a display image of a target object, such as Figure 4 As shown, the device includes:

[0110] The image acquisition module 401 is configured to acquire an original image of a target object and extract a main body area image of the target object from the original image.

[0111] The description information acquisition module 402 is configured to acquire description information of the target object and generate a candidate background image suitable for the target object based on the description information.

[0112] The fusion module 403 is used to fuse the subject area image with the candidate background image to generate a display image of the target object.

[0113] In some optional implementations, the description information acquisition module 402 may include:

[0114] The object information acquisition unit is used to acquire the object identification of the target object and query the object information matching the object identification, wherein the object information includes at least one of the image information, text information, audio information, and video information of the target object.

[0115] The parsing unit is configured to parse the object information to obtain one or more description contents related to the target object, and generate description information of the target object based on the one or more description contents.

[0116] In some optional implementations, the description information acquisition module 402 may further include:

[0117] The feature recognition unit is used to recognize the attribute features of the target object in the original image of the target object, and generate description information of the target object based on the recognized attribute features.

[0118] In some optional implementations, the description information acquisition module 402 may further include:

[0119] The background generation model training unit is used to train the background generation model.

[0120] Specifically, the background generation model training unit is used to: obtain an image sample set of the target object, the image sample set including multiple image samples of the target object; for any target image sample in the image sample set, extract a background area image other than the target object from the target image sample; generate a description sample for describing the target object in the target image sample, and use the background generation model to process the description sample to output a predicted background image corresponding to the description sample; compare the predicted background image and the background area image, and generate error information based on the comparison result to correct the background generation model through the error information.

[0121] In some optional implementations, the fusion module 403 may include:

[0122] The candidate region identification unit is used to identify one or more candidate regions for matching the target object in the candidate background image.

[0123] The attribute feature recognition unit is used to recognize the attribute features of the target object in the subject area image and determine the actual matching area in the candidate area based on the attribute features.

[0124] The fusion unit is used to fuse the subject area image to the actual matching area in the candidate background image.

[0125] In some optional implementations, the fusion module 403 may further include:

[0126] The scaling unit is configured to scale the subject area image according to size information of the target reference object in the candidate background image to generate a scaled subject area image.

[0127] Correspondingly, the fusion unit is further configured to fuse the scaled subject area image to the actual matching area in the candidate background image.

[0128] In some optional implementations, the device for generating a display image of the target object may further include:

[0129] The information determination unit is used to determine the segmentation degree of the target object in the main area image and determine the display effect information of the display image.

[0130] The quality determination unit is used to generate quality information of the display image according to the segmentation degree and the display effect information, so as to determine whether to retain the display image based on the quality information.

[0131] In some optional implementations, the information determining unit includes:

[0132] The segmentation degree determination subunit is used to detect the edge contour of the target object in the main area image; compare the edge contour with the standard contour of the target object to generate a contour comparison result; identify redundant image information other than the target object in the main area image, and determine the segmentation degree of the target object in the main area image based on the redundant image information and the contour comparison result.

[0133] The quality assessment model training subunit is used to train the quality assessment model to determine quality information.

[0134] Specifically, the quality assessment model training subunit is used to: obtain a display image sample set, the display samples in the display image sample set have labeled standard conversion information, wherein the standard conversion information is used to generate display effect information of the display samples; for any target display sample in the display image sample set, use the quality assessment model to process the target display sample to generate predicted conversion information corresponding to the target display sample; compare the predicted conversion information and the standard conversion information, and generate error information based on the comparison result, so as to correct the quality assessment model through the error information.

[0135] In some optional implementations, the device for generating a display image of the target object may further include:

[0136] The sorting module is used to sort the multiple display images according to the display effect information of each display image, and select the actual display image of the target object based on the sorting result.

[0137] The specific processing logic of each module and each unit can be found in the description of the aforementioned method implementation method, which will not be repeated here.

[0138] The various units described in the above embodiments can be implemented by computer chips or products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0139] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0140] The device for generating a display image of a target object provided by the embodiment of the present disclosure extracts the main area image of the target object from the original image and generates a corresponding candidate background image based on its description information. This can automatically generate a candidate background image for the target object based on the characteristics of the target object itself, facilitates batch generation of candidate background images of various styles, and then fuses the main area image with the candidate background image to obtain a batch of display images. This no longer relies on existing material libraries or picture templates and can be adapted to generation scenarios of display images of various categories.

[0141] An embodiment of the present invention further provides an electronic device having the above Figure 4 A device for generating a display image of the target object shown.

[0142] See also Figure 5 , Figure 5 is a structural diagram of an electronic device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0143] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0144] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0145] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0146] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0147] The electronic device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 5 The bus connection is taken as an example.

[0148] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0149] The electronic device also includes a communication interface for data communication between the electronic device and other devices or a communication network.

[0150] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0151] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant portions, refer to the descriptions of the method embodiments.

[0152] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

[0153] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for generating a display image of a target object, characterized in that: The method comprises: Acquire an original image of a target object, and extract a main body area image of the target object from the original image; Acquiring description information of the target object, and generating a candidate background image suitable for the target object based on the description information; fusing the subject area image with the candidate background image to generate a display image of the target object; Determining a segmentation degree of the target object in the main body area image and determining display effect information of the display image; the segmentation degree represents the completeness and accuracy of extraction of the target object in the main body area image; generating quality information of the display image according to the segmentation degree and the display effect information, so as to determine whether to retain the display image according to the quality information; Wherein, determining the segmentation degree of the target object in the main body area image includes: detecting the edge contour of the target object in the main body area image; comparing the edge contour with the standard contour of the target object to generate a contour comparison result; identifying redundant image information other than the target object in the main body area image, and determining the segmentation degree of the target object in the main body area image based on the redundant image information and the contour comparison result.

2. The method according to claim 1, characterized in that The acquiring the description information of the target object includes: Obtaining an object identifier of the target object, and searching for object information matching the object identifier, wherein the object information includes at least one of image information, text information, audio information, and video information of the target object; One or more description contents related to the target object are parsed from the object information, and description information of the target object is generated based on the one or more description contents.

3. The method according to claim 1, characterized in that The acquiring the description information of the target object includes: Attribute features of the target object are identified in the original image of the target object, and description information of the target object is generated based on the identified attribute features.

4. The method according to claim 1, wherein The step of generating a candidate background image suitable for the target object based on the description information is achieved by completing a trained background generation model; wherein the background generation model is trained in the following manner: Acquire an image sample set of the target object, where the image sample set includes multiple image samples of the target object; For any target image sample in the image sample set, extracting a background area image excluding the target object from the target image sample; generating a description sample for describing the target object in the target image sample, and processing the description sample using the background generation model to output a predicted background image corresponding to the description sample; The predicted background image and the background region image are compared, and error information is generated based on the comparison result, so as to correct the background generation model by using the error information.

5. The method according to claim 1, wherein The fusing of the subject area image with the candidate background image includes: In the candidate background image, identifying one or more candidate regions for matching the target object; Identifying attribute features of the target object in the subject area image, and determining an actual matching area in the candidate area based on the attribute features; The subject area image is fused to the actual matching area in the candidate background image.

6. The method according to claim 5, characterized in that Before fusing the subject area image to the actual matching area in the candidate background image, the method further includes: The main body region image is scaled according to the size information of the target reference object in the candidate background image to generate a scaled main body region image.

7. The method according to claim 1, characterized in that The step of determining the display effect information of the display image is achieved by completing the trained quality assessment model, and the quality assessment model is trained in the following manner: Acquire a display image sample set, wherein the display samples in the display image sample set have marked standard conversion information, wherein the standard conversion information is used to generate display effect information of the display samples; For any target display sample in the display image sample set, use the quality assessment model to process the target display sample to generate predicted conversion information corresponding to the target display sample; The predicted conversion information and the standard conversion information are compared, and error information is generated based on the comparison result, so as to calibrate the quality assessment model according to the error information.

8. The method according to claim 1, characterized in that If a plurality of display images of the target object are generated, the method further includes: The plurality of display images are sorted according to the display effect information of each display image, and the actual display image of the target object is selected based on the sorting result.

9. A device for generating a display image of a target object, characterized in that: The device comprises: An image acquisition module is used to acquire an original image of a target object and extract a main body area image of the target object from the original image; a description information acquisition module, configured to acquire description information of the target object and generate a candidate background image suitable for the target object based on the description information; a fusion module, configured to fuse the subject area image with the candidate background image to generate a display image of the target object; an information determining unit, configured to determine a segmentation degree of the target object in the main body region image and determine display effect information of the display image; the segmentation degree represents the completeness and accuracy of extraction of the target object in the main body region image; a quality determination unit, configured to generate quality information of the display image according to the segmentation degree and the display effect information, so as to determine whether to retain the display image based on the quality information; Among them, the information determination unit includes: a segmentation degree determination subunit, which is used to detect the edge contour of the target object in the main area image; compare the edge contour with the standard contour of the target object to generate a contour comparison result; identify redundant image information other than the target object in the main area image, and determine the segmentation degree of the target object in the main area image based on the redundant image information and the contour comparison result.

10. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory is used to store a computer program, and when the computer program is executed by the processor, the method for generating a display image of a target object as described in any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method for generating a display image of a target object according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Image synthesis method and device, equipment and storage medium

    CN114820292A

  • Background replacement method and device, computer equipment and storage medium

    CN116156092A