Image generation method and device adaptive to scene, equipment and medium
By acquiring object information from the target image and using the Transformer model to generate new objects that adapt to the scene, the problem of unreasonable semantic logic in existing technologies is solved, achieving high-quality data augmentation effects and improving the training effect and detection accuracy of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SHENZHOU EVERBRIGHT TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-08
AI Technical Summary
Existing data augmentation methods are difficult to meet the training requirements of models for complex scene samples, and have problems with unreasonable semantic logic, such as randomly embedding unreasonable new objects such as giraffes in office scenes.
By acquiring the object information of multiple first objects in the target image, the Transformer model is used to determine the object information of second objects that are adapted to the scene where the multiple first objects are located, a layout image is generated, and a second object matching the target image is generated in the layout image and added to the target image, using the visual features of the target image as a reference condition.
The generated images possess semantic logic that conforms to common sense, which can meet the training requirements of the model for complex scene samples, improve the realism and visual consistency of the generated images, avoid unreasonable semantic logic, and improve the detection accuracy and generalization ability of the object detection model.
Smart Images

Figure CN121999082A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image technology, and in particular to an image generation method, apparatus, device and medium adapted to a scene. Background Technology
[0002] Currently, mainstream data augmentation methods can be broadly categorized into two types. The first type involves simple image-level transformations, such as rotation, cropping, color dithering, or splicing. These methods offer only limited enhancement and are insufficient for training models on complex scene samples. The second type is instance embedding based on copy-paste. This approach attempts to enrich sample diversity by embedding new object instances into the original image, but suffers from serious semantic flaws: due to a lack of understanding of the scene's high-level semantics, the categories of newly added objects often clash with the scene context—for example, randomly embedding a giraffe into an office scene. Summary of the Invention
[0003] This application provides an image generation method that adapts to a specific scene. This method not only meets the training requirements of models for complex scene samples, but also effectively avoids the situation where data augmentation techniques in related technologies have unreasonable semantic logic, so that the generated images have semantic logic that conforms to common sense.
[0004] In a first aspect, embodiments of this application provide an image generation method adapted to a scene, the method comprising: The process involves: acquiring object information of multiple first objects in the target image; determining object information of second objects that are compatible with the scene containing the first objects; generating a layout image with the same size as the target image, containing the outlines and positional ranges of the second objects; using the visual features of the target image as a reference and the positional range of the second objects in the layout image as a spatial constraint, generating second objects within the positional range of the second objects contained in the layout image that match the visual features of the target image; and adding the second objects generated in the layout image that match the visual features of the target image to the target image to obtain a target image containing the second objects.
[0005] In one embodiment, the above-mentioned acquisition of object information of multiple first objects in a target image may specifically include: identifying the target image to obtain object information of each object in the target image; determining the confidence level of each object in the target image, wherein the confidence level of each object is used to characterize the reliability of the object information; and taking the object information of the N objects with the highest confidence levels as the object information of the first object, where N is an integer greater than 1.
[0006] In one embodiment, object information includes at least one of the following: object category, location information, or object size.
[0007] In one embodiment, the object information includes location information. Based on the object information of multiple first objects, the object information of a second object that is adapted to the scene where the multiple first objects are located is determined. Specifically, this may include: based on the location information of the multiple first objects, determining the spatial layout information of the multiple first objects, and based on the spatial layout information of the multiple first objects, determining the location information of the second object.
[0008] In one embodiment, the object information includes object category, location information, and object size. Based on the object information of the second object, a layout image is generated, which may specifically include: A layout image is generated based on the object category, size, and position information of the second object.
[0009] In one embodiment, the layout image is a binary mask image.
[0010] Secondly, embodiments of this application provide an image generation apparatus adapted to a scene, the image generation apparatus comprising: The acquisition module is used to acquire object information of multiple first objects in the target image; The determination module is used to determine the object information of a second object that is compatible with the scene where the plurality of first objects are located, based on the object information of the plurality of first objects. A generation module is used to generate a layout image based on the object information of the second object. The layout image is the same size as the target image and includes the outline of the second object and the position range of the second object. The generation module is further configured to generate a second object that matches the visual features of the target image, using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition. The generation module is further configured to add the second object, which is generated in the layout image and matches the visual features of the target image, to the target image to obtain the target image containing the second object.
[0011] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the scene-adaptive image generation method described in the first aspect.
[0012] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the scene-adaptive image generation method described in the first aspect.
[0013] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the scene-adaptive image generation method described in the first aspect.
[0014] The scene-adaptive image generation method provided in this application determines a second object that is adapted to the scene where multiple first objects are located. By utilizing the correlation between objects, it ensures that the newly added object is a reasonable existence in the scene. This not only meets the training requirements of the model for complex scene samples, but also effectively avoids the situation where the semantic logic of data augmentation techniques in related technologies is unreasonable, so that the generated image has a semantic logic that conforms to common sense. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating an image generation method adapted to a specific scene, as provided in an embodiment of this application.
[0017] Figure 2 This is a schematic diagram of an image generation method adapted to a specific scene, provided in an embodiment of this application.
[0018] Figure 3 This is a schematic diagram of the structure of the scene-adaptive image generation device provided in the embodiments of this application.
[0019] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] In the description of this invention, it should be understood that the terms "center," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," and "third" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0022] In the field of object detection, the generalization ability and robustness of a model directly determine its performance in real-world, complex environments, and sufficient and diverse training samples are the core foundation for improving these performance characteristics. Data augmentation technology, as a key support method in the training process of object detection models, aims to effectively alleviate the problem of model overfitting by reasonably expanding the diversity of training samples. This helps the model better adapt to the complex changes in object shape, position, and environment in real-world scenes, thereby improving the model's detection accuracy and environmental adaptability.
[0023] Currently, mainstream data augmentation methods can be divided into two main categories. The first category consists of simple image-level transformation techniques, such as rotating, cropping, color dithering, or stitching images. These methods are easy to operate and can increase the surface diversity of samples to some extent, but they do not fundamentally change the semantic information within the image or the inherent relationships between objects. They can only achieve limited enhancement effects and are difficult to meet the training requirements of models for complex scene samples.
[0024] The second category is instance embedding methods based on copy and paste. These methods attempt to enrich sample diversity by embedding new object instances into the original image, but the semantic logic is seriously unreasonable: due to the lack of understanding of the high-level semantics of the scene, the category of the newly added object is often seriously inconsistent with the scene context, such as randomly embedding "giraffe" in an office scene.
[0025] In conclusion, there is an urgent need for a data augmentation technique that can meet the training requirements of models for complex scene samples and accurately match scene semantics.
[0026] To address the aforementioned technical issues, this application provides an image generation method adapted to a specific scene. This method acquires object information of multiple first objects in a target image and, based on this object information, determines object information of second objects adapted to the scene containing the first objects. Then, based on the object information of the second objects, a layout image is generated. This layout image has the same size as the target image and includes the outline of the second objects and their positional range. Next, using the visual features of the target image as a reference and the positional range of the second objects in the layout image as a spatial constraint, a second object matching the visual features of the target image is generated within the positional range of the second objects contained in the layout image. The second object generated in the layout image and matching the visual features of the target image is then added to the target image to obtain a target image containing the second objects.
[0027] The scene-adaptive image generation method provided in this application determines a second object that is adapted to the scene where multiple first objects are located. By utilizing the correlation between objects, it ensures that the newly added object is a reasonable existence in the scene. This not only meets the training requirements of the model for complex scene samples, but also effectively avoids the situation where the semantic logic of data augmentation techniques in related technologies is unreasonable, so that the generated image has a semantic logic that conforms to common sense.
[0028] The following is combined Figures 1-4 The image generation method adapted to the scene provided in the embodiments of this application will be described in detail.
[0029] Figure 1 This is a flowchart illustrating a scene-adaptive image generation method provided in an embodiment of this application, applied to an electronic device. (Refer to...) Figure 1 As shown, the method includes the following steps S101-S105.
[0030] S101, Obtain object information of multiple first objects in the target image.
[0031] In this embodiment, object information is used to characterize the core attributes of existing objects in the target image. Its specific type is not uniquely limited. For example, it may include at least one of object category, location information and object size to meet the needs of subsequent scene adaptation judgment and spatial layout analysis.
[0032] In one example, the object's position can be represented by the coordinates of the center point of the bounding box corresponding to the object. In another example, the object's size can be quantitatively described by the width and height of the bounding box corresponding to the object.
[0033] To effectively improve the efficiency of acquiring object information of multiple first objects in a target image, a first model can be deployed in the electronic device. The first model is used to recognize the target image.
[0034] In one example, the first model could be an object detection model that incorporates contextual features for modeling.
[0035] Specifically, taking object information including object category, location information, and object size as an example, electronic devices can use a target detection model that incorporates contextual feature modeling to perform recognition operations on target images, so as to accurately extract object information of multiple objects contained in the image. The target detection model that incorporates contextual feature modeling can be improved by integrating non-local attention mechanisms or graph neural network modules into the target detection backbone network (for example, by optimizing the backbone network based on a faster region-based convolutional neural network (Faster R-CNN) or a detection transformer (DETR) model). It can improve the recognition accuracy of objects in the target image through contextual feature modeling, and then output the category labels (i.e., object categories) and bounding box coordinates corresponding to the multiple objects (i.e., the first object) identified in the target image, and the bounding box coordinates include the coordinates of the center point of the bounding box (i.e., the object's location information) and the width and height of the bounding box (i.e., the object's size).
[0036] To further ensure the accuracy and reliability of the target information extraction relied upon for subsequent processing, in an optional implementation, the electronic device acquires object information of multiple first objects in the target image. Specifically, this can be achieved by: recognizing the target image to obtain object information of each object in the target image. Then, the electronic device can determine the confidence level of each object in the target image and use the object information of the N objects with the highest confidence levels as the object information of the first object.
[0037] The confidence score for each object is used to characterize the reliability of the object's information. The confidence score for each object is positively correlated with the reliability of the object's information; the higher the confidence score, the more accurate the electronic device's recognition of the object's information, meaning the more reliable the object's information.
[0038] N is an integer greater than 1. This application does not limit the specific value of N; it can be flexibly set according to actual needs such as the scene complexity and object density of the target image. For example, N can be 3, 4, or a larger or smaller reasonable value.
[0039] Specifically, after obtaining the object information of each object in the target image using the method described above, the electronic device can determine the confidence level of each object. Subsequently, the electronic device can sort the confidence levels of each object, filter out the object information of the top N objects with the highest confidence levels, and use these as the object information of the multiple first objects finally obtained in this step, discarding the remaining object information with lower confidence levels to avoid interference.
[0040] S102, based on the object information of multiple first objects, determine the object information of a second object that is compatible with the scene where the multiple first objects are located.
[0041] Specifically, taking object information including object category, location information, and object size as an example, an electronic device can determine the object category of a second object based on the object categories of multiple first objects, determine the object size of a second object based on the object sizes of multiple first objects, determine the spatial layout information of multiple first objects based on the location information of multiple first objects, and determine the location information of a second object based on the spatial layout information of multiple first objects.
[0042] In order to effectively improve the efficiency and rationality of determining the object information of the second object and ensure that the generated second object is highly adapted to the scene semantics and spatial rules, in an optional implementation, the electronic device may be equipped with a second model with context-related reasoning ability. The second model is used to generate the object information of the second object that is adapted to the scene where the first object is located based on the object information of multiple first objects.
[0043] In one example, the second model can be a sequence-to-sequence model based on a Transformer. Its core advantage lies in its ability to deeply mine and learn spatial layout patterns through a self-attention mechanism, thereby achieving reasoning and generation from existing object information to reasonably added object information, rather than simply completing scene objects.
[0044] Specifically, continuing with the example of object information including object category, location information, and object size, the electronic device can first organize the object category, location information, and object size of multiple first objects obtained in step S101 into a unified input sequence (hereinafter referred to as input sequence 1) according to a preset format, and then input input sequence 1 into the trained Transformer model. After receiving the input sequence, the model does not fill in the missing objects in the scene, but rather performs a deep analysis of the contextual relationships of multiple first objects based on a self-attention mechanism. Through the learned spatial layout rules, it infers and generates object information of one or more second objects that conform to the semantic logic of the scene and the common sense of spatial distribution. The object information of the second object may include the object category (cls_new), location information (such as the center point coordinates x_new, y_new of the bounding box corresponding to the second object), and object size (such as the width w_new and height h_new of the bounding box corresponding to the second object).
[0045] In one example, suppose that the first objects include objects in the kitchen scene such as cabinets and stoves. The object category of the second object generated by the Transformer model above can be a kettle, and the position of the kettle can be a reasonable position on the cabinet. The size of the kettle can be adapted by referring to the size of the objects in the kitchen scene.
[0046] It is understandable that before using the Transformer model, a dedicated training set needs to be built to train the Transformer model so that it has the ability to infer and generate object information adapted to the second object based on the object information of the first object.
[0047] The training set can be constructed using targeted sample construction logic. Specifically, for a single sample image containing M objects (M is an integer greater than 1), M training samples can be constructed based on this sample image. The input sequence of each training sample is set as the complete object information of M-1 objects in the image. That is, by randomly masking any one object in the image, the category labels (cls_i) and bounding box information (bbox_i, including the object's position and size) of the remaining M-1 objects are arranged into an ordered input sequence. At the same time, the label of each training sample is set as the complete object information of the masked object, that is, the category label (cls_j) and bounding box information (BBox_j) of the masked object. Through this sample construction method, the model can continuously learn the core logic of "how to accurately predict the most likely other object and its reasonable spatial position in a given scene when the context information of some objects in the scene is used as a constraint" during the training process. In this way, it can gradually master the spatial layout rules between objects in different scenes, laying the foundation for receiving the sequence information of multiple first objects and inferring and generating the appropriate second object information.
[0048] S103, Generate a layout image based on the object information of the second object.
[0049] The layout image is the same size as the target image, and the layout image contains the outline of the second object and the position range of the second object.
[0050] Given that the object information includes object category, location information, and object size, a layout image is generated based on the object information of the second object. Specifically, this can be achieved by generating a layout image based on the object category, object size, and location information of the second object.
[0051] In one example, the layout image described above can be a binary mask image. The binary mask image can clearly mark the location range of the second object (i.e., the bounding box region corresponding to the second object) by distinguishing between black and white binary pixels, and mark the category outline of the second object within the bounding box region, providing clear visual guidance for the subsequent generation process.
[0052] S104, using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition, generate a second object that matches the visual features of the target image within the position range of the second object contained in the layout image.
[0053] S105, a second object generated in the layout image that matches the visual features of the target image is added to the target image to obtain a target image containing the second object.
[0054] The visual features of the target image may include information such as lighting conditions, texture style, and shadow distribution in the target image, but this application embodiment does not specifically limit these features.
[0055] By using the visual features of the target image as a reference and the layout image as a spatial constraint, it can be ensured that the second object and the background of the target image are highly coordinated in visual presentation.
[0056] Specifically, taking the layout image as a binary mask image as an example, the electronic device can determine the outline of the second object based on its object category and size, and determine its position range in the target image based on its position information. The electronic device can then generate a binary mask image of the same size as the target image, based on the outline of the second object and its position range in the target image, with the outline and position range of the second object marked in the mask image.
[0057] Electronic devices can use the visual features of the target image as a reference and the layout image as a spatial constraint to generate a second object that matches the visual features of the target image within the position range of the second object contained in the layout image. In this way, the second object generated in the layout image can not only be highly coordinated with the background of the target image in visual presentation, but also will not exceed the position range of the second object marked in the binary mask image.
[0058] Then, the electronic device can seamlessly fuse a second object generated in the layout image that matches the visual features of the target image with the target image to obtain a target image containing the second object.
[0059] To effectively improve the efficiency and quality of generating target images containing a second object, and to ensure that the newly added second object blends naturally with the background of the target image without any stitching marks, in an optional implementation, the electronic device may include a third model. The third model is used to generate image content that is consistent with the visual features of the target image and conforms to the category features of the second object within the position range defined by the layout image, using the target image and the layout image as dual constraints, and to achieve seamless integration of the image content with the target image.
[0060] In one example, the third model can be a conditionally controlled diffusion model. The core logic of this model is as follows: Using both the target image and the layout image as input conditions, it ensures that the generated second object visually matches the background by referencing the visual features of the target image (including lighting conditions, texture style, shadow distribution, etc.). Simultaneously, it uses the layout image as a strict spatial constraint, performing the generation operation only within the marked location range of the second object in the layout image, avoiding exceeding a reasonable area. During the generation process, this diffusion model progressively optimizes image details through an iterative denoising process, learning to generate image content of the second object within a specified bounding box that is highly consistent with the surrounding background in terms of lighting, texture, shadow, and style. Simultaneously, it automatically and correctly handles the occlusion relationship between the second object and the existing first object in the target image, ultimately outputting a highly realistic image seamlessly integrated with the second object—that is, the target image containing the second object.
[0061] The scene-adaptive image generation method provided in this application, through context-aware new object generation and precise image fusion design, has significant benefits in multiple dimensions, as follows: First, it significantly improves the realism and visual consistency of the generated images. This solution uses a conditionally controlled diffusion model to perform image generation. Compared to traditional generative adversarial networks, it can generate more refined and higher-quality image content, effectively avoiding common problems such as stitching artifacts and visual inconsistencies. At the same time, by using the original image as the core reference condition, it ensures that the newly generated objects are highly consistent with the original background in terms of visual features such as lighting, texture, shadow, and style, achieving seamless integration of the newly generated objects with the original image. The generated images have extremely high realism and avoid the negative interference of false samples on model training.
[0062] Secondly, it ensures semantic rationality and scene adaptability. The solution relies on the Transformer model to deeply model the contextual relationships between existing objects, enabling it to accurately learn the spatial layout rules of objects in different scenes. This ensures that the category and position of newly added objects perfectly match the scene logic—for example, automatically identifying the scene where a "television" is located and generating a "remote control" in a reasonable position. This fundamentally eliminates the generation of semantically absurd samples such as "a car floating in the air" or "a giraffe in the office," providing meaningful and common-sense high-quality training data for the object detection model.
[0063] Furthermore, the solution achieves full automation and intelligence in the data augmentation process, significantly improving efficiency and versatility. The entire solution employs an end-to-end integrated processing architecture, from extracting object information from the original image and intelligently inferring new objects to generating the final image, all without manual intervention. There is no need to manually design complex object pasting rules or pre-collect massive instance segmentation libraries. The system can autonomously complete scene semantic analysis, reasonable inference of new objects, and generation, greatly reducing labor costs. At the same time, it breaks through the limitations of traditional methods, significantly improving the efficiency and scalability of data augmentation.
[0064] Finally, this approach effectively enhances the training effect and performance of downstream object detection models. The high-quality and diverse training samples generated by this method can significantly enrich the training data distribution of the model, effectively alleviating the overfitting problem of object detection models. Especially in small dataset scenarios, these high-quality samples can help the model learn object features, scene semantics, and relationships between objects more comprehensively, thereby significantly improving the model's detection accuracy, recall, and generalization ability, making the model's performance more stable and reliable in real-world complex environments.
[0065] The following are Figure 2 The image A shown in the figure is an example of a target detection model that incorporates contextual features, a sequence-to-sequence model based on Transformer, and a diffusion model based on conditional control. In conjunction with the above embodiments, the image generation method adapted to the scene provided by the embodiments of this application will be further introduced.
[0066] Reference Figure 2 As shown, after the electronic device inputs image A into the object detection model, the object detection model can output the object information of each object contained in image A. Assuming image A contains P objects, the electronic device can determine the object information of the top N objects with the highest confidence. The electronic device can then input the object information of the top N objects into a Transformer-based sequence-to-sequence model, which can then output the object information of the second object.
[0067] Taking the object category of the second object as an airplane, and the location information of the second object indicating that the second object is located in the sky in image A as an example, the electronic device can generate a binary mask image of the same size as the target image based on the object information of the second object, such as... Figure 2 Image B is shown, and the binary mask image is marked with the outline of the second object and the location range of the second object.
[0068] An electronic device can input images A and B into a conditionally controlled diffusion model, which can then output a target image containing a second object, such as... Figure 2 Image C is shown, and the second object in the target image is visually adapted to the background height of the target image.
[0069] The following describes the scene-adaptive image generation apparatus provided in the embodiments of this application. The scene-adaptive image generation apparatus described below can be referred to in correspondence with the scene-adaptive image generation method described above.
[0070] Figure 3 This example illustrates a schematic diagram of an image generation device adapted to a specific scene. The device is located within an electronic device and may include: The acquisition module 301 is used to acquire object information of multiple first objects in the target image; The determining module 302 is used to determine the object information of a second object that is compatible with the scene where the plurality of first objects are located, based on the object information of the plurality of first objects. The generation module 303 is used to generate a layout image based on the object information of the second object. The size of the layout image is the same as the size of the target image. The layout image includes the outline of the second object and the position range of the second object. The generation module 303 is further configured to generate a second object that matches the visual features of the target image, using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition. The generation module 303 is further configured to add the second object, which is generated in the layout image and matches the visual features of the target image, to the target image to obtain the target image containing the second object.
[0071] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call a computer program in the memory 430 to execute steps of a scene-adaptive image generation method, such as: Obtain object information of multiple first objects in the target image; Based on the object information of the plurality of first objects, determine the object information of a second object that is adapted to the scene in which the plurality of first objects are located; Based on the object information of the second object, a layout image is generated. The size of the layout image is the same as the size of the target image. The layout image includes the outline of the second object and the position range of the second object. Using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition, a second object matching the visual features of the target image is generated within the position range of the second object contained in the layout image. The second object, which is generated in the layout image and matches the visual features of the target image, is added to the target image to obtain the target image containing the second object.
[0072] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0073] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the scene-adaptive image generation method provided in the above embodiments, such as including: Obtain object information of multiple first objects in the target image; Based on the object information of the plurality of first objects, determine the object information of a second object that is adapted to the scene in which the plurality of first objects are located; Based on the object information of the second object, a layout image is generated. The size of the layout image is the same as the size of the target image. The layout image includes the outline of the second object and the position range of the second object. Using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition, a second object matching the visual features of the target image is generated within the position range of the second object contained in the layout image. The second object, which is generated in the layout image and matches the visual features of the target image, is added to the target image to obtain the target image containing the second object.
[0074] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to execute the steps of the scene-adaptive image generation method provided in the above embodiments, such as including: Obtain object information of multiple first objects in the target image; Based on the object information of the plurality of first objects, determine the object information of a second object that is adapted to the scene in which the plurality of first objects are located; Based on the object information of the second object, a layout image is generated. The size of the layout image is the same as the size of the target image. The layout image includes the outline of the second object and the position range of the second object. Using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition, a second object matching the visual features of the target image is generated within the position range of the second object contained in the layout image. The second object, which is generated in the layout image and matches the visual features of the target image, is added to the target image to obtain the target image containing the second object.
[0075] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0076] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating images that adapts to a scene, characterized in that, The method includes: Obtain object information of multiple first objects in the target image; Based on the object information of the plurality of first objects, determine the object information of a second object that is adapted to the scene in which the plurality of first objects are located; Based on the object information of the second object, a layout image is generated. The size of the layout image is the same as the size of the target image. The layout image includes the outline of the second object and the position range of the second object. Using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition, a second object matching the visual features of the target image is generated within the position range of the second object contained in the layout image. The second object, which is generated in the layout image and matches the visual features of the target image, is added to the target image to obtain the target image containing the second object.
2. The image generation method according to claim 1, characterized in that, The step of obtaining object information of multiple first objects in the target image includes: The target image is identified to obtain object information for each object in the target image; The confidence score of each object in the target image is determined, and the confidence score of each object is used to characterize the reliability of the object information. The object information of the N objects with the highest confidence is used as the object information of the first object, where N is an integer greater than 1.
3. The image generation method according to claim 1 or 2, characterized in that, The object information includes at least one of the following: object category, location information, or object size.
4. The image generation method according to claim 3, characterized in that, The object information includes location information. The step of determining the object information of a second object adapted to the scene where the plurality of first objects are located, based on the object information of the plurality of first objects, includes: Based on the position information of the plurality of first objects, the spatial layout information of the plurality of first objects is determined; Based on the spatial layout information of the plurality of first objects, the position information of the second object is determined.
5. The image generation method according to claim 3, characterized in that, The object information includes object category, location information, and object size. Generating a layout image based on the object information of the second object includes: The layout image is generated based on the object category, object size, and position information of the second object.
6. The image generation method according to claim 1, characterized in that, The layout image is a binary mask image.
7. An image generation apparatus adapted to a scene, characterized in that, The image generation device includes: The acquisition module is used to acquire object information of multiple first objects in the target image; The determination module is used to determine the object information of a second object that is compatible with the scene where the plurality of first objects are located, based on the object information of the plurality of first objects. A generation module is used to generate a layout image based on the object information of the second object. The layout image is the same size as the target image and includes the outline of the second object and the position range of the second object. The generation module is further configured to generate a second object that matches the visual features of the target image, using the visual features of the target image as a reference condition and the position range of the second object in the layout image as a spatial constraint condition. The generation module is further configured to add the second object, which is generated in the layout image and matches the visual features of the target image, to the target image to obtain the target image containing the second object.
8. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the scene-adaptive image generation method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the scene-adaptive image generation method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the scene-adaptive image generation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image editing method and device, equipment and storage medium
CN121033227A