Image generation method and storage medium

By segmenting images and processing anchor box parameters, and combining text information to generate images, this technology solves the problems of lack of physical mechanism and insufficient scene adaptation in the generation of infrared animal images in existing technologies, and achieves high-quality image generation that is suitable for applications such as autonomous driving.

CN122115616APending Publication Date: 2026-05-29HEFEI YINGJU INNOVATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI YINGJU INNOVATION TECHNOLOGY CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing image generation technologies lack domain-specific physical mechanism modeling, geometric consistency constraints, and scene diversity adaptability, resulting in distortions in the size, perspective, and occlusion relationships of generated infrared animal images, making it difficult to meet the high-fidelity requirements of applications such as autonomous driving.

Method used

By segmenting the first image and processing the anchor box parameters, and combining the text information, a target image is generated. By using object detection models, segmentation models, preset depth estimation models, and conditional diffusion models, the realism and credibility of the generated image are ensured.

Benefits of technology

The generated images have significantly improved realism and credibility, meeting the quality standards for practical applications. They can generate a large number of rich and diverse visible light or infrared images, adapting to complex and ever-changing background environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115616A_ABST
    Figure CN122115616A_ABST
Patent Text Reader

Abstract

The present application relates to the field of image generation, in particular to an image generation method and a storage medium, the image generation method comprising: respectively performing image processing on a first image and a second image to obtain a segmentation image dataset corresponding to the first image and an anchor box parameter set corresponding to the second image; and performing image generation processing based on the segmentation image dataset and the anchor box parameter set according to text information to obtain a first target image. The present application improves the diversity and controllability of the generated image, enhances the layout rationality of the generated image, ensures that the output first target image has high precision and high reliability, and meets the quality standards of practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image generation technology, and more specifically to an image generation method and storage medium. Background Technology

[0002] Currently, although advanced image generation methods have shown great potential beyond specific fields and can be extended to various visual tasks, they still have limitations in specific application scenarios.

[0003] Taking the generation of in-vehicle infrared animal images, which is crucial for autonomous driving, as an example, existing technologies suffer from the following problems: First, most of them are based on visible light image data and lack a dedicated generation mechanism for infrared imaging mechanisms, resulting in the synthesized results failing to accurately reflect the thermal radiation characteristics and temperature gradient information of the target. Second, existing methods typically ignore the geometric constraints of the physical world, such as depth estimation and camera pose, making the generated animal targets prone to severe distortions in size, perspective, and occlusion relationships. Finally, due to the difficulty in modeling the presentation patterns of animals of different species and postures in complex and varied background environments, the images generated by existing technologies are insufficient in terms of scene diversity, resulting in a large gap between the synthesized results and the distribution of real data. These problems collectively restrict the quality of the generated images, making it difficult to meet the requirements of downstream perception models for high-fidelity training data, thereby limiting the full realization of the target perception potential. Summary of the Invention

[0004] One objective of this invention is to provide an image generation method and storage medium that addresses the technical problems of existing image generation technologies lacking domain-specific physical mechanism modeling, geometric consistency constraints, and scene diversity adaptability.

[0005] In a first aspect, embodiments of the present invention provide an image generation method, comprising: Image processing is performed on the first image and the second image respectively to obtain the segmented image dataset corresponding to the first image and the anchor box parameter set corresponding to the second image; Based on the text information, image generation processing is performed using the segmented image dataset and the anchor box parameter set to obtain the first target image.

[0006] In a second aspect, an image generation apparatus is provided, comprising: The processing unit is used to perform image processing on the first image and the second image respectively to obtain the segmented image dataset corresponding to the first image and the anchor box parameter set corresponding to the second image; The processing unit is further configured to perform image generation processing based on the text information, the segmented image dataset, and the anchor box parameter set to obtain a first target image.

[0007] In a third aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the image generation method as described in the first aspect.

[0008] In the embodiment implemented by the above image generation method and storage medium, image processing is first performed on the first image and the second image respectively to obtain the segmented image dataset corresponding to the first image and the anchor frame parameter set corresponding to the second image. The first image is any image in the first image database, and the second image is any image in the second image database. Then, based on the text information, image generation processing is performed on the segmented image dataset and the anchor frame parameter set to obtain the first target image. Finally, the first target image is filtered to obtain the first target image that meets the preset conditions as the second target image. This embodiment, by processing independent databases containing the target (first image) and the background (second image) respectively and combining them with text information, can flexibly generate a large number of rich and diverse visible light or infrared images. At the same time, the text information provides precise control over the generated content, meeting customization needs. Furthermore, the anchor frame parameter set derived from the second image (background) is used to guide the placement of the target object, ensuring that the generated target matches the physical laws of the background scene (such as "near objects appear larger and far objects appear smaller") in terms of size, position, and perspective, thereby significantly improving the realism and credibility of the final image and meeting the quality standards of practical applications. Attached Figure Description

[0009] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0010] Figure 1 This is a schematic diagram of the framework of an image generation system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an image generation method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of multiple individual unit segmentation images of a first image in an embodiment of the present invention; Figure 4 This is a schematic diagram of a first target image after superposition and fusion in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an image generation device according to an embodiment of the present invention. Detailed Implementation

[0011] To facilitate understanding of the present invention, a more detailed description is provided below with reference to the accompanying drawings and specific embodiments. It should be noted that when an element is described as being "fixed to" another element, it can be directly on the other element, or one or more intermediate elements may exist between them. When an element is described as being "electrically connected" to another element, it can be directly connected to the other element, or one or more intermediate elements may exist between them. The terms "upper," "lower," "inner," "outer," "bottom," etc., used in this specification indicate orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0012] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0013] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0014] Please see Figure 1 , Figure 1 This is a schematic diagram of the image generation system. Figure 1 In the image generation system 10, the overall architecture includes an object detection model 20, a segmentation model 30, a preset depth estimation model 40, a conditional diffusion model 50, a preset detection model 60, a processor 70, and a memory 80.

[0015] Specifically, the object detection model 20 can be YOLO or Faster R-CNN, without being limited to any particular model. The input to the object detection model 20 is real-world photographic data containing outdoor scene targets (such as animals, vehicles, pedestrians, etc.). Further, as a preprocessing step for the segmentation model 30, this model first quickly locates all instances of interest in the input image and accurately marks them with bounding boxes; then it outputs the bounding box coordinates for each image. These coordinates provide accurate regions of interest for the subsequent segmentation model, improving the efficiency and specificity of the segmentation process.

[0016] Specifically, segmentation model 30 can be U-Net or Mask R-CNN, but this is not a strict limitation. The input to segmentation model 30 can be the original outdoor scene photo and the bounding box coordinates from object detection model 20. Further, within the specified bounding box region, precise pixel-level recognition of the target pixels is performed, separating them from the original background to achieve a matting effect, thereby outputting high-quality individual segmentation images. The backgrounds of these images are transparent, retaining only the target itself (such as an animal), preparing core materials for subsequent image synthesis.

[0017] Specifically, the preset depth estimation model 40 can be a hybrid CNN + Transformer structure, without being limited to a single architecture. The input to the preset depth estimation model 40 is a background image (such as a vehicle infrared background image, a visible light driving scene image, etc.). The preset depth estimation model 40 analyzes the texture, lighting, and structural information in the 2D background image to infer the relative distance from each pixel in the scene to the camera, and then outputs a dense depth map of the same size as the background image. The value of each pixel in the map represents the depth of that point, providing crucial geometric constraints for subsequent anchor box generation and the proper placement of targets in 3D space.

[0018] Specifically, the conditional diffusion model 50 can be a diffusion model generator based on the U-Net structure, without being limited to a single specific model. The input to the conditional diffusion model 50 can be a single-unit segmentation image (from the segmentation model), a background image (from external input), text instructions (such as "a deer is in the left front"), fusion features, and anchor box parameters (derived from depth maps, text, etc.). Further, using text instructions and depth information as conditions, through a progressive denoising iterative process, the single-unit segmentation image is drawn onto the specified position in the background image, and intelligent fusion of details such as lighting and edges is performed to make it appear as if it truly exists in the scene. Therefore, a preliminary synthesized first target image is output.

[0019] Specifically, the preset detection model 60 can be a pre-trained image quality assessment network or a classification network; this is not a strict limitation. The input to the preset detection model 60 can be the first target image generated by the conditional diffusion model 50. The preset detection model 60 performs a comprehensive evaluation of the input image to determine whether it meets the preset quality standards. Evaluation dimensions may include: image sharpness, physical plausibility (e.g., whether the target is "floating" or sized incorrectly), semantic consistency with the text instruction, and the presence of obvious artifacts. Finally, it outputs a quality score or a pass / fail binary judgment. This output is used for automatic filtering; only images with a score higher than the threshold or judged as pass are retained as the final second target image.

[0020] The processor 70 is responsible for running the computational tasks of all the above models, while the memory 80 is used to store model parameters, input / output data, and intermediate calculation results to ensure that the entire process can be executed smoothly.

[0021] Specifically, processor 70 is configured to support the image generation system in executing the corresponding functions of the image generation method in the embodiments of the image generation method provided by the present invention. The processor 70 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0022] Specifically, the processor 70 may include a transmitting card, a receiving card, and a driver chip.

[0023] The memory 80 is used to store storage components such as program code, ensuring data storage and management. The memory 80 may include volatile memory (VM), such as random access memory (RAM); the memory 80 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 80 may also include a combination of the above types of memory.

[0024] Therefore, the above-mentioned multiple models work together within the same system framework to ensure the generation of high-quality target images from input images and to effectively evaluate and control their quality.

[0025] See Figure 2 , Figure 2This is a schematic flowchart of an image generation method provided in an embodiment of the present invention, which includes the following steps S10-S30: S10. Perform image processing on the first image and the second image respectively to obtain the segmented image dataset corresponding to the first image and the anchor box parameter set corresponding to the second image.

[0026] The first image refers to any outdoor image selected from the first image database. These images may include natural landscapes, buildings, animals, etc., and are not limited to a single image. For example, a clear photo of a tiger taken at a zoo, a photo of a car driving on a road, or a photo of a person walking on the street.

[0027] The first image database refers to a collection that stores tens of thousands of first images. For example, a database called "Real Animals" contains tens of thousands of high-resolution photos of different animals in different environments.

[0028] Among them, the segmented image dataset refers to the dataset in which a first image is decomposed into multiple regions with specific meanings (such as foreground and background) using image segmentation technology, and the data of these regions are organized into a dataset.

[0029] For example, from a photo of a tiger, a PNG image is obtained that contains only the tiger's outline and has a completely transparent background. This PNG image is a sample in the segmented image dataset.

[0030] The second image refers to any background image selected from the second image database. These images may be used to provide environmental information or background reference, and are not limited to a single image. For example, an infrared image of the road ahead taken from inside a car, a photograph of an open grassland, or a distant view of a city street.

[0031] The second image database refers to a collection that stores a large number of second images. For example, a database called "vehicle infrared background" contains infrared video frames taken under various weather conditions (sunny, rainy) and road conditions (highway, urban). In some embodiments, the second image database contains visible light images taken under various weather conditions (sunny, rainy) and road conditions (highway, urban).

[0032] The second image database can be an image database acquired by a vehicle-mounted image acquisition device, or an image database acquired by an image acquisition device from a drone or gimbal.

[0033] The first image can be a single frame or multiple frames, and there is no specific limitation here; the second image can be a single frame or multiple frames, and there is no specific limitation here.

[0034] The anchor frame parameters can include information such as the center coordinates, width, and height of the anchor frame. Therefore, the anchor frame parameter set refers to a set of suggested parameters describing where and how to place the target object in the second image; each parameter includes position (x, y coordinates), width, height, and aspect ratio. It is calculated based on the 3D spatial information of the image.

[0035] For example, for a road infrared image, the anchor frame parameter set is [x=100,y=600, w=120, h=80] and [x=80,y=500, w=130, h=70], which respectively means "place an anchor frame with a width of 120 pixels and a height of 80 pixels at coordinates (100,600)" and "place an anchor frame with a width of 130 pixels and a height of 70 pixels at coordinates (80,500)".

[0036] In a specific implementation, S10 can perform image processing on the first image and the second image in parallel.

[0037] The specific implementation process of S10 can be found in the detailed descriptions of S101-S103, and will not be repeated here.

[0038] S101. In one embodiment, the step of image processing the second image to obtain the anchor box parameter set corresponding to the second image includes: inputting the second image into a preset depth estimation model for prediction to obtain a dense depth map output by the preset depth estimation model; calculating the depth distribution features of the second image based on the dense depth map to obtain a depth distribution histogram; dividing the depth distribution histogram into distance intervals according to a preset depth value layering strategy to obtain multiple different distance intervals; obtaining a first anchor box parameter corresponding to each distance interval based on the multiple different distance intervals; adjusting the first anchor box parameter of each distance interval to obtain the adjusted first anchor box parameter of each distance interval; and summarizing the adjusted first anchor box parameters of each distance interval to obtain the anchor box parameter set corresponding to the second image.

[0039] The preset depth estimation model is a trained neural network capable of inferring the 3D structure of a scene from a single 2D image. This preset depth estimation model can be a hybrid structure combining CNN and Transformer.

[0040] The dense depth map is the output of a preset depth estimation model. It is a grayscale image of the same size as the original image. The model calculates a depth value for each pixel in the image, and the brightness value of each pixel represents its distance from the camera. Generally, closer objects have brighter (or darker, depending on the definition) pixel values, while farther objects have darker (or brighter) pixel values.

[0041] The depth distribution histogram displays the distribution of different depth values ​​in an image, used to analyze the depth features of a scene. The X-axis represents the depth value (distance), and the Y-axis represents the number of pixels corresponding to that depth value.

[0042] Specifically, in the specific implementation of calculating the depth distribution features of the dense depth map to obtain the depth distribution histogram, each pixel of the depth map is traversed and its depth value is recorded; the depth value range (e.g., 0-100 meters) is divided into several small "buckets"; the number of pixels falling into each "bucket" is counted; and a histogram is drawn with the depth value as the X-axis and the number of pixels as the Y-axis.

[0043] For example, for road images, the depth distribution histogram shows a high peak in the 0-20 meter range (near-ground road surface), a lower peak in the 20-60 meter range (mid-ground road surface), and a gentle peak above 60 meters (far-ground and sky).

[0044] The preset depth value layering strategy refers to predefined rules used to divide continuous depth values ​​into several meaningful distance intervals. Furthermore, multiple different distance intervals refer to the result of layering, which divides the scene into several logical levels.

[0045] Specifically, the preset depth value layering strategy can divide the entire scene into multiple distance intervals based on depth values ​​(i.e., distance), such as: [0-10 meters], [10-30 meters], [30-60 meters], [above 60 meters]. A corresponding set of anchor frame dimensions and aspect ratios is automatically generated for each distance interval, with different anchor sizes set for different distance segments based on the depth estimation results.

[0046] The anchor box parameter set refers to the set of bounding box parameters for the region of interest in the image, including position, size, and aspect ratio. Specifically, the first anchor box parameter refers to the basic anchor box size and aspect ratio corresponding to each distance interval, calculated based on physical rules.

[0047] Specifically, the real-world dimensions of the target object are preset (e.g., an adult wild boar: 1.5 meters long and 0.8 meters high) and camera intrinsic parameters (focal length, etc.). Based on the perspective projection formula pixel size = (real size × focal length) / distance, the pixel size that the real object should appear in the image at different distances is calculated. Therefore, for the [0-10 meters] range, the calculated pixel size may be very large, for example, 180px wide and 96px high. For the [10-30 meters] range, the size may become 60px wide and 32px high. For the [60 meters and above] range, the size may only be 15px wide and 8px high.

[0048] Specifically, the parameters of the first anchor frame for each distance interval are adjusted to adapt to different scenario requirements and target sizes.

[0049] Optionally, the adjustment process can be tailored to analyze gradient changes in the dense depth map (i.e., places where depth changes abruptly, typically at the edges of objects), tending to place the anchor frame in areas with relatively flat and reasonable depth (such as a road surface) rather than in areas with dramatic depth changes (such as in the air or on tree trunks). If the background is a slope, the aspect ratio of the anchor frame can be slightly adjusted based on the continuous changes in the dense depth map, making it appear more "attached" to the slope rather than floating.

[0050] Furthermore, the specific implementation process of adjusting the first anchor frame parameters for each distance interval to obtain the adjusted first anchor frame parameters for each distance interval can be referred to the description in S102, and will not be repeated here.

[0051] This involves aggregating the adjusted anchor frame parameters for each distance range to form a complete set of anchor frame parameters. Specifically, it involves collecting finely adjusted anchor frames from all ranges, including foreground, midground, and background, and formatting them uniformly to form a complete data structure.

[0052] Therefore, the set of anchor box parameters corresponding to the second image may include: [anchor box 1: position (x1,y1), size (w1,h1)], [anchor box 2: position (x2,y2), size (w2,h2)], [anchor box 3: position (x3,y3), size (w3,h3)], etc., and each anchor box has its own reasonable attributes in 3D space.

[0053] As can be seen, in this embodiment, the 2D background image is transformed into a 3D anchor box parameter. That is, the second image is processed to generate a set of anchor box parameters. These parameters are used for further image analysis and applications, such as object detection and scene understanding, to ensure the accuracy and flexibility of the final synthesized image.

[0054] S102. In one embodiment, adjusting the first anchor frame parameter of each distance interval to obtain the adjusted first anchor frame parameter of each distance interval includes: performing boundary gradient analysis on the dense depth map to obtain the edge information of each candidate body in the dense depth map; performing spatial aggregation feature analysis on the dense depth map to obtain the spatial aggregation feature of each candidate body in the dense depth map; and adjusting the first anchor frame parameter of each distance interval according to the edge information and spatial aggregation feature of each candidate body to obtain the adjusted first anchor frame parameter of each distance interval.

[0055] Boundary gradient analysis refers to identifying the edges of objects in an image by calculating the rate of change of depth values. These edges are typically locations where depth values ​​change significantly. Boundary gradient analysis is used to detect locations in an image where pixel values ​​undergo drastic changes.

[0056] Candidate objects refer to potential target objects identified through analysis in dense depth maps.

[0057] Edge information refers to the result obtained after analysis; it is an edge map that marks the locations of depth abrupt changes. These edges delineate the contours of different objects or planes in the scene.

[0058] In practice, edge detection operators (such as Sobel and Canny operators) are applied to the dense depth map. These operators calculate the depth difference between each pixel in the dense depth map and its surrounding pixels. If a pixel's depth value differs significantly from its neighbors (e.g., a sudden change from a road surface (5 meters) to a tree trunk (4 meters)), that pixel will have a high gradient value. A further threshold is set to connect all pixels with high gradient values, forming lines, i.e., edge information.

[0059] Spatial aggregation features refer to the spatial distribution features of objects in a dense depth map, including their shape, size, and positional relationships.

[0060] Spatial aggregation feature analysis refers to identifying the features of candidate objects by analyzing the spatial distribution and aggregation patterns of pixels.

[0061] This involves applying region growing, clustering algorithms (such as DBSCAN), or semantic segmentation networks to the dense depth map. The algorithm starts with a seed pixel and aggregates all spatially contiguous pixels with similar depth values ​​to form a region. This process is repeated throughout the dense depth map, ultimately segmenting the image into multiple candidate regions with different depth characteristics. For each aggregated region, its area, average depth, and other features are calculated.

[0062] Therefore, edge information can be understood as indicating which are the boundaries of an object, and the anchor frame should not cross these edges, otherwise it would lead to the absurd result that a target is "half on the wall and half outside the wall"; spatial aggregation features indicate which areas are flat, open and suitable for placing objects (such as candidate A), and which areas are narrow or irregular (such as candidate C) and not suitable for placing objects.

[0063] In practice, a first anchor frame parameter (e.g., an anchor frame located in the mid-field) is placed on a dense depth map. The fit score is calculated as follows: Penalty (from edge information): Check if the anchor frame has significant overlap with "edge information." If the overlap is high, the score is significantly reduced. This prevents the anchor frame from crossing object boundaries. Reward (from spatial aggregation features): Check if the center and body of the anchor frame fall on a high-quality candidate object (e.g., a large, flat candidate object A). If so, the score is increased. Additionally, if the size of the anchor frame matches the size of the candidate object it belongs to (e.g., a medium-sized anchor frame placed on a wide road), additional points are awarded. Further small-scale searches and fine-tuning are performed near the original anchor frame position (e.g., moving it a few pixels up, down, left, or right, slightly scaling the size) to find the position and size that maximizes the fit score; this optimal position and size are recorded as the adjusted anchor frame parameter.

[0064] For example, the initial state is a mid-range anchor frame generated based on physics rules, with its center falling precisely on a tree trunk (candidate C). Edge analysis shows that the anchor frame area highly overlaps with the tree trunk edge, resulting in a high penalty. Spatial aggregation analysis shows that candidate C, where the anchor frame is located, has a small area, resulting in a low reward. Therefore, the anchor frame is automatically shifted to the side, so that it falls entirely on a wide road (candidate A). Thus, the anchor frame's position is successfully adjusted, moving from an unreasonable tree trunk position to a reasonable road surface position.

[0065] As can be seen, through the analysis and adjustment process in this embodiment, candidate objects in the dense depth map are accurately identified and located, thereby optimizing the anchor frame parameters to support various visual recognition and processing tasks.

[0066] S103. In one embodiment, the step of performing image processing on the first image to obtain a segmented image dataset corresponding to the first image includes: performing target detection on the first image to obtain detection information for each target in the first image; and performing image segmentation processing on the first image based on the detection information for each target to obtain a segmented image dataset corresponding to the first image, wherein the segmented image dataset includes a single-unit segmentation map of each target.

[0067] Object detection is used to identify and locate target objects in an image and generate a bounding box for each object, while also performing classification processing. Object detection can be performed using object detection models, which may include, but are not limited to, YOLO (You Only Look Once), Faster R-CNN, or improvements to the backbone network of YOLO.

[0068] Furthermore, classification processing refers to classifying the results output by the target detection model to determine which preset category each detected object belongs to (for example, classifying detected animals as "lion" or "tiger").

[0069] In practice, the first image is input into the object detection model. The further object detection model extracts image features through a convolutional neural network (CNN), and then uses these features to identify and locate objects. Finally, the object detection model outputs detection information for each object, including object category, bounding box coordinates, and confidence score.

[0070] In this context, the target object is the specific object in the first image that you want to identify and separate. For example, in a photo taken on the African savanna, there are two lions and a tree. If the focus is on the animals, then the two lions are the target objects, while the tree and the grassland are the background.

[0071] The detection information may include, but is not limited to, the target object's category (such as people, vehicles, or animals), location (bounding box coordinates), and confidence score.

[0072] Specifically, the bounding box can be a rectangle, using four numbers (usually the center coordinates x and y, the width w, and the height h) to precisely locate the target's position in the image.

[0073] For example, for the first lion, the detection information might be a bounding box [x1, y1, w1, h1] and a category label "lion" with a confidence level of 98%. For the second lion, it would be another bounding box [x2, y2, w2, h2] and the same label.

[0074] Image segmentation is the process of dividing an image into multiple meaningful regions or objects, typically used to extract the contours of specific targets to obtain the outline shape of a particular animal. Image segmentation can be performed using segmentation models, which may include, but are not limited to, U-Net or Mask R-CNN.

[0075] In this context, a single-entity segmentation image refers to the independent segmentation result of each target object, showing the specific shape and location of the target. It is an image that perfectly matches the shape of the target object, but with a transparent background.

[0076] For example, from the detection box of the first lion, we obtain a PNG image that only contains the outline of the lion. The lion itself is preserved as is, while the pixels outside the lion are transparent, resulting in a single-unit segmentation map.

[0077] Specifically, the bounding box information obtained from object detection is used to narrow the processing range of image segmentation, improve segmentation efficiency and accuracy, and further use the segmentation model to perform fine segmentation of the target objects and extract the contour of each target object; the segmentation results of each target object are organized into a segmented image dataset, which contains the individual segmentation map of each target object.

[0078] In practice, the first image and the detection information (i.e., bounding boxes) of each target obtained in the previous step are input into an instance segmentation model. The segmentation model receives the image and two bounding boxes. It ignores all regions outside the boxes and focuses its computational resources only on the regions inside the boxes. Within the first bounding box, the model analyzes each pixel to determine whether it belongs to the "lion" or the "background." The model's output is a mask, which is a black-and-white binary image of the same size as the bounding box region. The pixels corresponding to the lion are white (value 1), and the background pixels are black (value 0). This mask is then used to extract the lion from the original first image, retaining the original image pixels corresponding to the white areas in the mask and setting the pixels corresponding to the black areas to transparent. This process is repeated for the second bounding box to obtain the individual segmentation image of the second lion. Organizing, naming, and storing these two lion segmentation images constitutes the segmented image dataset corresponding to the first image.

[0079] In some embodiments, when performing detection and segmentation processing on real animal images, image segmentation operations are performed based on the obtained animal detection information (detection boxes and animal categories). Thresholding, contour detection, and texture analysis techniques are used to process information such as animal contour edges, textures, and colors, ultimately obtaining high-quality individual animal segmentation images of different species.

[0080] The individual segmentation diagrams for each target body can be referenced. Figure 3 , Figure 3 This is a schematic diagram of multiple individual segmentation images of the first image. The left side is the first image, and the right side is two individual segmentation images obtained after image processing.

[0081] It should be noted that in some embodiments, S103 and S101-S102 are not sequential and can be processed in parallel.

[0082] As can be seen, by detecting first and then segmenting, this embodiment can efficiently and accurately transform the target object in the original image into a segmented image dataset, laying a solid foundation for the subsequent image generation steps.

[0083] S20. Based on the text information, perform image generation processing on the segmented image dataset and the anchor box parameter set to obtain the first target image.

[0084] Text information refers to the textual descriptions used to guide image generation, which may include information such as location, object type, and color.

[0085] Text information can be obtained in various ways, including direct input and speech-to-text conversion. Users can specify the location and content of the generated image. For example, specifying that an animal image be generated on the left side of the road.

[0086] Image generation processing can be performed using a conditional diffusion model (or a similar generative model).

[0087] Optionally, the text information (such as "a deer is on the left side of the road") is converted into a high-dimensional mathematical vector by a text encoder (such as CLIP's TextEncoder). This vector captures the semantics of the text, such as the visual concept of the deer, the visual concept of the road, and the spatial relationship to the left. Individual segmentation images related to the deer are selected from the segmented image dataset. Simultaneously, the anchor box parameter set is loaded. Finally, the text vector, the selected deer segmentation image, the background image, and the anchor box parameter set are all input into the conditional diffusion model. The model then receives the instruction: "On this background image, referring to these anchor box positions, and based on the semantics of 'a deer is on the left side of the road,' fuse the image of the deer into it." Therefore, the model outputs the first target image. The first target image already contains the target object.

[0088] The first target image is the preliminary composite result output from the image generation process. It already contains the target object, but may still have some minor flaws, such as unnatural edge blending or deviations in lighting and shadow.

[0089] For example, consider a vehicle-mounted infrared background image. Its anchor box parameter set contains candidate boxes from multiple locations, including near-field, mid-field, and far-field. User command 1 (simple): "Generate a deer on the road." Understanding "deer" and "road," it selects a suitable anchor box (possibly mid-field or near-field) located within the road area from the anchor box set to generate the deer. User command 2 (precise): "Generate a running deer on the left side of the road." Resolving to "left side of the road," it filters out all anchor boxes located on the right or in the center, selecting only from the candidate anchor boxes on the left. Resolving to "running," it prioritizes segmentation images of running deer from the "segmentation image dataset." Finally, a conditional diffusion model generates a running deer within a reasonable anchor box on the left side of the road, matching it with motion blur (if present) in the background. User instruction 3 (via voice): The user speaks into the microphone, "Hey, add a little fox not far in front of the car." The speech is converted into text "Add a little fox not far in front of the car," and then the above parsing and generation process is repeated, finally generating a little fox in the foreground area below the image.

[0090] For details on the specific implementation process of S20, please refer to the descriptions in S201-S205, which will not be repeated here.

[0091] As can be seen, the image generation process that combines text information in this embodiment can achieve precise control and generation of the target image.

[0092] In one embodiment, step S20, performing image generation processing on the segmented image dataset and the anchor box parameter set based on the text information to obtain a first target image, further includes S201-S205: S201. Perform text recognition processing on the text information to obtain the corresponding speech embedding vector.

[0093] The text recognition processing involves converting text information into a format that can be used for further processing, such as a speech embedding vector.

[0094] Among them, the speech embedding vector is a numerical representation that captures the speech features in the text information and is used to fuse with other features.

[0095] Specifically, text information (such as "a little fox is hiding in the grass") is input into a pre-trained text encoder (such as the Text Encoder of BERT or CLIP), which then outputs a speech embedding vector. This speech embedding vector not only encodes nouns such as "fox" and "grass", but also encodes the size, state and spatial relationship implied by adjectives and verbs such as "little" and "hiding in".

[0096] Optionally, the input text information is first parsed to extract key semantics and instructions. Then, text-to-speech (TTS) technology is used to convert the text into a speech signal. Finally, a speech embedding model (such as Wav2Vec or other deep learning models) is used to convert the speech signal into a speech embedding vector.

[0097] S202. Obtain the dense depth map corresponding to the second image; perform image processing on the dense depth map to obtain the corresponding depth estimation features.

[0098] The depth estimation features refer to features extracted from dense depth maps, which are used to describe the depth information of the scene.

[0099] Specifically, the dense depth map corresponding to the second image is input into a feature extraction network (such as a small CNN or Vision Transformer); then the depth estimation features are output. This feature map is no longer a simple grayscale, but an abstract representation that marks high-level geometric information such as "flat regions", "vertical planes", and "depth abrupt boundaries".

[0100] It is important to note that feature extraction can be performed on both the text and the background image in parallel.

[0101] S203. Perform dynamic feature fusion processing on the speech embedding vector and the depth estimation features to obtain fused features.

[0102] The dynamic feature fusion process can use a learnable mechanism (such as an attention mechanism) to allow text information to actively query and focus on the most relevant parts of deep features.

[0103] Among them, fusion features refer to features obtained through the dynamic fusion of multimodal data (such as speech embeddings and deep features), which contain rich contextual information.

[0104] Optionally, the speech embedding vector and the depth estimation features can be dynamically fused using a gating-based dynamic attention module to obtain fused features.

[0105] In one embodiment, step S203 further includes: acquiring preset semantic information; calculating the semantic similarity between the text information and the preset semantic information to obtain a target semantic similarity; adjusting the weights of the speech embedding vector and the depth estimation feature according to the target semantic similarity to obtain a first weight corresponding to the speech embedding vector and a second weight corresponding to the depth estimation feature; and performing dynamic feature fusion processing on the speech embedding vector and the depth estimation feature according to the first weight and the second weight respectively to obtain fused features.

[0106] Among them, preset semantic information refers to predefined semantic content or concepts used to make semantic comparisons with the input text. It is usually an existing knowledge base or semantic tags.

[0107] Specifically, a database or knowledge base containing various semantic information is constructed, covering relevant concepts in the target application domain. Pre-defined semantic information relevant to the current task or application scenario is extracted from this semantic base. In natural language processing tasks, this pre-defined semantic information can be key terms or concepts in a specific domain.

[0108] Semantic similarity measures the semantic closeness of two texts, obtained by calculating their distance in vector space (such as cosine similarity). The value range is usually between [-1, 1] or [0, 1], with values ​​closer to 1 indicating greater similarity.

[0109] Among them, target semantic similarity refers to the similarity score between text information and preset semantic information.

[0110] Specifically, the pre-defined semantic information and text information are converted into vector representations, which can be achieved through word embedding (such as Word2Vec) or sentence embedding (such as BERT). Furthermore, similarity measurement methods (such as cosine similarity and Euclidean distance) are used to calculate the semantic similarity between the two, and the target semantic similarity is obtained based on the calculation results, which is used for subsequent weight adjustment.

[0111] In this context, weight refers to the degree of importance of different features in feature fusion, which is reflected by the weight value.

[0112] Specifically, the first weight refers to the weight assigned to the text information (speech embedding vector). The higher the weight, the more strictly the generation process will follow the semantics of the text.

[0113] Specifically, the second weight refers to the weight assigned to the scene's geometric information (depth estimation features). The higher the weight, the more the generation process will prioritize the physical plausibility of the scene.

[0114] In practical implementation, based on the target semantic similarity, the first weight and the second weight can be calculated using a learnable function or a simple linear mapping. Taking a simple linear mapping as an example: First weight = Target semantic similarity; Second weight = 1 - Target semantic similarity. When semantic information is ambiguous, the system can automatically reduce the semantic feature weight and increase the depth estimation feature weight; conversely, when the semantic instructions are clear and the depth estimation is uncertain, the semantic feature weight is increased to ensure the stability and accuracy of the generated conditions.

[0115] Specifically, the speech embedding vector is multiplied by a first weight; the depth estimation features are multiplied by a second weight, and these two weighted features are fused (e.g., through concatenation or addition) to generate fused features. Alternatively, neural networks or other fusion algorithms (such as weighted averaging) can be used to achieve feature fusion.

[0116] As can be seen, this embodiment can effectively utilize semantic information and similarity to optimize the feature fusion process, thereby improving the accuracy and robustness of the final image generation method.

[0117] S204. Adjust the target anchor frame parameters in the anchor frame parameter set according to the fusion features to obtain multiple adjusted second anchor frame parameters, wherein the target anchor frame parameters are the anchor frame parameters corresponding to the text information.

[0118] The target anchor box parameters are the anchor box parameters corresponding to the text information. Specifically, in the set of anchor box parameters, candidate anchor box parameters that match the semantic content of the text information are selected, or an adjustment benchmark for the candidate anchor box parameters is determined. Specifically, the semantic features of the text (e.g., 'a running dog') are analyzed, and specific anchor box parameters suitable for carrying the semantic object (i.e., suitable for placing 'dog') are identified from the set of anchor box parameters generated based on the second image. These identified anchor box parameters serve as target anchor box parameters. Subsequently, based on the fusion features (including text and depth information), these specific target anchor box parameters are finely adjusted in position, size, or shape to ensure that the object described in the text accurately fits the appropriate position in the second image in the generated image. Specifically, in the process of adjusting the target anchor frame parameters in the anchor frame parameter set according to the fusion features to obtain multiple adjusted second anchor frame parameters, information related to anchor frame adjustment is extracted from the fusion features. This information may include the object's positional tendency, size adjustment suggestions, etc. Further, a neural network model (such as a fully connected layer) is used to map the fusion features onto the anchor frame parameter adjustment values. These values ​​may be offsets or scaling factors used to adjust the center position and size of the anchor frames. Further, the center coordinate adjustment value of each anchor frame is calculated using the fusion features to determine the precise position of the anchor frame in the image. The width and height adjustment values ​​of the anchor frames are calculated to adapt to the actual size of the object. Combining the information in the fusion features, the adjustment factor of each anchor frame is dynamically determined to improve detection accuracy. Therefore, the calculated adjustment values ​​are applied to the target anchor frame parameters to update the center coordinates and size of the anchor frames, ensuring that the adjusted anchor frames remain within the image boundaries and avoid exceeding the image range. The effectiveness of the adjusted second anchor frame parameters can be evaluated using a loss function for object detection tasks (such as intersection-union loss). Backpropagation can be used to optimize the anchor frame adjustment model, minimize the loss function, and improve the accuracy of anchor frame localization.

[0119] In the specific implementation of adjusting the target anchor frame parameters in the anchor frame parameter set according to the fusion feature to obtain multiple adjusted second anchor frame parameters: Step 1: Spatially map each abstract coordinate parameter in the target anchor frame parameter set to the fusion feature. Specifically, iterate through each target anchor frame (e.g., anchor frame B) in the anchor frame parameter set. The parameters of anchor frame B are [x, y, w, h] (center coordinates and width and height). Then, find the part on the fusion feature that completely corresponds to the rectangular area of ​​[x, y, w, h]. Step 2: Calculate the fit score for each aligned anchor frame area. Specifically, calculate the average pixel value of the area covered by each anchor frame on the fusion feature. Fit score = mean(fusion feature [all pixels in the anchor frame area]). Check whether the anchor frame area contains the local peak point (i.e., the brightest point) of the fusion feature. An ideal anchor frame should cover a bright area, and its center should be as close as possible to the peak point. The fit score is used to measure the degree of matching between an original anchor frame and the best intention expressed by the fusion feature. The higher the score, the more ideal the position and size of the anchor frame. Step 3: Set one or more thresholds and classify the anchor frames according to their fit scores. Specifically, when the fit score > high threshold (e.g., > 0.7), move the anchor frame around its current position in a small range (e.g., left, right, up, down, expand, shrink by a few pixels), calculate the fit score of the new position after each movement, move the center of the anchor frame and fine-tune its size along the direction that increases the score the fastest, until a local optimum is reached, and obtain the adjusted second anchor frame parameters. When the fit score is between low and high thresholds (e.g., 0.3 < fit score < 0.7), the anchor frame may not be far from the ideal position, but requires a large displacement. Therefore, directly find the region with the nearest score exceeding the high threshold on the fusion feature, and then jump the anchor frame to the center of that region to obtain the adjusted second anchor frame parameters. When the fit score < low threshold (e.g., < 0.3), the anchor frame is located in a region that does not conform to the instructions at all (e.g., let the fox "hide" in the sky or on the road). Keeping them will interfere with the generation or produce incorrect results. Therefore, the system will directly remove them from the candidate list, thus eliminating this anchor box. Step four: All the anchor box parameters that have been adjusted and corrected are summarized to obtain multiple adjusted second anchor box parameters.

[0120] For example, suppose the text information is "A deer is standing on the left side of the road". The fused features are bright in the left side of the road and dark elsewhere. The original set of anchor boxes is: {anchor box A: [x=100, y=600, w=120, h=80] (lower left of the image), anchor box B: [x=400, y=300, w=100, h=100] (front of the image), anchor box C: [x=700, y=500, w=80, h=80] (lower right of the image)}. Therefore, anchor box A (front left) scores 0.85 (high score), anchor box B (front) scores 0.15 (low score), and anchor box C (right) scores 0.10 (low score). Further, anchor box A is judged to have a high score. Through gradient ascent, its center is fine-tuned to the brightest position on the left, and its size is also fine-tuned according to the depth value of that point. It becomes an adjusted second anchor frame. Anchor frames B and C are judged to be low scores and are discarded directly. The final set of parameters for multiple adjusted second anchor frames is: { Anchor frame A': [x=105, y=605, w=125, h=82] (the finely adjusted anchor frame located on the left side of the road)}.

[0121] Therefore, this scheme can use fused features to precisely adjust the anchor frame parameters, thereby improving the accuracy and robustness of target detection.

[0122] The adjusted second anchor frame parameters refer to the exact position and size where each object should be placed.

[0123] As can be seen, this embodiment can achieve accurate generation and adjustment of target images by combining text information with image depth features. Furthermore, the use of dynamic feature fusion processing can flexibly handle multimodal data and improve the quality and accuracy of image generation.

[0124] S205. In one embodiment, obtaining the first target image based on the segmented image dataset and the adjusted plurality of second anchor box parameters includes: obtaining at least one target segmented image corresponding to the text information in the segmented image dataset; superimposing each target segmented image onto the second image according to the adjusted plurality of second anchor box parameters to obtain a superimposed image; and performing image local fusion processing on the superimposed image to obtain the first target image.

[0125] Here, image segmentation refers to dividing the original image into multiple regions or objects using image segmentation techniques, with each region corresponding to a segmented image. This can correspond to the single-unit segmentation image in S103.

[0126] Furthermore, a target segmented image refers to a single-object segmentation image that matches the text description, selected and extracted from the segmented image dataset of the first image based on the semantic content of the text information (e.g., "a deer").

[0127] For example, a segmented image dataset = {segmentation image A (a deer), segmentation image B (a tiger), segmentation image C (a deer), ...}. The text information = "A deer is standing on the left side of the road". Therefore, obtaining the target segmented image = identifying objects related to "deer" from the above segmented image dataset, thus the target segmented image = {segmentation image A, segmentation image C}.

[0128] Image overlay involves placing one image on top of another, which usually requires adjustments to the position and size.

[0129] Specifically, the position and size of each target segmented image in the second image are determined using the adjusted anchor frame parameters. The position and scale of the segmented images are adjusted according to the anchor frame parameters to ensure that they are accurately superimposed on the second image. The adjusted segmented images are then superimposed on the second image to generate a superimposed image.

[0130] In practice, a suitable segmentation image is selected from the segmentation image dataset based on the text information. For example, if the instruction is "a deer," the system will select a segmentation image of a deer. If there are multiple anchor boxes, multiple deer in different poses may be selected. For each second anchor box parameter [x, y, w, h] and the corresponding segmentation image, the size of the segmentation image is scaled so that its width and height perfectly match the w and h of the anchor box; further, the scaled segmentation image is moved onto the background image so that its center point is aligned with the (x, y) coordinates of the anchor box; each pixel in the overlay region is traversed, and if the alpha value of the segmentation image at that pixel is 1, the pixel color of the segmentation image is used directly. If the alpha value is 0, the pixel color of the background image is used. (For anti-aliased edges, the alpha value may be between 0 and 1, in which case weighted blending is performed).

[0131] Image local fusion processing refers to the process of intelligently identifying the boundary areas (fusion boundaries) between objects and the background, rather than processing the entire image uniformly, and then performing fine pixel-level repair and harmonization only in these local areas.

[0132] The first target image refers to the final output synthetic image with a high degree of realism.

[0133] Specifically, the image regions that need to be locally fused are determined, usually the boundary regions between the segmented image and the second image; image fusion algorithms (such as Gaussian mixture and Laplacian pyramid) are used to process the selected regions to smooth the transition; based on the visual effect, the fused regions are finely adjusted to ensure a natural transition of color and texture; after fusion processing, the first target image is generated.

[0134] Among them, you can refer to Figure 4 , Figure 4 This is a schematic diagram of the first target image after overlay and fusion.

[0135] As can be seen, this embodiment can effectively combine the segmented image with the base image through the two steps of overlay and local fusion, generating a visually natural and consistent target image. The image is generated based on physical consistency and semantic rationality, thereby improving the accuracy and credibility of the final generated image.

[0136] In some embodiments, the second image and the first target image are visible light images or infrared images.

[0137] In some embodiments, when the second image and the first target image are infrared images, after obtaining the first target image, the method further includes step S30: performing image filtering on the first target image, and obtaining a first target image that meets preset conditions as the second target image.

[0138] The second target image refers to the image that meets specific conditions after screening, and it is the final output result.

[0139] Among them, preset conditions refer to a set of standards or rules used to filter images, which are usually related to the target application.

[0140] In one embodiment, the step of filtering the first target image to obtain a first target image that meets preset conditions as a second target image includes: inputting the first target image into a preset detection model for detection to obtain the detection result output by the preset detection model; comparing the detection result with the text information to obtain a target similarity; when the target similarity is greater than or equal to the preset similarity, determining that the first target image meets the preset conditions, and saving the first target image that meets the preset conditions as the second target image; or, when the target similarity is less than the preset similarity, deleting the first target image.

[0141] The preset detection model can be a multimodal model, capable of simultaneously understanding images and text and determining the correlation between them. Examples include CLIP (Contrastive Language-Image Pre-Training) model, YOLO, and Faster R-CNN.

[0142] The detection result refers to the output of the preset detection model. The detection result can include information such as the object category, location, and confidence level identified in the image.

[0143] For example, input a pre-defined detection model with a newly synthesized "first target image" containing a deer on the left side of a road; output a 512-dimensional image embedding vector whose position in mathematical space is very close to other image vectors in the training set that contain concepts such as "deer," "road," and "left side."

[0144] Similarity calculation refers to converting the detection results and text information into vector representations using natural language processing techniques (such as BERT and Word2Vec), and then calculating the similarity between the two.

[0145] Among them, target similarity refers to the similarity value obtained from the calculation results, which reflects the degree of matching between image content and text information.

[0146] The preset similarity refers to a pre-set value. It is an empirical value used to judge whether the generated quality meets the standard. For example, it can be set to 0.85.

[0147] Specifically, if the target similarity is greater than or equal to a preset similarity, the first target image is considered to meet the criteria and is saved as the second target image. If the target similarity is less than the preset similarity, the image is considered to not meet the criteria and is deleted or excluded.

[0148] As can be seen, in this embodiment, image filtering can effectively filter out images that meet specific conditions, thereby improving the accuracy and relevance of the image generation method.

[0149] This embodiment processes independent databases containing the target (first image) and background (second image) separately, and combines them with text information to flexibly generate a large number of rich and diverse images. At the same time, the text information provides precise control over the generated content, meeting customization needs. Furthermore, the anchor frame parameter set derived from the second image (background) guides the placement of the target object, ensuring that the generated target matches the physical laws of the background scene (such as "near objects appear larger and far objects appear smaller") in terms of size, position, and perspective. This significantly improves the realism and credibility of the final image, meeting the quality standards for practical applications.

[0150] It is understood that the image generation method provided in this embodiment of the invention can be used in vehicle-mounted products, as well as in drones or gimbal products. This invention does not limit the specific application scenarios.

[0151] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.

[0152] As another aspect of the embodiments of this application, this application provides an image generation apparatus. The image generation apparatus can be a software module, which includes several instructions stored in a memory. A processor can access the memory, invoke the instructions, and execute them to complete the image generation methods described in the various embodiments above.

[0153] See Figure 5 , Figure 5 This is a schematic diagram of the structure of an image generation device provided in an embodiment of this application. For example... Figure 5 As shown, the image generation apparatus 500 includes: Processing unit 501 is used to perform image processing on the first image and the second image respectively to obtain the segmented image dataset corresponding to the first image and the anchor box parameter set corresponding to the second image; The processing unit 501 is further configured to perform image generation processing based on the text information, the segmented image dataset, and the anchor box parameter set to obtain a first target image.

[0154] This embodiment processes independent databases containing the target (first image) and background (second image) separately, and combines them with text information to flexibly generate a large number of rich and diverse images. At the same time, the text information provides precise control over the generated content, meeting customization needs. Furthermore, the anchor frame parameter set derived from the second image (background) guides the placement of the target object, ensuring that the generated target matches the physical laws of the background scene (such as "near objects appear larger and far objects appear smaller") in terms of size, position, and perspective. This significantly improves the realism and credibility of the final image, meeting the quality standards for practical applications.

[0155] In one embodiment, in the image processing of the second image to obtain the anchor frame parameter set corresponding to the second image, the processing unit 501 is further configured to: input the second image into a preset depth estimation model for prediction to obtain a dense depth map output by the preset depth estimation model; calculate the depth distribution features of the second image based on the dense depth map to obtain a depth distribution histogram; classify the depth distribution histogram into distance intervals according to a preset depth value layering strategy to obtain multiple different distance intervals; obtain a first anchor frame parameter corresponding to each distance interval based on the multiple different distance intervals; adjust the first anchor frame parameter of each distance interval to obtain an adjusted first anchor frame parameter for each distance interval; and summarize the adjusted first anchor frame parameters for each distance interval to obtain the anchor frame parameter set corresponding to the second image.

[0156] In one embodiment, in the process of adjusting the first anchor frame parameters of each distance interval to obtain the adjusted first anchor frame parameters for each distance interval, the processing unit 501 is further configured to: perform boundary gradient analysis on the dense depth map to obtain edge information of each candidate body in the dense depth map; perform spatial aggregation feature analysis on the dense depth map to obtain spatial aggregation features of each candidate body in the dense depth map; and adjust the first anchor frame parameters of each distance interval according to the edge information and spatial aggregation features of each candidate body to obtain the adjusted first anchor frame parameters for each distance interval.

[0157] In one embodiment, in the step of performing image generation processing on the segmented image dataset and the anchor box parameter set based on the text information to obtain a first target image, the processing unit 501 is further configured to: perform text recognition processing on the text information to obtain a corresponding speech embedding vector; obtain a dense depth map corresponding to the second image; perform image processing on the dense depth map to obtain a corresponding depth estimation feature; perform dynamic feature fusion processing on the speech embedding vector and the depth estimation feature to obtain a fused feature; adjust the target anchor box parameters in the anchor box parameter set according to the fused feature to obtain a plurality of adjusted second anchor box parameters, wherein the target anchor box parameters are the anchor box parameters corresponding to the text information; and obtain the first target image based on the segmented image dataset and the plurality of adjusted second anchor box parameters.

[0158] In one embodiment, in the step of performing dynamic feature fusion processing on the speech embedding vector and the depth estimation feature to obtain fused features, the processing unit 501 is further configured to: acquire preset semantic information; calculate the semantic similarity between the text information and the preset semantic information to obtain a target semantic similarity; adjust the weights of the speech embedding vector and the depth estimation feature according to the target semantic similarity to obtain a first weight corresponding to the speech embedding vector and a second weight corresponding to the depth estimation feature; and perform dynamic feature fusion processing on the speech embedding vector and the depth estimation feature according to the first weight and the second weight respectively to obtain fused features.

[0159] In one embodiment, when obtaining a first target image based on the segmented image dataset and the adjusted plurality of second anchor box parameters, the processing unit 501 is further configured to: obtain at least one target segmented image corresponding to the text information in the segmented image dataset; superimpose each target segmented image onto the second image based on the adjusted plurality of second anchor box parameters to obtain a superimposed image; and perform image local fusion processing on the superimposed image to obtain the first target image.

[0160] In one embodiment, the first image is an image from a first image database, and the second image is an image from a second image database; the first image database includes multiple real animal images, and the second image database includes multiple images acquired by a vehicle-mounted image acquisition device, a drone image acquisition device, or a gimbal image acquisition device; the second image and the first target image are visible light images or infrared images.

[0161] In one embodiment, the second image and the first target image are infrared images. After obtaining the first target image, the method further includes: performing image filtering on the first target image to obtain a first target image that meets preset conditions as the second target image. In the process of filtering the first target image to obtain a first target image that meets preset conditions as a second target image, the filtering unit 502 is configured to: input the first target image into a preset detection model for detection to obtain the detection result output by the preset detection model; compare the detection result with the text information to obtain a target similarity; when the target similarity is greater than or equal to the preset similarity, determine that the first target image meets the preset conditions and save the first target image that meets the preset conditions as the second target image; or, when the target similarity is less than the preset similarity, delete the first target image.

[0162] In one embodiment, when performing image processing on the first image to obtain a segmented image dataset corresponding to the first image, the processing unit 501 is further configured to: perform target detection on the first image to obtain detection information for each target in the first image; and perform image segmentation processing on the first image based on the detection information for each target to obtain a segmented image dataset corresponding to the first image, wherein the segmented image dataset includes a single segmentation map of each target.

[0163] It should be noted that the image generation apparatus described above can execute the image generation method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of executing the method. Technical details not described in detail in the embodiments of the image generation apparatus can be found in the image generation method provided in the embodiments of this application.

[0164] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the image generation method as described in the foregoing embodiments.

[0165] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0166] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An image generation method, characterized in that, include: Image processing is performed on the first image and the second image respectively to obtain the segmented image dataset corresponding to the first image and the anchor box parameter set corresponding to the second image; Based on the text information, image generation processing is performed using the segmented image dataset and the anchor box parameter set to obtain the first target image.

2. The image generation method according to claim 1, characterized in that, The image processing of the second image to obtain the anchor frame parameter set corresponding to the second image includes: The second image is input into a preset depth estimation model for prediction, and a dense depth map output by the preset depth estimation model is obtained. The depth distribution features of the second image are calculated based on the dense depth map to obtain a depth distribution histogram; According to a preset depth value layering strategy, the depth distribution histogram is divided into distance intervals to obtain multiple different distance intervals; Based on the multiple different distance intervals, the first anchor frame parameters corresponding to each distance interval are obtained; The first anchor frame parameters for each distance interval are adjusted to obtain the adjusted first anchor frame parameters for each distance interval. By summarizing the first anchor frame parameters after adjustment for each distance interval, the set of anchor frame parameters corresponding to the second image is obtained.

3. The image generation method according to claim 2, characterized in that, The process of adjusting the first anchor frame parameters for each distance interval to obtain the adjusted first anchor frame parameters for each distance interval includes: Boundary gradient analysis is performed on the dense depth map to obtain the edge information of each candidate body in the dense depth map; Spatial aggregation feature analysis is performed on the dense depth map to obtain the spatial aggregation features of each candidate body in the dense depth map; Based on the edge information and spatial aggregation features of each candidate body, the first anchor frame parameters of each distance interval are adjusted to obtain the adjusted first anchor frame parameters for each distance interval.

4. The image generation method according to claim 2, characterized in that, The step of performing image generation processing on the segmented image dataset and the anchor box parameter set based on the text information to obtain the first target image includes: The text information is processed by text recognition to obtain the corresponding speech embedding vector; Obtain the dense depth map corresponding to the second image; Image processing is performed on the dense depth map to obtain the corresponding depth estimation features; The speech embedding vector and the depth estimation features are subjected to dynamic feature fusion processing to obtain fused features; The target anchor frame parameters in the anchor frame parameter set are adjusted according to the fusion features to obtain multiple adjusted second anchor frame parameters, wherein the target anchor frame parameters are the anchor frame parameters corresponding to the text information; The first target image is obtained based on the segmented image dataset and the adjusted multiple second anchor box parameters.

5. The image generation method according to claim 4, characterized in that, The dynamic feature fusion process of the speech embedding vector and the depth estimation features to obtain fused features includes: Obtain preset semantic information; Calculate the semantic similarity between the text information and the preset semantic information to obtain the target semantic similarity; Based on the target semantic similarity, the speech embedding vector and the depth estimation feature are weighted and adjusted to obtain the first weight corresponding to the speech embedding vector and the second weight corresponding to the depth estimation feature. The speech embedding vector and the depth estimation features are dynamically fused according to the first weight and the second weight, respectively, to obtain fused features.

6. The image generation method according to claim 4, characterized in that, The step of obtaining the first target image based on the segmented image dataset and the adjusted multiple second anchor box parameters includes: From the segmented image dataset, obtain the target segmented image corresponding to the text information; Based on the adjusted parameters of the multiple second anchor boxes, each of the target segmented images is superimposed onto the second image to obtain a superimposed image; The superimposed images are subjected to local image fusion processing to obtain the first target image.

7. The image generation method according to any one of claims 1 to 6, characterized in that, The first image is an image from a first image database, and the second image is an image from a second image database; the first image database includes multiple real animal images, and the second image database includes multiple images acquired by a vehicle-mounted image acquisition device, a drone image acquisition device, or a gimbal image acquisition device; The second image and the first target image are visible light images or infrared images.

8. The image generation method according to claim 1, characterized in that, The second image and the first target image are infrared images. After obtaining the first target image, the image generation method further includes: performing image filtering on the first target image, and obtaining a first target image that meets preset conditions as the second target image; The step of performing image filtering on the first target image to obtain a first target image that meets preset conditions as the second target image includes: The first target image is input into a preset detection model for detection, and the detection result output by the preset detection model is obtained; The detection results are compared with the text information to obtain the target similarity. When the target similarity is greater than or equal to a preset similarity, the first target image is determined to meet the preset conditions, and the first target image that meets the preset conditions is saved as the second target image; or... When the target similarity is less than the preset similarity, the first target image is deleted.

9. The image generation method according to claim 1, characterized in that, The step of performing image processing on the first image to obtain the segmented image dataset corresponding to the first image includes: Target detection is performed on the first image to obtain detection information for each target in the first image; Based on the detection information of each target, the first image is segmented to obtain a segmented image dataset corresponding to the first image. The segmented image dataset includes a single segmentation map of each target.

10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, cause the processor to perform the image generation method as described in any one of claims 1-9.