Image Processing Method, Apparatus, Electronic Device, and Storage Medium

By extracting the foreground object from the initial image and generating a synthetic image with the depth information of the scene background image, the problem of difficult to label traffic scene images is solved, and efficient downstream model training and detection accuracy are achieved.

CN119477721BActive Publication Date: 2025-07-04ZHEJIANG SUPCON INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510053164.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-07-04
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Traffic scene images are difficult to collect and label, resulting in poor detection results in downstream object detection models. The existing technology relies on high cost and poor accuracy on manual labeling.

Method used

By extracting the foreground object image from the initial image, combining the depth information of the scene background image and the target area, a synthetic image is generated, and instead of manual annotation, accurate label information is provided for downstream model training.

Benefits of technology

It reduces the time-consuming and labor-intensive manual labeling, improves the training accuracy of downstream detection models, and the generated synthetic images have clear identification information, which can effectively improve the detection effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477721B_ABST
    Figure CN119477721B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, apparatus, electronic device and storage medium, relating to the technical field of image processing. The method includes: extracting at least one foreground object image from each acquired initial image; determining the scene depth information corresponding to each scene background image according to each acquired scene background image; determining a target area in each scene background image according to each scene background image; and embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate a plurality of composite images. This method realizes replacing expensive manual annotation data with the way of generating composite images. Since the foreground objects in the composite images have clear identifiers, accurate label information can be provided for the training of downstream detection models. On the premise of reducing the time and effort of manual annotation, the training accuracy of downstream detection models can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and more particularly, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Since it is difficult to collect and annotate images of special traffic scenes, the effect of the trained downstream object detection model in the detection task is not good. Therefore, how to generate high-quality traffic scene content images to provide training data for the training of downstream models has become particularly important.

[0003] Currently, most of the existing traffic scene background images in the acquisition database are obtained through manual annotation to obtain training sample data.

[0004] However, the annotation process is highly dependent on manual labor, with high costs and poor accuracy, thus affecting the detection effect of downstream models. Summary of the Invention

[0005] The purpose of this application is to provide an image processing method, apparatus, electronic device, and storage medium for the above-mentioned deficiencies in the prior art. By synthesizing images to replace manual annotation, it reduces human labor and time consumption and improves the accuracy of downstream model detection.

[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, an embodiment of this application provides an image processing method, including:

[0008] Extracting at least one foreground object image from each of the acquired initial images;

[0009] Determining the scene depth information corresponding to each of the acquired scene background images;

[0010] Determining the target area in each of the scene background images according to the respective scene background images;

[0011] Embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each of the scene background images and the target area in each of the scene background images to generate a plurality of synthetic images.

[0012] Optionally, the extracting at least one foreground object image from each of the acquired initial images includes:

[0013] Collecting a plurality of initial images, each of the initial images containing at least one foreground object;

[0014] According to each initial image and at least one foreground object description information corresponding to each initial image, using a large language vision model, extract each foreground object image corresponding to the foreground object description information from each initial image.

[0015] Optionally, the step of extracting each foreground object image corresponding to the foreground object description information from each initial image according to each initial image and at least one foreground object description information corresponding to each initial image, includes:

[0016] Encode the initial image and the target foreground object description information corresponding to the initial image respectively by the large language vision model to generate fused encoded information;

[0017] Perform segmentation decoding on the fused encoded information, extract a binary mask image of the target foreground object corresponding to the target foreground object description information from the initial image, and perform conversion on the binary mask image to obtain a bounding box of the target foreground object;

[0018] Extract the target foreground object image according to the bounding box of the target foreground object.

[0019] Optionally, the step of determining the scene depth information corresponding to each scene background image according to the collected scene background images, includes:

[0020] Encode the scene background image respectively by using a depth estimation model and a segmentation model to obtain a feature map of the scene background image;

[0021] Decode the feature map of the scene background image to obtain a depth map of the scene background image;

[0022] Determine the depth information of the scene background image according to the depth map.

[0023] Optionally, the step of determining the target area in each scene background image according to the scene background images, includes:

[0024] Generate a road surface mask image in the scene background image according to the scene background image;

[0025] Determine the embeddable area in the scene background image according to the road surface mask image;

[0026] Determine region ratio information according to the boundary information of the embeddable area and the width information of the scene background image; the region ratio information is used to represent the ratio of the target area to be selected in the embeddable area.

[0027] Determine the region segmentation position according to the region ratio information and the boundary information of the embeddable region, and based on the region segmentation position, segment the embeddable region to obtain the target region.

[0028] Optionally, the determining the embeddable region in the scene background image according to the road surface mask image includes:

[0029] According to the road surface mask image, respectively determine a first distance from the road surface to the bottom edge of the scene background image and a second distance from the road surface to the top edge of the scene background image;

[0030] Determine a first boundary line according to the first distance and the boundary of the road surface mask image;

[0031] Determine a second boundary line according to the second distance and the boundary of the road surface mask image;

[0032] According to the first boundary line, the second boundary line and the image region of the road surface mask image, respectively determine a third boundary line and a fourth boundary line;

[0033] The embeddable region is formed by the first boundary line, the second boundary line, the third boundary line and the fourth boundary line.

[0034] Optionally, the determining the region ratio information according to the boundary information of the embeddable region and the width information of the scene background image includes:

[0035] Determine the region ratio information according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line and the width information of the scene background image.

[0036] Optionally, before the determining the region segmentation position according to the region ratio information and the boundary information of the embeddable region, and segmenting the embeddable region based on the region segmentation position to obtain the target region, further includes:

[0037] If the region ratio information is greater than a preset threshold, adjust the region ratio information to obtain the adjusted region ratio information.

[0038] Optionally, the determining the region segmentation position according to the region ratio information and the boundary information of the embeddable region includes:

[0039] Determine the region segmentation position according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line and the region ratio information.

[0040] Optionally, embedding at least one foreground object image into at least one scene background image according to scene depth information corresponding to each scene background image and a target area in each scene background image to generate a plurality of composite images includes:

[0041] Determine a target embedding position from a target area in a target scene background image; the target embedding position is any position in the target area;

[0042] Determining the depth information of the target embedding position according to the scene depth information corresponding to the target scene background image;

[0043] Determine the image size of the target foreground object to be embedded at the target embedding position in the target scene background image according to the depth information of the target embedding position and the mapping relationship between the depth information and the image embedding size;

[0044] According to the image size, the target foreground object image is embedded in the target embedding position in the target scene background image to generate a composite image.

[0045] Optionally, embedding at least one foreground object image into at least one scene background image to generate a plurality of composite images comprises:

[0046] embedding one or more of the foreground object images into a current scene background image to generate at least one synthetic image corresponding to the current scene background image, wherein the current scene background image is any scene background image;

[0047] The multiple composite images are obtained according to at least one composite image corresponding to each scene background image.

[0048] Optionally, embedding at least one foreground object image into at least one scene background image to generate a plurality of composite images comprises:

[0049] If the number of foreground object images currently embedded in the target scene background image does not exceed the preset upper limit of the embedding number, the foreground object image to be embedded is embedded into the target scene background image according to the embedding position of each currently embedded foreground object image to generate a current composite image.

[0050] Optionally, after generating a plurality of composite images, the method further includes:

[0051] The pre-trained image harmonization model is used to repair the inharmonious areas in each composite image to obtain an optimized composite image.

[0052] Optionally, the training process of the image harmonization model includes:

[0053] Collect a training sample data set, where the training sample data set includes: multiple groups of sample data, and each group of sample data includes: a real captured image and a synthetic image constructed based on the real captured image;

[0054] Use the training sample data set to train and obtain the image harmonization model.

[0055] In a second aspect, an embodiment of the present application further provides an image processing device, including: an acquisition module, a determination module, and an embedding module;

[0056] The acquisition module is used to extract at least one foreground object image from each acquired initial image;

[0057] The determination module is used to determine the scene depth information corresponding to each scene background image according to each acquired scene background image;

[0058] The determination module is used to determine the target area in each scene background image according to each scene background image;

[0059] The embedding module is used to embed at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image, and generate multiple synthetic images.

[0060] Optionally, the acquisition module is specifically used to collect multiple initial images, and each initial image contains at least one foreground object;

[0061] According to each initial image and the at least one foreground object description information corresponding to each initial image, use a large language vision model to extract each foreground object image corresponding to each foreground object description information from each initial image.

[0062] Optionally, the acquisition module is specifically used to encode the initial image and the target foreground object description information corresponding to the initial image respectively by the large language vision model to generate fused encoding information;

[0063] Perform segmentation decoding on the fused encoding information, extract a binary mask image of the target foreground object corresponding to the target foreground object description information from the initial image, and perform conversion on the binary mask image to obtain the bounding box of the target foreground object;

[0064] Extract the target foreground object image according to the bounding box of the target foreground object.

[0065] Optionally, the determination module is specifically used to encode the scene background image by using a depth estimation model and a segmentation model respectively to obtain a feature map of the scene background image;

[0066] Decode the feature map of the scene background image to obtain the depth map of the scene background image;

[0067] Determine the depth information of the scene background image according to the depth map.

[0068] Optionally, the determining module is specifically configured to generate a road surface mask image in the scene background image according to the scene background image;

[0069] Determine the embeddable area in the scene background image according to the road surface mask image;

[0070] Determine the area ratio information according to the boundary information of the embeddable area and the width information of the scene background image; the area ratio information is used to characterize the ratio of the to-be-selected target area to the embeddable area;

[0071] Determine the area segmentation position according to the area ratio information and the boundary information of the embeddable area, and segment the embeddable area based on the area segmentation position to obtain the target area.

[0072] Optionally, the determining module is specifically configured to respectively determine a first distance between the road surface and the bottom edge of the scene background image and a second distance between the road surface and the top edge of the scene background image according to the road surface mask image;

[0073] Determine a first boundary line according to the first distance and the boundary of the road surface mask image;

[0074] Determine a second boundary line according to the second distance and the boundary of the road surface mask image;

[0075] Determine a third boundary line and a fourth boundary line respectively according to the first boundary line, the second boundary line and the image area of the road surface mask image;

[0076] The embeddable area is formed by the first boundary line, the second boundary line, the third boundary line and the fourth boundary line.

[0077] Optionally, the determining module is specifically configured to determine the area ratio information according to the endpoint coordinate information of the first boundary line of the embeddable area, the endpoint coordinate information of the second boundary line and the width information of the scene background image.

[0078] Optionally, it further includes an adjustment module;

[0079] The adjustment module is used to adjust the area ratio information if the area ratio information is greater than a preset threshold to obtain the adjusted area ratio information.

[0080] Optionally, the determination module is specifically configured to determine the region segmentation position according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line, and the region ratio information.

[0081] Optionally, the embedding module is specifically used to determine the target embedding position from the target area in the target scene background image; the target embedding position is any position in the target area;

[0082] Determining the depth information of the target embedding position according to the scene depth information corresponding to the target scene background image;

[0083] Determine the image size of the target foreground object to be embedded at the target embedding position in the target scene background image according to the depth information of the target embedding position and the mapping relationship between the depth information and the image embedding size;

[0084] According to the image size, the target foreground object image is embedded in the target embedding position in the target scene background image to generate a composite image.

[0085] Optionally, the embedding module is specifically used to embed one or more of the foreground object images into a current scene background image to generate at least one synthetic image corresponding to the current scene background image, where the current scene background image is any scene background image;

[0086] The multiple composite images are obtained according to at least one composite image corresponding to each scene background image.

[0087] Optionally, the embedding module is specifically used to embed the foreground object image to be embedded into the target scene background image according to the embedding position of each foreground object image currently embedded, if the number of foreground object images currently embedded in the target scene background image does not exceed a preset embedding number upper limit, so as to generate a current composite image.

[0088] Optionally, the adjustment module is further used to use a pre-trained image harmonization model to repair the inharmonious areas in each composite image to obtain an optimized composite image.

[0089] Optionally, it also includes: a training module;

[0090] The training module is used to collect a training sample data set, wherein the training sample data set includes: a plurality of groups of sample data, each group of sample data includes: a real shot image and a synthetic image constructed based on the real shot image;

[0091] The image harmonization model is obtained by training using the training sample data set.

[0092] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to implement the image processing method provided in the first aspect.

[0093] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the image processing method provided in the first aspect.

[0094] The beneficial effects of the present application are as follows:

[0095] The present application provides an image processing method, device, electronic device, and storage medium, including: extracting at least one foreground object image from each acquired initial image; determining the scene depth information corresponding to each scene background image according to each acquired scene background image; determining the target area in each scene background image according to each scene background image; embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate a plurality of synthetic images. Based on the foreground object images extracted from the initial images, combined with the depth information of the scene background images and the target areas of the scene background images, this method can determine the image size and embedding position of the foreground object images when embedding them into the scene background images, so as to embed the foreground object images into the scene background images according to reasonable sizes and positions to obtain synthetic images in a traffic scene. It realizes replacing high-cost manual annotation data with the method of generating synthetic images. Since the foreground objects in the synthetic images have clear identifiers, accurate label information can be provided for the training of downstream detection models. On the premise of reducing the time and effort-consuming of manual annotation, the training accuracy of downstream detection models can also be improved.

[0096] In addition, extracting foreground object images from initial images based on description information can improve the accuracy of the extraction results of foreground object images and facilitate the generation of annotation information for foreground object images. Description of the Drawings

[0097] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0098] Figure 1Schematic diagram of the architecture of an image processing system provided by an embodiment of the present application;

[0099] Figure 2 Schematic diagram of the process of an image processing method provided by an embodiment of the present application;

[0100] Figure 3 Schematic diagram of the process of another image processing method provided by an embodiment of the present application;

[0101] Figure 4 Schematic diagram of the generation of a foreground object binary map mask provided by an embodiment of the present application;

[0102] Figure 5 Schematic diagram of the process of yet another image processing method provided by an embodiment of the present application;

[0103] Figure 6 Schematic diagram of the process of another image processing method provided by an embodiment of the present application;

[0104] Figure 7 Schematic diagram of the process of another image processing method provided by an embodiment of the present application;

[0105] Figure 8 Schematic diagram of the process of yet another image processing method provided by an embodiment of the present application;

[0106] Figure 9 Schematic diagram of the road mask of a scene background image provided by an embodiment of the present application;

[0107] Figure 10 Schematic diagram of the embeddable area of a scene background image provided by an embodiment of the present application;

[0108] Figure 11 Schematic diagram of the process of another image processing method provided by an embodiment of the present application;

[0109] Figure 12 Schematic diagram of the process of yet another image processing method provided by an embodiment of the present application;

[0110] Figure 13 Schematic diagram of an image processing device provided by an embodiment of the present application;

[0111] Figure 14 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0112] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application are only for the purposes of illustration and description, and are not used to limit the protection scope of this application. Additionally, it should be understood that the schematic drawings are not drawn to actual scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. Moreover, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.

[0113] Furthermore, the described embodiments are only some embodiments of this application, rather than all of the embodiments. The components of the embodiments of this application described and illustrated in the accompanying drawings here can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application that is claimed, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the protection scope of this application.

[0114] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.

[0115] Since it is difficult to collect and annotate images in special traffic scenarios, the performance quality of existing models in downstream detection tasks is generally not high. Most current solutions in the market train expert models using the method of multi-person annotation, and then use the expert models to perform inference on large-scale datasets, and obtain high-quality training datasets by manually screening the inference results. There are also some solutions that directly adopt image fusion and generation technologies to increase the number of training samples.

[0116] Existing image fusion technologies introduce an instruction-based object addition pipeline, which can automatically insert objects into real-world scenarios in a reasonable size and position. It includes two parts: creating text instructions and fine-tuning the diffusion model to achieve reasonable generation. Specifically, first, an instruction-based image pair is created by using object removal technology, and a large background dataset is constructed. Then, an instruction-based stable diffusion model is trained on this dataset to learn diverse and reasonable object generation, avoiding relying on costly manual annotations such as bounding boxes.

[0117] The disadvantages of the prior art are as follows: 1. For extracting special categories of objects, ordinary text instructions will cause the object detection boxes generated by the vision-language model to be inaccurate, thus directly affecting the effect of image generation; 2. Usually, the elements in traffic scenes are relatively complex, and there is more than one foreground element generated in a single image. The perspective relationship between elements and between elements and the background is particularly important. Existing solutions for generating images by fine-tuning the diffusion model often ignore physical laws and are difficult to adapt to multi-scene depth modeling; 3. During the image synthesis process, the adaptation of lighting, shadows, perspective, and background is crucial for the final synthesis effect. In traffic scenes, the objects to be detected are usually small objects. Therefore, traditional synthesis methods such as Poisson fusion and Gaussian fusion, as well as deep learning methods such as TF-ICON (Training-Free Image COmpositioN, a cross-domain image synthesis framework based on the diffusion model) and Blended (Blended Learning) rely too much on image gradients, resulting in the loss of the original edge and color information of small objects after they are incorporated into dark backgrounds such as highway pavements.

[0118] Based on this, this solution provides an image processing method. By extracting the foreground object image and the depth information of the scene background image, and based on the road surface mask image of the scene background image, the target area in the scene background image is determined, so as to embed the extracted foreground object image into the target area in the scene background image. Among them, by combining the depth information of the scene background image and the specific embedding position of the foreground object image in the target area, the image size of the foreground object image during embedding can be determined, so that the foreground object image can be embedded into the scene background image according to a reasonable size and position to obtain a synthetic image in a traffic scene. It realizes replacing expensive manual annotation data with the generated synthetic image. Since the foreground object in the synthetic image has a clear identifier, accurate label information can be provided for the training of downstream detection models. On the premise of reducing the time and effort consumed by manual annotation, the training accuracy of downstream detection models can also be improved.

[0119] Figure 1 It is a schematic diagram of the architecture of an image processing system provided by an embodiment of this application; as Figure 1As shown, based on the crawled initial images containing foreground objects, the initial images and the description information of the foreground objects to be extracted are input into the large language vision model. The large language vision model performs image encoding on the initial images and text encoding on the description information to obtain a fused encoding, and then through segmentation decoding of the fused encoding, the foreground object images can be obtained. Similarly, processing the collected scene background images based on the large language vision model can obtain the road surface mask images in the scene background images. In addition, processing the scene background images based on the depth estimation model can extract the depth images corresponding to the scene background images, and based on the depth images, the scene depth information of the scene background images can be obtained. Thus, by combining the foreground object images, the road surface mask images in the scene background images, and the scene depth information, image embedding processing is performed to embed the foreground object images into reasonable positions in the scene background images, obtaining synthetic images labeled with foreground object identifiers. The synthetic images can be used as training data to train downstream detection models. Since the foreground objects are known objects with clear identifiers, the synthetic images have accurate identifier information, which can improve the accuracy of the trained models.

[0120] Figure 2 It is a schematic flowchart of an image processing method provided by an embodiment of the present application; the execution subject of this method is a computer device such as a terminal, a server, etc., as Figure 2 shown, this method may include:

[0121] S101. Extract at least one foreground object image from each of the obtained initial images.

[0122] The initial images can be images crawled from websites that contain foreground objects to be detected. In a traffic scenario, the foreground objects to be detected can be objects of special categories that are not easily detected accurately. Suppose the foreground object to be detected is a "roadside pier", then the crawled initial images are images containing various types of "roadside piers" in various scenarios.

[0123] Of course, an initial image does not necessarily contain only one foreground object to be detected. Usually, the number and type of foreground objects contained in the initial image are not limited, and can be one or multiple.

[0124] Based on the crawled initial images, at least one foreground object image can be extracted from the initial images. Among them, when the initial images contain multiple different foreground objects, at least one foreground object image can be extracted from the initial images according to requirements.

[0125] S102. Determine the scene depth information corresponding to each of the collected scene background images.

[0126] The scene background image can also be collected in advance, which can be collected from the historical images stored in the camera. In the traffic scene, the scene background image can be an image containing the road surface scene.

[0127] Based on the collected scene background image, the scene depth distance relationship map corresponding to the scene background image can be estimated, which can also be called the scene depth map.

[0128] The scene depth map is a special form of image representation, which records the distance information between each point in the scene and the observer (camera). In the depth map, the grayscale value or color value of each pixel represents the distance from the corresponding point in the three-dimensional space to the camera. This representation method enables the depth information to be visually presented in the form of an image. Therefore, the scene depth information can be obtained based on the scene depth map.

[0129] S103. Determine the target area in each scene background image according to each scene background image.

[0130] For the scene background image, the road surface mask image in the scene background image can also be extracted, and then the maximum bounding box of the road surface can be determined based on the road surface mask image. The maximum bounding box defines the boundary when the foreground object image is embedded in the scene background image, and the foreground object image is allowed to be embedded within the maximum bounding box.

[0131] However, in order to further improve the rationality and accuracy of the embedding position, the target area can also be determined within the maximum bounding box through a preset algorithm, and the target area is the reasonable area where the foreground object image is allowed to be embedded.

[0132] S104. Embed at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image, and generate multiple composite images.

[0133] In some embodiments, according to the target area in the scene background image, the embedding position of the foreground object image can be determined. According to the scene depth information corresponding to the scene background image, based on the principle of objects being larger when closer and smaller when farther away, the embedding size of the foreground object image can be determined, and then the foreground object image can be embedded into the embedding position in the scene background image according to the embedding size to generate a composite image.

[0134] Since multiple foreground object images can be extracted and multiple scene background images are also included, one or more foreground object images can be embedded in one scene background image, and the same foreground object image can also be embedded into different scene background images, so that multiple composite images can be combined and generated, and the scene background or foreground object in each composite image is not completely the same.

[0135] Optionally, the multiple generated synthetic images can be used as sample data for training a downstream detection model. When generating the synthetic images, the foreground objects embedded in the scene background images have clear and accurate identifications, so that the synthetic images have accurate annotation information and do not need to be manually annotated. The synthetic images can be directly used to train the model, and the trained model has a good detection effect.

[0136] In summary, the image processing method provided in this embodiment includes: extracting at least one foreground object image from each acquired initial image; determining the scene depth information corresponding to each scene background image according to each acquired scene background image; determining the target area in each scene background image according to each scene background image; and embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate multiple synthetic images. Based on the foreground object images extracted from the initial images, combined with the depth information of the scene background images and the target areas of the scene background images, this method can determine the image size and embedding position of the foreground object images when embedding them into the scene background images, so as to embed the foreground object images into the scene background images according to reasonable sizes and positions to obtain synthetic images in the traffic scene. It realizes replacing expensive manual annotation data with the method of generating synthetic images. Since the foreground objects in the synthetic images have clear identifications, accurate label information can be provided for the training of the downstream detection model, and the training accuracy of the downstream detection model can be improved while reducing the time and effort of manual annotation.

[0137] Figure 3 It is a schematic flowchart of another image processing method provided by an embodiment of the present application; optionally, referring to Figure 3 As shown, in step S101, extracting at least one foreground object image from each acquired initial image may include:

[0138] S201. Collect a plurality of initial images, and each initial image contains at least one foreground object.

[0139] In some embodiments, a plurality of initial images can be collected from websites. Based on the above description, each initial image may include one or more foreground objects to be detected.

[0140] S202. According to each initial image and at least one foreground object description information corresponding to each initial image, use a large language vision model to extract each foreground object image corresponding to each foreground object description information from each initial image.

[0141] According to requirements, at least one foreground object can be extracted from each initial image. That is, one or more foreground objects can be extracted from one initial image.

[0142] According to the requirements, it is possible to determine which foreground object is desired to be extracted from the initial image, so that the description information of the foreground object to be extracted can be generated. When the description information of the foreground object and the initial image are input into the large language vision model together, the foreground object image described by the description information of the foreground object can be extracted from the initial image.

[0143] Exemplarily, assume that the foreground objects included in the initial image are: a trash box and a puppy; if it is desired to extract the trash box in the initial image according to the requirements, the description information of the foreground object "trash box" and the initial image can be input into the large language vision model together, and the large language vision model extracts the "trash box" image from the initial image based on the semantics described by the description information.

[0144] Figure 4 It is a schematic diagram for generating a binary mask map of a foreground object provided by an embodiment of the present application. Using the description information of the foreground object and the initial image as the input of the large language vision model, and through the processing of the large language vision model, the binary mask map of the foreground object described by the description information of the foreground object can be extracted.

[0145] Figure 5 It is a schematic flowchart of another image processing method provided by an embodiment of the present application; referring to Figure 5 As shown, optionally, in step S202, according to each initial image and at least one piece of description information of the foreground object corresponding to each initial image, using the large language vision model to extract each foreground object image corresponding to each piece of description information of the foreground object from each initial image may include:

[0146] S301. The large language vision model encodes the initial image and the target foreground object description information corresponding to the initial image respectively to generate fused encoding information.

[0147] In a realizable manner, the large language vision model can perform image encoding on the initial image and text encoding on the target foreground object description information to obtain fused encoding information.

[0148] Among them, the target foreground object description information describes the target foreground object in the initial image, and the target foreground object can be any foreground object in the initial image.

[0149] S302. Perform segmentation decoding on the fused encoding information, extract the binary mask map of the target foreground object corresponding to the target foreground object description information from the initial image, and perform conversion on the binary mask map to obtain the bounding box of the target foreground object.

[0150] By splitting and decoding the fused coding information, a binary mask image of the target foreground object described by the target foreground object description information can be extracted from the initial image. By converting the binary mask image into a circumscribed rectangle, the bounding box of the target foreground object can be obtained.

[0151] S303. Extract the target foreground object image according to the bounding box of the target foreground object.

[0152] Based on the bounding box of the target object, the image within the bounding box can be extracted to obtain the target foreground object image.

[0153] In another implementable way, image processing techniques such as convolutional neural network (CNN) can be used to extract features from the input initial image. The extracted features include low-level features such as the color, texture, and shape of the image, as well as more advanced semantic features.

[0154] Then, parse the input target foreground object description information to understand its semantic content. This usually involves natural language processing (NLP) techniques such as word segmentation, part-of-speech tagging, and syntactic analysis.

[0155] Match the parsed description information with the extracted image features. By comparing the object features in the description information with the features in the image, the target foreground object image that matches the target foreground object description information is found.

[0156] Figure 6 It is a schematic flowchart of another image processing method provided by the embodiments of the present application; referring to Figure 6 As shown, optionally, in step S102, according to the collected background images of each scene, determining the scene depth information corresponding to each background image of the scene may include:

[0157] S401. Respectively use a depth estimation model and a segmentation model to encode the background image of the scene to obtain a feature map of the background image of the scene.

[0158] In some embodiments, for the background image of the scene, the background image of the scene can be used as input, and after being encoded by the depth estimation model Depth-Anything and the segmentation model DINOV2, a feature map is obtained.

[0159] S402. Decode the feature map of the background image of the scene to obtain the depth map of the background image of the scene.

[0160] Then input Depth-Anything to decode the above-obtained feature map to obtain the depth map of the background image of the scene. The purpose is to use the discrete information of semantic segmentation to guide the depth estimation model to learn the depth representations of different objects, thereby enhancing the accuracy of monocular depth estimation.

[0161] S403. Determine the depth information of the scene background image based on the depth map.

[0162] Since each pixel point in the depth map contains depth information and the depth map is an image representation of depth information, the depth information of the scene background image can be determined based on the depth map. The depth information of the scene background image includes the depth information of each pixel point in the scene background image, and this depth information indicates the distance of the pixel point from the camera emission source.

[0163] Figure 7 It is a schematic flowchart of another image processing method provided by an embodiment of the present application; refer to Figure 7 As shown, optionally, in step S103, to determine the target area in each scene background image according to each scene background image, it may include:

[0164] S501. Generate a road surface mask image in the scene background image according to the scene background image.

[0165] In some embodiments, the Grounding-DINO (advanced open detection model) model can be used to process the scene background image to generate a road surface mask image in the scene background image.

[0166] Alternatively, referring to the above extraction method of the foreground object image, a road surface mask image can be extracted from the scene background image based on the large language vision model.

[0167] S502. Determine the embeddable area in the scene background image according to the road surface mask image.

[0168] According to the road surface mask image, the embeddable area in the scene background image can be calculated and generated. The embeddable area can indicate the maximum inscribed bounding box of the road surface mask image. The embeddable area can be used as the effective area of the road surface. In this area, foreground objects can be embedded, while the remaining area outside the embeddable area in the scene background image can be excluded. The purpose of determining the embeddable area is to exclude non-effective areas such as the road surface edge.

[0169] S503. Determine the area ratio information according to the boundary information of the embeddable area and the width information of the scene background image.

[0170] Among them, the area ratio information is used to characterize the proportion of the to-be-selected target area in the embeddable area.

[0171] Based on the determined boundary information of the embeddable area and the width information of the scene background image, the area ratio information (also referred to as ROI ratio ) can be calculated.

[0172] The boundary information here may refer to the information of the bounding box constituting the embeddable region, and the width information of the scene background image may refer to the side length of the scene background image in the vertical direction.

[0173] S504. Determine the region segmentation position according to the region ratio information and the boundary information of the embeddable region, and segment the embeddable region based on the region segmentation position to obtain the target region.

[0174] Next, according to the calculated region ratio information combined with the boundary information of the embeddable region, the region segmentation position can be determined. Among them, the coordinates of the segmentation point during region segmentation can be determined first, and then a horizontal segmentation line is made based on the segmentation point. Based on the horizontal segmentation line, the embeddable region is segmented into two parts. Since objects that are too far from the camera do not need to be detected, the foreground object does not need to be embedded in a far region during embedding. Therefore, the region closer to the camera among the two segmented parts can be used as the target region, that is, the part closer to the bottom edge of the scene background image is used as the target region.

[0175] Figure 8 It is a schematic flowchart of another image processing method provided by an embodiment of the present application; Figure 9 It is a schematic diagram of a road surface mask of a scene background image provided by an embodiment of the present application, Figure 10 It is a schematic diagram of an embeddable region of a scene background image provided by an embodiment of the present application. Figure 10 The first boundary line, the second boundary line, the third boundary line, and the fourth boundary line are respectively marked. Refer to Figures 8 - 10 As shown, optionally, in step S502, determining the embeddable region in the scene background image according to the road surface mask image may include:

[0176] S601. According to the road surface mask image, respectively determine the first distance between the road surface close to the bottom edge of the scene background image and the second distance between the road surface close to the top edge of the scene background image.

[0177] According to the road surface mask image, on the basis of removing the non-road surface information in the mask image, the nearest first distance between the road surface close to the bottom edge of the scene background image and the second distance between the road surface closest to the top edge of the scene background image can be determined.

[0178] Taking Figure 9 the road surface mask image shown as an example, since there is text information in the lower right corner of the mask image, the text information needs to be removed from the mask image. After removing the text information, the determined first distance can be the distance y1 marked in Figure 10 . Of course, if there is no non-road surface information in the lower right corner of the road surface mask image, the ordinate of the bottom edge boundary line of the road surface mask image can be directly determined as the first distance.

[0179] The second distance is determined according to the outermost boundary line of the road surface mask image, and the second distance can be, for example, Figure 10 the distance y2 marked in. Wherein, the coordinate system of the image takes the lower left corner of the image as the origin, the direction of the length of the image is the x-axis, and the direction of the width of the image is the y-axis.

[0180] S602. Determine the first boundary line according to the first distance and the boundary of the road surface mask image.

[0181] According to the first distance, make a parallel line to the upper and lower sides of the scene background image, and the length of the parallel line needs to be inscribed within the bottom boundary of the road surface mask image, that is, make a horizontal line that inscribes the bottom boundary line of the road surface mask image to obtain the first boundary line.

[0182] S603. Determine the second boundary line according to the second distance and the boundary of the road surface mask image.

[0183] Similarly, according to the second distance, make a horizontal line that inscribes the top boundary line of the road surface mask image to obtain the second boundary line.

[0184] S604. Determine the third boundary line and the fourth boundary line respectively according to the first boundary line, the second boundary line and the image area of the road surface mask image.

[0185] Emit a ray from the left end point of the second boundary line to the first boundary line, and determine the maximum ray that can cover the left boundary line of the road surface mask image as the third boundary line. Similarly, emit a ray from the right end point of the second boundary line to the first boundary line, and determine the maximum ray that can cover the right boundary line of the road surface mask image as the fourth boundary line.

[0186] S605. The first boundary line, the second boundary line, the third boundary line and the fourth boundary line form an embeddable area.

[0187] Connect the first boundary line, the second boundary line, the third boundary line and the end points of the third boundary line in sequence to form a closed figure, then the embeddable area can be obtained.

[0188] Figure 10 The embeddable area in Figure 9 is determined based on the road surface mask image in

[0189] Optionally, in step S503, determining the region ratio information according to the boundary information of the embeddable region and the width information of the scene background image may include: determining the region ratio information according to the endpoint coordinate information of the first boundary line and the second boundary line of the embeddable region and the width information of the scene background image.

[0190] Optionally, the region ratio information may be calculated by using a pre-designed calculation formula according to the abscissa of the left endpoint of the first boundary line, the abscissa of the left endpoint of the second boundary line, and the width information of the scene background image.

[0191] Continuing to refer to Figure 10 , assuming that the abscissa of the left endpoint of the first boundary line is , the abscissa of the left endpoint of the second boundary line is , and the width information of the scene background image is , then the region ratio information ROI can be calculated by using the following formula ratio :

[0192]

[0193] Optionally, in step S504, before determining the region segmentation position according to the region ratio information and the boundary information of the embeddable region and segmenting the embeddable region based on the region segmentation position to obtain the target region, it further includes: if the region ratio information is greater than a preset threshold, adjusting the region ratio information to obtain the adjusted region ratio information.

[0194] In some embodiments, when the endpoint of the second boundary line of the generated embeddable region is close to the left edge or the right edge, Figure 10 the embeddable region shown in ratio will change from a trapezoid to a shape closer to a rectangle. This will cause the ROI

[0195] calculated by using the above formula to become larger, resulting in a smaller target region finally calculated.

[0196]

[0197] That is, if it is determined that the calculated ROI ratio exceeds 0.5, then perform the operation (1 - ROI ratio ) to effectively compensate for the problem that the target region gradually shrinks as the ROI ratio increases.

[0198] Optionally, in step S504, determining the region segmentation position according to the region ratio information and the boundary information of the embeddable region may include: determining the region segmentation position according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line, and the region ratio information.

[0199] Next, according to the ordinate of the left endpoint of the first boundary line of the embeddable region and the ordinate of the left endpoint of the second boundary line, combined with the adjusted region ratio information ROI ratio , the ordinate of the region segmentation position point can be calculated.

[0200] Assume that the ordinate of the left endpoint of the first boundary line is , and the ordinate of the left endpoint of the second boundary line is , then the ordinate of the region segmentation position point can be calculated using the following formula:

[0201]

[0202] Based on the ordinate of the region segmentation position point, a horizontal line can be drawn to generate a region segmentation line. Continuing to refer to Figure 10 , assume that the region segmentation line is ROI top , then, the region formed by ROI top and the first boundary line, the third boundary line, and the third boundary line of the embeddable region is the target region. The target region is as shown in Figure 10 the shaded area in.

[0203] Figure 11 is a schematic flowchart of another image processing method provided by an embodiment of the present application; referring to Figure 11 shown, optionally, in step S104, embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target region in each scene background image to generate a plurality of composite images may include:

[0204] S701. Determine a target embedding position from the target region in the target scene background image; the target embedding position is any position in the target region.

[0205] It should be noted that any position in the determined target region can be used as the target embedding position, that is to say, the foreground object image can be embedded at any position in the target region.

[0206] S702. Determine the depth information of the target embedding position according to the scene depth information corresponding to the target scene background image.

[0207] Based on the determined target embedding position, the depth information of the target embedding position can be determined according to the scene depth information.

[0208] S703. Determine the size of the target foreground object image at the target embedding position in the target scene background image to be embedded according to the depth information of the target embedding position and the mapping relationship between the depth information and the image embedding size.

[0209] Optionally, based on the display rule of objects getting larger when closer and smaller when farther away from the camera, that is, the closer an object is to the camera, the larger its display size, and the farther an object is from the camera, the smaller its display size. The mapping relationship between the depth information and the image embedding size can be pre-constructed.

[0210] Then, based on the depth information of the target embedding position, according to the mapping relationship between the depth information and the image embedding size, the size that the foreground object image should have when being embedded in the target embedding position of the scene background image can be determined.

[0211] S704. Embed the target foreground object image at the target embedding position in the target scene background image according to the image size to generate a composite image.

[0212] Based on the determined image size, the target foreground object image can be embedded at the target embedding position with this image size to generate a composite image.

[0213] Among them, the target foreground object image can be any one of the extracted foreground object images, and the target scene background image can be any one of the scene background images.

[0214] In some embodiments, in order to increase the number of generated composite images, each foreground object image can also be rotated according to a preset rotation angle to expand and obtain more foreground object images.

[0215] Thus, the expanded foreground object images can also be embedded in the scene background image to generate a composite image.

[0216] Figure 12 It is a schematic flowchart of another image processing method provided by the embodiments of this application; optionally, in step S104, embedding at least one foreground object image in at least one scene background image to generate multiple composite images may include:

[0217] S801. Embed one or more of each foreground object image in the current scene background image to generate at least one composite image corresponding to the current scene background image, where the current scene background image is any one of the scene background images.

[0218] In one feasible manner, for a scene background image, one or more of all the extracted foreground object images may be used as embedded images and embedded into the scene background image to obtain at least one composite image corresponding to the scene background image.

[0219] That is, a foreground object image can be embedded in one scene background image or in different scene background images, and a scene background image can be embedded with only one foreground object image or with multiple identical or different foreground object images, as long as the final synthesized images are not completely identical.

[0220] S802: Obtain multiple composite images according to at least one composite image corresponding to each scene background image.

[0221] Each scene background image corresponds to at least one synthetic image, thereby forming a plurality of synthetic images.

[0222] Optionally, in step S104, at least one foreground object image is embedded in at least one scene background image to generate multiple composite images, which may include: if the number of foreground object images currently embedded in the target scene background image does not exceed a preset upper limit on the number of embeddings, then according to the embedding positions of the currently embedded foreground object images, the foreground object image to be embedded is embedded in the target scene background image to generate a current composite image.

[0223] If multiple foreground object images need to be embedded in a scene background image, in order to avoid overlap and occlusion between the embedded foreground object images, an upper limit on the number of foreground object images that can be embedded in the target area can be set. Each time the current foreground object image is embedded, the number of foreground object images already embedded in the scene background image can be first determined. If the upper limit is not reached, the current foreground object image can be embedded, and the current foreground object image needs to be embedded in a new non-overlapping position according to the embedding positions of the embedded foreground object images. If the upper limit is reached, the current foreground object image cannot be embedded.

[0224] Optionally, in step S104, after generating a plurality of composite images, the method further includes: using a pre-trained image harmonization model to repair inharmonious regions in each composite image to obtain an optimized composite image.

[0225] In some embodiments, the composite image obtained by the above method may have defects, such as the foreground and background look disharmonious, the interlacing is unnatural, etc. To solve this problem, a pre-trained image harmonization model can be used to repair the composite image to adjust the foreground object or background in the composite image to make the composite image more harmonious.

[0226] Optionally, the training process of the image harmonization model may include: collecting a training sample data set, where the training sample data set includes: multiple groups of sample data, and each group of sample data includes: a real captured image and a synthetic image constructed based on the real captured image; using the training sample data set to train an image harmonization model.

[0227] In some embodiments, multiple real images can be collected from a database, and based on the real images, synthetic images corresponding to the real images can be constructed. Thus, a real image and a synthetic image can form a group of sample data, and an image harmonization model can be trained from multiple groups of sample data.

[0228] Among them, in the process of training the image harmonization model, the concept of domain verification was attempted, and DoveNet (DOmain VErification Network, a deep learning model for image harmonization) based on domain verification was used. Specifically, we believe that the foreground object and the background of the synthetic image are respectively collected in different shooting environments, and each shooting environment can be regarded as a domain. There may be countless domains in the actual situation, and we do not know the domain labels of the foreground object and the background, and can only migrate the foreground object to the same domain as the background.

[0229] In summary, the image processing method provided in this embodiment includes: extracting at least one foreground object image from each obtained initial image; determining the scene depth information corresponding to each scene background image according to the collected scene background images; determining the target area in each scene background image according to the scene background images; embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate multiple synthetic images. Based on the foreground object images extracted from the initial images, combined with the depth information of the scene background images and the target areas of the scene background images, this method can determine the image size and embedding position of the foreground object images when embedding them into the scene background images, so as to embed the foreground object images into the scene background images according to reasonable sizes and positions to obtain synthetic images in the traffic scene. It realizes replacing expensive manually labeled data with the method of generating synthetic images. Since the foreground objects in the synthetic images have clear identifiers, accurate label information can be provided for the training of downstream detection models. On the premise of reducing the time and effort consumed by manual annotation, the training accuracy of downstream detection models can also be improved.

[0230] In addition, extracting the foreground object image from the initial image based on the description information can improve the accuracy of the foreground object image extraction result and facilitate the generation of the annotation information of the foreground object image.

[0231] The following describes the apparatus, device, storage medium, etc. for implementing the image processing method provided in this application. For the specific implementation process and technical effects, please refer to the above, and will not be elaborated below.

[0232] Figure 13 The figure is a schematic diagram of an image processing apparatus provided in an embodiment of this application. The functions implemented by this image processing apparatus correspond to the steps executed by the above method. This apparatus can be understood as the above-mentioned terminal, server, or the processor of the server, or can also be understood as a component independent of the above-mentioned server or processor that implements the functions of this application under the control of the server, such as Figure 13 As shown, the apparatus includes: an acquisition module 130, a determination module 131, and an embedding module 132;

[0233] The acquisition module 130 is configured to extract at least one foreground object image from each acquired initial image;

[0234] The determination module 131 is configured to determine the scene depth information corresponding to each scene background image according to each acquired scene background image;

[0235] The determination module 131 is configured to determine the target area in each scene background image according to each scene background image;

[0236] The embedding module 132 is configured to embed at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image, and generate a plurality of composite images.

[0237] Optionally, the acquisition module 130 is specifically configured to collect a plurality of initial images, and each initial image contains at least one foreground object;

[0238] According to each initial image and the at least one foreground object description information corresponding to each initial image, using a large language vision model, extract each foreground object image corresponding to the foreground object description information from each initial image.

[0239] Optionally, the acquisition module 130 is specifically configured to encode the initial image and the target foreground object description information corresponding to the initial image respectively by the large language vision model to generate fusion encoding information;

[0240] Segment and decode the fusion encoding information, extract the binary mask image of the target foreground object corresponding to the target foreground object description information from the initial image, and perform conversion on the binary mask image to obtain the bounding box of the target foreground object;

[0241] Extract the target foreground object image according to the bounding box of the target foreground object.

[0242] Optionally, the determining module 131 is specifically configured to encode the scene background image by using a depth estimation model and a segmentation model respectively, so as to obtain a feature map of the scene background image;

[0243] Decode the feature map of the scene background image to obtain a depth map of the scene background image;

[0244] Determine the depth information of the scene background image according to the depth map.

[0245] Optionally, the determining module 131 is specifically configured to generate a road surface mask image in the scene background image according to the scene background image;

[0246] Determine an embeddable area in the scene background image according to the road surface mask image;

[0247] Determine region ratio information according to the boundary information of the embeddable area and the width information of the scene background image; the region ratio information is used to represent the ratio of the to-be-selected target area to the embeddable area;

[0248] Determine a region segmentation position according to the region ratio information and the boundary information of the embeddable area, and segment the embeddable area based on the region segmentation position to obtain a target area.

[0249] Optionally, the determining module 131 is specifically configured to respectively determine a first distance between the road surface and the bottom edge of the scene background image and a second distance between the road surface and the top edge of the scene background image according to the road surface mask image;

[0250] Determine a first boundary line according to the first distance and the boundary of the road surface mask image;

[0251] Determine a second boundary line according to the second distance and the boundary of the road surface mask image;

[0252] Determine a third boundary line and a fourth boundary line respectively according to the first boundary line, the second boundary line and the image area of the road surface mask image;

[0253] The first boundary line, the second boundary line, the third boundary line and the fourth boundary line form an embeddable area.

[0254] Optionally, the determining module 131 is specifically configured to determine region ratio information according to the endpoint coordinate information of the first boundary line of the embeddable area, the endpoint coordinate information of the second boundary line and the width information of the scene background image.

[0255] Optionally, it further includes an adjustment module;

[0256] The adjustment module is configured to adjust the region ratio information if the region ratio information is greater than a preset threshold to obtain adjusted region ratio information.

[0257] Optionally, the determination module 131 is specifically configured to determine the region segmentation position according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line, and the region ratio information.

[0258] Optionally, the embedding module 132 is specifically used to determine the target embedding position from the target area in the target scene background image; the target embedding position is any position in the target area;

[0259] Determine the depth information of the target embedding position according to the scene depth information corresponding to the target scene background image;

[0260] Determine the image size of the target foreground object at the target embedding position in the target scene background image to be embedded according to the depth information of the target embedding position and the mapping relationship between the depth information and the image embedding size;

[0261] According to the image size, the target foreground object image is embedded at the target embedding position in the target scene background image to generate a composite image.

[0262] Optionally, an embedding module 132 is specifically configured to embed one or more of the foreground object images into a current scene background image to generate at least one synthetic image corresponding to the current scene background image, where the current scene background image is any one of the scene background images;

[0263] A plurality of composite images are obtained according to at least one composite image corresponding to each scene background image.

[0264] Optionally, the embedding module 132 is specifically used to embed the foreground object image to be embedded into the target scene background image according to the embedding position of each currently embedded foreground object image if the number of foreground object images currently embedded in the target scene background image does not exceed a preset embedding number upper limit, so as to generate a current composite image.

[0265] Optionally, the adjustment module is further used to use a pre-trained image harmonization model to repair the inharmonious areas in each composite image to obtain an optimized composite image.

[0266] Optionally, it also includes: a training module;

[0267] A training module is used to collect a training sample data set, the training sample data set includes: multiple groups of sample data, each group of sample data includes: a real shot image and a synthetic image constructed based on the real shot image;

[0268] The image harmonization model is trained using a training sample data set.

[0269] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar, so details are not described herein again.

[0270] The above modules may be one or more integrated circuits configured to implement the above method. For example, one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0271] The above modules may be connected or communicate with each other via a wired connection or a wireless connection. The wired connection may include metal wires, optical fibers, hybrid wires, etc., or any combination thereof. The wireless connection may include connections in the form of LAN, WAN, Bluetooth, ZigBee, or NFC, etc., or any combination thereof. Two or more modules may be combined into a single module, and any one module may be divided into two or more units. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems and devices described above may refer to the corresponding processes in the method embodiments, and details are not described herein again in this application.

[0272] Figure 14 FIG. is a schematic structural diagram of an electronic device provided in an embodiment of the present application, and the device may be a computing device with data processing capabilities.

[0273] The device may include: a processor 801 and a storage medium 802.

[0274] The storage medium 802 is used to store a program, and the processor 801 calls the program stored in the storage medium 802 to execute the above method embodiment. The specific implementation manners and technical effects are similar, and details are not described herein again.

[0275] Among them, the storage medium 802 stores program code, and when the program code is executed by the processor 801, the processor 801 is caused to execute various steps in the image processing method according to various exemplary embodiments of the present application described in the "Exemplary Method" section of this specification.

[0276] The processor 801 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0277] The storage medium 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The storage medium may include at least one type of storage medium, for example, it may include flash memory, a hard disk, a multimedia card, a card-type storage medium, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic storage medium, a magnetic disk, an optical disk, and so on. The storage medium is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The storage medium 802 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0278] Optionally, the present application also provides a program product, such as a computer-readable storage medium, including a program that is used to execute the above method embodiments when executed by a processor.

[0279] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0280] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0281] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0282] The above-mentioned integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only storage media (English: Read-Only Memory, abbreviated as: ROM), random access storage media (English: Random Access Memory, abbreviated as: RAM), magnetic disks or optical discs and other various media that can store program codes.

Claims

1. An image processing method, characterized in that, Including: Extracting at least one foreground object image from each acquired initial image; Determining the scene depth information corresponding to each scene background image according to each acquired scene background image; Determining a target area in each scene background image according to each of the scene background images, where the target area is a reasonable area allowing the foreground object image to be embedded; Embedding at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate a plurality of composite images; The embedding of at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate a plurality of composite images includes: Determining a target embedding position from the target area in the target scene background image; the target embedding position is any position in the target area; Determining the depth information of the target embedding position according to the scene depth information corresponding to the target scene background image; Determining the size of the target foreground object image to be embedded at the target embedding position in the target scene background image according to the depth information of the target embedding position and the mapping relationship between the depth information and the image embedding size, where the mapping relationship is pre-constructed based on the display rule of objects appearing larger when closer; Embedding the target foreground object image at the target embedding position in the target scene background image according to the image size to generate a composite image; The determining of the target area in each scene background image according to each of the scene background images includes: Generating a road surface mask image in the scene background image according to the scene background image; Respectively determining a first distance between the road surface and the bottom edge of the scene background image and a second distance between the road surface and the top edge of the scene background image according to the road surface mask image; Determining a first boundary line according to the first distance and the boundary of the road surface mask image; Determining a second boundary line according to the second distance and the boundary of the road surface mask image; Respectively determining a third boundary line and a fourth boundary line according to the first boundary line, the second boundary line, and the image area of the road surface mask image; Forming an embeddable area by the first boundary line, the second boundary line, the third boundary line, and the fourth boundary line; Determining the target area in each scene background image according to the embeddable area.

2. The method according to claim 1, characterized in that, The extracting of at least one foreground object image from each acquired initial image includes: Collecting a plurality of initial images, each initial image containing at least one foreground object; Extracting each foreground object image corresponding to each foreground object description information from each initial image by using a large language vision model according to each initial image and at least one foreground object description information corresponding to each initial image.

3. The method according to claim 2, characterized in that, The extracting of each foreground object image corresponding to each foreground object description information from each initial image by using a large language vision model according to each initial image and at least one foreground object description information corresponding to each initial image includes: The large language vision model encodes the initial image and the target foreground object description information corresponding to the initial image respectively to generate fused encoding information; Performing segmentation decoding on the fused encoding information, extracting a binary mask image of the target foreground object corresponding to the target foreground object description information from the initial image, and converting the binary mask image to obtain a bounding box of the target foreground object; According to the bounding box of the target foreground object, a target foreground object image is extracted.

4. The method according to claim 1, wherein The determining the scene depth information corresponding to each scene background image according to the collected scene background images includes: Encoding the scene background image by using a depth estimation model and a segmentation model respectively to obtain a feature map of the scene background image; Decoding the feature map of the scene background image to obtain a depth map of the scene background image; Determining the depth information of the scene background image according to the depth map.

5. The method according to claim 1, wherein The determining the target region in each scene background image according to the embeddable region includes: Determining region ratio information according to the boundary information of the embeddable region and the width information of the scene background image; the region ratio information is used to represent the ratio of the to-be-selected target region to the embeddable region; Determining a region segmentation position according to the region ratio information and the boundary information of the embeddable region, and segmenting the embeddable region based on the region segmentation position to obtain the target region.

6. The method according to claim 5, wherein The determining the region ratio information according to the boundary information of the embeddable region and the width information of the scene background image includes: Determining the region ratio information according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line, and the width information of the scene background image.

7. The method according to claim 5, wherein Before the determining the region segmentation position according to the region ratio information and the boundary information of the embeddable region, and segmenting the embeddable region based on the region segmentation position to obtain the target region, further includes: If the region ratio information is greater than a preset threshold, adjusting the region ratio information to obtain adjusted region ratio information.

8. The method according to claim 5, wherein The determining the region segmentation position according to the region ratio information and the boundary information of the embeddable region includes: Determining the region segmentation position according to the endpoint coordinate information of the first boundary line of the embeddable region, the endpoint coordinate information of the second boundary line, and the region ratio information.

9. The method according to claim 1, characterized in that The generating a plurality of synthetic images by embedding at least one foreground object image into at least one scene background image includes: Embedding one or more of each foreground object image into the current scene background image to generate at least one synthetic image corresponding to the current scene background image, where the current scene background image is any one of the scene background images; Obtaining the plurality of synthetic images according to at least one synthetic image corresponding to each scene background image.

10. The method according to claim 1, characterized in that, The generating a plurality of synthetic images by embedding at least one foreground object image into at least one scene background image includes: If the number of foreground object images currently embedded in the target scene background image does not exceed the preset upper limit of the embedding number, the foreground object image to be embedded is embedded into the target scene background image according to the embedding position of each currently embedded foreground object image to generate a current composite image.

11. The method according to any one of claims 1 to 10, characterized in that, After generating a plurality of composite images, the method further includes: The pre-trained image harmonization model is used to repair the inharmonious areas in each composite image to obtain an optimized composite image.

12. The method according to claim 11, wherein The training process of the image harmonization model includes: Collecting a training sample data set, the training sample data set comprising: a plurality of groups of sample data, each group of sample data comprising: a real shot image and a synthetic image constructed based on the real shot image; The image harmonization model is obtained by training using the training sample data set.

13. An image processing apparatus, characterized in that, include: Get modules, determine modules, embed modules; The acquisition module is used to extract at least one foreground object image from each acquired initial image; The determination module is used to determine the scene depth information corresponding to each scene background image according to the collected scene background images; The determination module is used to determine a target area in each scene background image according to each scene background image, wherein the target area is a reasonable area allowing the foreground object image to be embedded; The embedding module is used to embed at least one foreground object image into at least one scene background image according to the scene depth information corresponding to each scene background image and the target area in each scene background image to generate a plurality of composite images; The embedding module is specifically used to determine the target embedding position from the target area in the target scene background image; The target embedding position is any position in the target area; Determining the depth information of the target embedding position according to the scene depth information corresponding to the target scene background image; Determine the image size of the target foreground object to be embedded in the target scene background image at the target embedding position according to the depth information of the target embedding position and the mapping relationship between the depth information and the image embedding size, wherein the mapping relationship is pre-constructed based on the display rule that near objects are larger and far objects are smaller; embedding the target foreground object image at the target embedding position in the target scene background image according to the image size to generate a composite image; The determination module is specifically used for: According to the scene background image, a road surface mask image in the scene background image is generated; Determining, according to the road mask image, a first distance of the road surface close to the bottom edge of the scene background image and a second distance of the road surface close to the top edge of the scene background image; Determining a first boundary line according to the first distance and a boundary of the road surface mask image; determining a second boundary line according to the second distance and a boundary of the road surface mask image; Determine a third boundary line and a fourth boundary line respectively according to the first boundary line, the second boundary line and the image area of ​​the road surface mask image; The first boundary line, the second boundary line, the third boundary line and the fourth boundary line constitute an embeddable area; Determine the target regions in the background images of each scenario according to the embeddable regions.

14. An electronic device, characterized in that, Including: A processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus. The processor executes the program instructions to implement the image processing method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, it implements the image processing method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Portrait segmentation method and device, storage medium and electronic equipment

    CN112085002A

  • Live broadcast image synthesis method and device, terminal equipment and readable storage medium

    CN113837979A

  • Sample data processing method, related equipment and storage medium

    CN116071614A