Training data semantic synthesis method, device and equipment based on camera parameter model
Through the semantic synthesis method of training data based on the camera parameter model, the problem of poor matching between target objects and background images in the existing technology is solved, and high-quality image synthesis effects are achieved.
Patent Information
- Application Number
- CN202510714134.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-05
AI Technical Summary
When processing complex scenes, existing image fusion technology has difficulty in accurately locating and synthesizing target objects and background images, resulting in unnatural and unrealistic synthesized images.
A semantic synthesis method of training data based on camera parameter model is adopted to obtain target object and scene images, perform image segmentation, semantic segmentation and camera parameter model superposition to generate high-quality synthetic images.
The matching accuracy between the target object and the background image is improved, ensuring the natural and realistic effect of the synthesized image.
Smart Images

Figure CN120599261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a method, device and equipment for semantic synthesis of training data based on a camera parameter model. Background Art
[0002] Image fusion is an important technology that aims to combine images from different sources into a complete, harmonious image. This technology has broad application in many fields, such as augmented reality, virtual reality, medical image processing, photography, and advertising design. In these fields, image fusion can be used to merge objects from different scenes to create richer, more realistic composite images.
[0003] Current image fusion techniques typically use methods such as layer overlay, masking, edge detection, and alignment to combine the target object with the background image. These methods create different layers on the image and use various blending modes to overlay the target object into the target scene. Some existing solutions use basic object positioning methods, such as manual selection or simple boundary detection algorithms, to determine the location of the target object in the scene.
[0004] However, these methods may not perform well when dealing with complex scenes or locating objects. Existing methods have limitations in matching features such as angle and size between the target object and the background image, resulting in unnatural and unrealistic synthesized images. Therefore, existing technologies struggle to accurately locate and synthesize target objects in complex scenes, ensuring a visually harmonious match between the target object and the background image. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a method, apparatus and device for semantic synthesis of training data based on a camera parameter model to solve the problem of limitations in the prior art in processing the matching of features such as angle and size between the target object and the background image.
[0006] In a first aspect, an embodiment of the present invention provides a method for semantic synthesis of training data based on a camera parameter model, the method comprising:
[0007] Acquire a first image of a target object in a first scene and a second image including the target scene, wherein the target scene includes a plurality of scene objects, and the second image is acquired by an imaging device;
[0008] performing image segmentation on the first image according to the target object to obtain an initial template image;
[0009] Performing image change processing on the initial template image to obtain a first preset number of standard template images, wherein the image change processing includes image scaling and / or image rotation;
[0010] Inputting the second image into a pre-built semantic large model to perform semantic segmentation on the scene objects therein to obtain a contour image of the scene objects associated with the target object;
[0011] According to the contour image, the rotation angle and the scaling factor of the standard template image and the camera parameter model of the imaging device, the standard template image is superimposed on a designated area of the second image to obtain a target composite image.
[0012] Preferably, the target object is located at the center of the first image, and the step of segmenting the first image according to the target object to obtain an initial template image includes:
[0013] Obtaining a seed point according to the image center of the first image;
[0014] Determine a preset range in the first image according to the seed point and the length or width of the first image, wherein the preset range is as follows;
[0015] Randomly selecting a second preset number of random points within the preset range as first prompt points;
[0016] Performing grayscale equalization processing according to the grayscale value of the first prompt point to obtain a plurality of second prompt points;
[0017] Obtaining a segmented region of the first image according to the seed point, the second prompt point, and an image segmentation model;
[0018] Pixels in an area outside the segmented area in the first image are set according to a preset transparency and a first preset color value to obtain an initial template image.
[0019] Preferably, grayscale values of the first prompting points are subjected to grayscale equalization processing to obtain a plurality of second prompting points, including:
[0020] grayscale the first image to obtain a grayscale value of each of the first prompt points;
[0021] Dividing the grayscale value range evenly to obtain a third preset number of grayscale value intervals, wherein the grayscale value range is 0 to 255;
[0022] placing the first prompt point in the corresponding grayscale value interval according to the grayscale value of the first prompt point;
[0023] If a grayscale value interval includes two or more first prompt points, remove the first prompt points that are farther away from the center point according to the coordinates of the first prompt points;
[0024] A plurality of second prompt points are obtained according to the first prompt points in the grayscale value interval.
[0025] Preferably, the step of inputting the second image into a pre-built semantic large model to perform semantic segmentation on scene objects therein to obtain a contour image of the scene objects associated with the target object comprises:
[0026] Acquire scene items associated with the target object to obtain the target scene object;
[0027] Performing semantic segmentation on the target scene object in the second image according to the semantic large model to obtain a segmented image;
[0028] Performing contour extraction on the segmented image to obtain a contour area of the target scene object;
[0029] Pixel points within the contour area of the second image are set according to a second preset color value to obtain a contour image.
[0030] Preferably, the step of superimposing the standard template image onto a designated area of the second image based on the contour image, the rotation angle and the scaling factor of the standard template image, and the camera parameter model of the imaging device to obtain a target composite image includes:
[0031] Obtaining a minimum circumscribed rectangular frame of the contour image in the direction of the rotation angle;
[0032] Sequentially acquiring left boundary points and right boundary points of the contour image according to the upper boundary and the lower boundary of the minimum circumscribed rectangular frame to obtain a left boundary point sequence and a right boundary point sequence;
[0033] Calculating the width of each row of the contour image according to the left boundary point sequence and the right boundary point sequence to obtain a width sequence;
[0034] A height sequence is acquired according to the width sequence, the rotation angle, and the scaling factor, wherein the height sequence is acquired based on the following formula:
[0035] Height(i)=R*Width(i)*α
[0036] Wherein, Height(i) is the height sequence, R is the aspect ratio of the standard template image, α is the rotation angle, and i is an integer greater than 0;
[0037] Acquire a rectangular region sequence according to the height sequence and the width sequence, wherein the rectangular region sequence includes a plurality of rectangular regions, the upper edges of the rectangular regions are determined according to the width sequence, and the heights of the rectangular regions are determined based on height values corresponding to the upper edges in the height sequence;
[0038] Acquire a designated area of the second image according to the rectangular area sequence;
[0039] The standard template image is superimposed on the designated area according to the camera parameter model to obtain a target composite image.
[0040] Preferably, acquiring the designated area of the second image according to the rectangular area sequence includes:
[0041] Obtaining a rectangular region that meets a preset condition in the rectangular region sequence to obtain a plurality of candidate rectangular regions, wherein the preset condition includes: the color value of each pixel in the rectangular region is the second preset color value;
[0042] The candidate rectangular region with the largest area is obtained to obtain the designated region of the second image.
[0043] Preferably, superimposing the standard template image onto the designated area according to the camera parameter model to obtain a target composite image comprises:
[0044] Adjusting the size of the standard template image to be the same as the size of the designated area;
[0045] The standard template image is superimposed on the designated area according to a preset fusion rule and the camera parameter model to obtain a target composite image, wherein the preset fusion rule includes:
[0046] If the pixels of the designated area correspond to the pixels of the outline area in the standard template image, then the pixels of the designated area are replaced with the corresponding pixels of the outline area;
[0047] If the pixel points in the designated area correspond to the pixel points outside the contour area in the standard template image, the pixel points in the designated area are retained.
[0048] Preferably, superimposing the standard template image onto the designated area according to a preset fusion rule and the camera parameter model to obtain a target composite image comprises:
[0049] Acquire calibration parameters and an imaging mode of an imaging device used to capture the second image based on the camera parameter model, wherein the calibration parameters include an intrinsic parameter matrix and a distortion coefficient, and the imaging mode includes a visible light mode and an infrared night vision mode;
[0050] Correcting the standard target image according to the calibration parameters and the imaging mode to obtain a corrected standard template image;
[0051] extracting color histogram features and texture features of the second image to obtain a target style feature vector;
[0052] Calculating style normalization conversion parameters based on the target style feature vector and the corresponding feature vector of the corrected standard template image;
[0053] The standard template image corrected by the style normalization conversion parameters is subjected to normalization mapping processing to obtain a standard template image consistent with the style of the second image. In a second aspect, an embodiment of the present invention provides a training data semantic synthesis device based on a camera parameter model, the device comprising:
[0054] an image acquisition module, configured to acquire a first image of a target object in a first scene and a second image including the target scene, wherein the target scene includes a plurality of scene objects, and the second image is acquired by an imaging device;
[0055] an image segmentation module, configured to segment the first image according to the target object to obtain an initial template image;
[0056] an image transformation module, configured to perform image change processing on the initial template image to obtain a first preset number of standard template images, wherein the image change processing includes image scaling and / or image rotation;
[0057] a semantic segmentation module, configured to input the second image into a pre-built semantic model to perform semantic segmentation on scene objects therein, and obtain a contour image of the scene objects associated with the target object;
[0058] An image synthesis module is used to superimpose the standard template image on a specified area of the second image based on the contour image, the rotation angle and scaling factor of the standard template image and the camera parameter model of the imaging device to obtain a target synthetic image.
[0059] In a third aspect, an embodiment of the present invention provides an electronic device comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect of the above-mentioned embodiment.
[0060] In a fourth aspect, an embodiment of the present invention provides a storage medium having computer program instructions stored thereon, which implements the method of the first aspect of the above-mentioned embodiment when the computer program instructions are executed by a processor.
[0061] In summary, the beneficial effects of the present invention are as follows:
[0062] The embodiment of the present invention provides a method, device and apparatus for semantic synthesis of training data based on a camera parameter model, which separates the target object and the target scene by acquiring a first image of the target object in a first scene and a second image including the target scene, thereby helping to better handle the fusion between the two and improve the quality and authenticity of the synthesized image; performing image segmentation on the first image according to the target object to obtain an initial template image, and the image segmentation operation can accurately extract the target object and eliminate the interference between the target object and the background; performing image change processing on the initial template image to obtain a first preset number of standard template images, which can generate standard template images of different angles and sizes, and these standard template images provide the target object. Provide diversified presentation forms so that the target object can adapt to the designated areas in different scenes: perform semantic segmentation on the scene objects in the second image, obtain the contour image of the scene objects associated with the target object, be able to identify different objects in the target scene, and extract the contour image of the scene objects associated with the target object. This contour image can provide an accurate reference for the synthesis process, so that the target object can be naturally integrated into the target scene; according to the rotation angle and scaling factor of the contour image and the standard template image, the standard template image is superimposed on the designated area of the second image, which can ensure that the size, angle and position of the target object match the target scene, and help ensure the naturalness and reality of the synthesized image. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work, and these are all within the scope of protection of the present invention.
[0064] Figure 1 4 is a flow chart of a method for semantic synthesis of training data based on a camera parameter model according to an embodiment of the present invention.
[0065] Figure 2 This is another flowchart of the method for semantic synthesis of training data based on a camera parameter model according to an embodiment of the present invention.
[0066] Figure 3 This is another flowchart of the method for semantic synthesis of training data based on a camera parameter model according to an embodiment of the present invention.
[0067] Figure 4 4 is a structural diagram of a device for semantic synthesis of training data based on a camera parameter model according to an embodiment of the present invention.
[0068] Figure 5 2 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0069] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present invention by illustrating examples of the present invention.
[0070] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0071] Example 1
[0072] See Figure 1 , an embodiment of the present invention provides a method for semantic synthesis of training data based on a camera parameter model, the method comprising:
[0073] S1. Acquire a first image of a target object in a first scene and a second image including the target scene, wherein the target scene includes a plurality of scene objects, and the second image is acquired by an imaging device;
[0074] Specifically, the purpose of this step is to obtain the two images required for image fusion: one is a first image containing the target object, and the other is a second image containing the target scene. These two images are the basis of the entire image fusion process. The target object is the object that needs to be fused into the second image, such as a person, object, animal, etc. The target scene is the scene to be fused with the target object, such as a room, a landscape, etc. The scene usually includes multiple scene objects, such as furniture, buildings, natural landscapes, etc.; assuming that the image to be synthesized is a person lying on a sofa, in this scene, the target object in the first image is a person, and the first scene can be any scene, which can be the same as the target scene or any scene. Preferably, the first scene should be simple enough, such as a white wall, to avoid a complex background for subsequent processing. In another preferred embodiment, the first scene is the same as the target scene, which helps to maintain the consistency of the image's lighting, color, and other visual features, and improve the naturalness of the synthesis effect. In one embodiment, the second image is obtained by an imaging device.
[0075] S2. Segmenting the first image according to the target object to obtain an initial template image;
[0076] Specifically, image segmentation is a computer vision technology used to divide an image into different regions (or segmented regions) to identify and extract specific objects in the image. A suitable image segmentation model, such as a semantic segmentation model or an instance segmentation model, is used to process the first image. The image segmentation model identifies the target object based on specific features in the first image, distinguishes the target object area from the background and other parts in the first image, forms a segmented region, and thus obtains an initial template image.
[0077] Preferably, see Figure 2 , the target object is located at the center of the first image, and segmenting the first image according to the target object to obtain an initial template image includes:
[0078] S21. Acquire a seed point according to the image center of the first image;
[0079] Specifically, in this embodiment, the target object is located at the center of the first image and occupies a prominent position in the image. The model can better distinguish between the target object and the background, reducing the risk of segmentation errors.
[0080] Furthermore, since the target object is located at the center of the first image, the image center point of the first image must be on the target object. A pixel point at the center of the image is selected as a seed point. A seed point is typically an initial point or starting point used to initiate image processing processes such as region growing or image segmentation. Therefore, the image center can serve as an initial reference position to facilitate distinguishing between the target object and the background. Analyzing the image starting from the center point can reduce deviations caused by noise or uneven background in other parts. Furthermore, the center point provides a relatively stable starting point, helping to reduce segmentation errors caused by the intersection of the target object and the background.
[0081] S22. Determine a preset range in the first image according to the seed point and the length or width of the first image, wherein the preset range is centered on the seed point;
[0082] Specifically, the preset range is determined in order to perform random point selection and image segmentation operations in the subsequent steps. In order to focus on processing near the target object and improve the accuracy of segmentation, the preset range is centered on the seed point. The preset range can be a rectangular area or an area of other shapes such as a circle. The size of the preset range is related to the length or width of the first image. The size of the range depends on the needs of the actual application. For example, the width of the range can be a part of the width of the first image (such as 1 / 4 or 1 / 2), and the height of the range can be a part of the height of the first image.
[0083] In a preferred embodiment, the preset range is a circular area centered at the seed point and having a radius of one-quarter of the smaller of the length and width of the first image. This means that the radius of the preset range will be a fraction of the smaller of the length and width of the first image, thereby ensuring a moderate range. The circular area is concentrated around the seed point, which helps to process the area where the target object is located and improves the accuracy of the processing. The circular area is symmetrical in shape, ensuring uniformity and consistency of the processed area. The circular area is suitable for target objects of different shapes (such as people, objects, animals, etc.) because the circular area is more adaptable to the shape of the target object.
[0084] S23, randomly selecting a second preset number of random points within the preset range as first prompt points;
[0085] Specifically, based on the previously determined preset range, a second preset number of random points are randomly selected within the range. The purpose of random selection is to avoid bias towards a specific area and to help ensure that these points are evenly distributed in different areas of the target object, thereby covering different features of the target object. The second preset number is determined based on actual needs. For example, the second preset number is set based on the complexity and size of the target object. For more complex or larger target objects, more prompt points may be required to cover different parts of the target object and improve the accuracy of the model; or the second preset number is determined based on the size of the preset range. If the preset range is large, more prompt points may be required to ensure uniform coverage throughout the range.
[0086] The randomly selected points are considered the first cue points, which serve as reference points for the model to identify and extract the target object. The first cue points serve as reference information for the model, providing the model with initial direction and positioning, enabling the model to better locate and identify the target object.
[0087] In a specific embodiment, the image segmentation model used in subsequent image segmentation is the EdgeSAM segmentation model. EdgeSAM is an accelerated variant of the SegmentAnything Model (SAM) and aims to provide a semantic segmentation model that runs efficiently on edge devices. Compared with the original SAM, EdgeSAM has significant improvements in speed and performance, and is especially suitable for real-time operation on edge devices. When selecting the second preset number, it is necessary to find a balance between the efficiency and accuracy of the model. The second preset number is 15, which is set to 15 based on factors such as the capabilities of the model, experimental results, and balancing efficiency and accuracy. This number can ensure that the model is evenly covered in different grayscale ranges, thereby improving the accuracy and stability of the model in identifying and segmenting target objects.
[0088] S24, performing grayscale equalization processing on the grayscale values of the first prompting points to obtain a plurality of second prompting points;
[0089] Specifically, the purpose of this step is to process the grayscale value of the selected first prompt point so that it is evenly distributed within the grayscale range and covers different grayscale ranges. This helps to improve the recognition and extraction capabilities of the subsequent segmentation model on target objects in different grayscale ranges. First, the grayscale value of each first prompt point is obtained, and the grayscale values of the first prompt point are sorted from low to high according to the grayscale. The sorted grayscale value list is traversed, and points with close grayscale values are screened. If the difference between adjacent grayscale values is less than the grayscale threshold, only one of the points is retained, and other points with close grayscale values are eliminated. The threshold can be a fixed value or adaptively adjusted according to the grayscale value data set of the first prompt point. This uniform distribution helps the subsequent segmentation model to better cover different grayscale ranges and improve the recognition and extraction capabilities of the target object. By eliminating points with close grayscale values, the model avoids bias towards points with too similar grayscale values, so that the model has a better recognition effect on target objects in different grayscale ranges.
[0090] Preferably, see Figure 3 , performing grayscale equalization processing according to the grayscale values of the seed point and the first prompt point, to obtain a plurality of second prompt points including:
[0091] S241, performing grayscale processing on the first image to obtain a grayscale value of each of the first prompt points;
[0092] Specifically, the first image is grayscaled and converted into a grayscale image, and grayscale values of the first prompt point in the first image are obtained from the grayscale image. These grayscale values are in the range of 0 to 255 and represent the brightness of the first prompt point in the image.
[0093] S242, dividing the grayscale value range evenly to obtain a third preset number of grayscale value intervals, wherein the grayscale value range is 0 to 255;
[0094] Specifically, the grayscale value range of 0 to 255 is evenly divided into a third preset number of grayscale value intervals, and the number of intervals is set according to the overall coverage of the grayscale value range and specific needs. Setting a reasonable number of intervals helps to ensure that the grayscale values of the first prompt point are evenly distributed. In the first embodiment, the second preset number is 16, that is, divided into 16 intervals, each interval has a range of 16 grayscale values (for example, 0-15, 16-31, etc.), and these intervals are used to group the first prompt point according to the grayscale value; the grayscale value range is evenly divided into 16 intervals, each interval has a range of 16 grayscale values. This division is fine and can better distinguish different grayscale ranges. 16 intervals provide a reasonable balance point, which can ensure the uniform classification of grayscale values without causing too few intervals to affect the recognition ability of the model.
[0095] S243, placing the seed point and the first prompt point in the corresponding grayscale value interval according to the grayscale values of the seed point and the first prompt point;
[0096] Specifically, according to the grayscale values of the seed point and the first hint point, these points are placed in the corresponding grayscale value intervals. For example, if the grayscale value of a point is between 0 and 15, it is placed in the first grayscale value interval.
[0097] S244: If the grayscale value interval includes two or more first prompt points, remove the first prompt points that are farther away from the center point according to the coordinates of the first prompt points;
[0098] Specifically, if a grayscale value interval includes two or more first prompt points, the first prompt points with a larger distance from the center point are removed based on their coordinates. That is, for all prompt points in each grayscale value interval, the distance between the prompt point and the center point (i.e., the seed point) is calculated using the Euclidean distance formula, and the prompt point closest to the center point is selected and retained, while the prompt points with a larger distance are removed.
[0099] The purpose of this step is to select cue points that are close to the center point in each grayscale value interval to ensure that the cue points are evenly distributed within the preset range while avoiding redundancy. The uniform distribution of cue points helps the model more accurately identify and extract the target object during subsequent segmentation.
[0100] S245: Obtain a plurality of second prompt points according to the first prompt points in the grayscale value interval.
[0101] Specifically, based on the retained first cue points, the first cue points in each interval are selected as second cue points based on the distribution of grayscale value intervals. This selection ensures that the second cue points are evenly distributed across different grayscale ranges. Through these steps, the grayscale values of the first cue points are preferably homogenized, ensuring that the second cue points are evenly distributed across the grayscale range. This helps the subsequent model better identify and extract the target object across different grayscale ranges, improving segmentation and processing accuracy.
[0102] S25. Obtain a segmented region of the first image according to the seed point, the second prompt point, and an image segmentation model;
[0103] Specifically, the seed point and the second cue point are input into the image segmentation model, and the model begins to segment the first image based on these points as initial reference information. The image segmentation model can be many different types of models, including convolutional neural network (CNN) models, DeepLab models, Segment Anything Model (SAM) models, etc. When selecting a model, it is necessary to select the most suitable model based on the requirements of the specific task and actual conditions to ensure high-quality completion of the semantic segmentation task. After the seed point and the second cue point are input into the image segmentation model, the model begins to distinguish the target object from the background by identifying the features around the seed point and the second cue point, such as color, texture, shape, etc. The model identifies the area of the target object by analyzing the features in the image and separates it from the background. The segmentation result generated by the model is a mask or segmentation area, which identifies the part of the first image containing the target object and distinguishes it from the background and other irrelevant parts.
[0104] In one embodiment, the image segmentation model is an EdgeSAM model, which is an accelerated variant of SAM specifically optimized for execution on edge devices. Compared to the original SAM, EdgeSAM significantly improves speed and efficiency while maintaining high accuracy.
[0105] S26 . Setting pixel points in an area outside the segmented area in the first image according to a preset transparency and a first preset color value to obtain an initial template image.
[0106] Specifically, a segmented region, i.e., the region containing the target object, is identified and extracted from the first image. Regions outside the segmented region are considered non-segmented regions. Pixels in the non-segmented region of the first image are set to a certain transparency based on a preset transparency. In one embodiment, the preset transparency can be set to completely transparent so that these regions do not affect the presentation of the target object during subsequent synthesis. Simultaneously, pixels in the non-segmented region of the first image are set to a specific color based on a first preset color value. This color value can be black or another color that does not affect the target object, ensuring that the non-segmented region does not interfere with subsequent synthesis.
[0107] In a specific embodiment, the preset transparency is 1, and the first preset color value is black. By setting the pixels that do not belong to the segmented area to black and completely opaque, while retaining the original features of the pixels that belong to the segmented area, it helps to maintain the clarity of the target object, reduce interference, and facilitate subsequent synthesis processes.
[0108] S3. Performing image change processing on the initial template image to obtain a first preset number of standard template images, wherein the image change processing includes image scaling and / or image rotation;
[0109] Specifically, by performing image scaling and / or image rotation on the initial template image, standard template images of different angles and sizes are generated. Image scaling refers to resizing the initial template image, which can enlarge or reduce the size of the target object. Image rotation refers to rotating the initial template image to change the angle of the target object. Standard template images of different sizes and angles are generated through image change processing. These standard template images are helpful for subsequent image synthesis, so that the target object can be integrated into the target scene at different angles and sizes. The first preset number can be selected according to the needs of the synthesis task. If the target object needs to be integrated with the scene at different angles, viewing angles, and sizes, more standard template images are required. It can also be determined based on the complexity and diversity of the target object. For complex target objects (such as a multi-jointed human body) or objects with large morphological changes, more standard template images may be required. The specific number is not limited here.
[0110] By generating standard template images at different angles and sizes, the position, angle, and size of the target object can be more flexibly adjusted in subsequent synthesis to adapt to different scenarios. Different image variation processing can enrich the presentation of the target object, making it more diverse and natural in the synthesized image.
[0111] S4, inputting the second image into a pre-built semantic large model to perform semantic segmentation on the scene objects therein, and obtaining a contour image of the scene objects associated with the target object;
[0112] Specifically, the "semantic large model" in this step refers to a large-scale visual model capable of semantically identifying and pixel-level segmenting multiple regions in an image, such as the SegmentAnything Model (SAM), DeepLab, Segformer, or Grounding DINO+SAM combined with textual cues. A notable feature of this type of model is that it can automatically identify semantically meaningful object regions in an image and convert them into pixel-level masks or boundary contours without requiring extensive manual annotation. Semantic segmentation is performed on the scene objects in the second image to obtain contour images of the scene objects associated with the target object. The purpose of this step is to identify and extract the scene objects associated with the target object and generate contour images of these scene objects. This helps to better integrate the target object with the relevant scene during subsequent image synthesis. The semantic segmentation model used in this step can be the second semantic segmentation model mentioned above, or another model. In the segmentation results, the scene objects associated with the target object are identified. These associated scene objects may be the region into which the target object is to be synthesized. The contours of the identified associated scene objects are extracted. The contour images are generated by extracting the edges or boundaries of the scene objects.
[0113] Preferably, the step of inputting the second image into a pre-built semantic large model to perform semantic segmentation on scene objects therein to obtain a contour image of the scene objects associated with the target object comprises:
[0114] S41, acquiring scene items associated with the target object to obtain the target scene object;
[0115] Specifically, the scene items associated with the target object are determined before image synthesis. They are the area where the target object is to be synthesized and can be directly acquired. For example, if the image to be synthesized is a person lying on a sofa, then the person is the target object and the scene items associated with the person are the sofa.
[0116] S42, performing semantic segmentation on the target scene object in the second image according to the pre-built semantic large model to obtain a segmented image;
[0117] Specifically, the second image is provided as input to a pre-built semantic large model. The model analyzes the image to identify the target scene object and other scene items. The model distinguishes the target scene object from other parts to generate a segmented image. The segmented image is an image with a marked target scene object area, where the target scene object is marked as a specific category (for example, represented by a different color or label). A segmented image can be generated by performing semantic segmentation on the target scene object in the second image using a second semantic segmentation model. This segmented image clearly identifies the area of the target scene object in the image, providing important reference information for subsequent image processing and synthesis.
[0118] In one embodiment, the large semantic model is SAM, which has a more complex model structure and can handle richer data and complex scenes. The image segmentation model uses EdgeSAM to quickly and accurately extract the target object area, laying the foundation for subsequent processing. The second semantic segmentation model uses SAM to better identify scene objects associated with the target object and complex scenes, improving the model's performance in the second image. The combined use of EdgeSAM and SAM can improve the model's adaptability to different scenarios and tasks.
[0119] S43, performing contour extraction on the segmented image to obtain a contour area of the target scene object;
[0120] Specifically, the segmented image is generated after semantic segmentation of the target scene object in the second image by the second semantic segmentation model. The segmented image distinguishes the area of the target scene object from other parts, and contour extraction is performed on the segmented image to identify the boundaries and edges of the target scene object in the image. Contour extraction usually uses edge detection algorithms (such as Canny edge detection) or other methods to identify boundary lines in the image. The boundaries of the extracted target scene object are formed into a contour area. The contour area can be a closed area formed by continuous boundary lines, which clearly indicates the shape and position of the target scene object.
[0121] S44. Setting pixel points within the contour area of the second image according to a second preset color value to obtain a contour image.
[0122] Specifically, the pixels within the outline region of the second image are set according to a second preset color value to obtain a contour image. This step aims to make the contour region clearly visible in the image by setting the pixels within the contour region to a specific color value. A clear contour image facilitates precise matching of the target object and target scene during the synthesis process, ensuring a natural synthesis effect. This processing method helps accurately identify and match the contours of the target object and target scene during subsequent image synthesis and processing.
[0123] The second preset color value is typically a bright color or a color that is significantly different from the background color (such as white or another contrasting color). Preferably, the second preset color value is white, which often creates a high contrast with other colors in the image. This high contrast helps clearly highlight the outline area, making it clearly visible in the image. This is because white has the highest brightness in the image, making the outline area easily identifiable and distinguishable in the image.
[0124] S5. Based on the contour image, the rotation angle and the scaling factor of the standard template image, and the camera parameter model of the imaging device, superimpose the standard template image on a designated area of the second image to obtain a target composite image.
[0125] Specifically, the camera parameter model of the imaging device includes the intrinsic parameter matrix (focal length, principal point, etc.) and distortion parameters of the device, which represent the actual imaging behavior of the target device. According to the rotation angle and scaling factor of the contour image and the standard template image, the standard template image is superimposed on the specified area of the second image to obtain the target composite image. The purpose of this step is to accurately superimpose the standard template image on the predetermined specified area in the second image, thereby realizing the fusion of the target object and the target scene. The specified area is determined based on the contour image, and the angle and size of the standard template image are adjusted according to the determined rotation angle and scaling factor. The adjusted standard template image matches the expected position, angle and size of the target object. The adjusted standard template image is superimposed on the specified area of the second image. The superposition process accurately fuses the standard template image with the specified area of the second image, and finally generates the target composite image.
[0126] Preferably, the step of superimposing the standard template image onto a designated area of the second image according to the rotation angle and scaling factor of the contour image and the standard template image to obtain a target composite image comprises:
[0127] S51, obtaining the minimum circumscribed rectangular frame of the contour image in the direction of the rotation angle;
[0128] Specifically, a minimum bounding rectangle is calculated based on the contour image in the direction of the rotation angle. The contour image can be rotated according to a given rotation angle, and the contour is extracted from the rotated image to obtain the coordinate points of the contour line. A computational geometry method is then used to find the minimum bounding rectangle that surrounds the extracted contour line in the rotated image. The vertex coordinates of the minimum bounding rectangle are recorded, including the coordinates of the upper left corner, upper right corner, lower right corner, and lower left corner.
[0129] S52, sequentially acquiring left boundary points and right boundary points of the contour image according to the upper boundary and the lower boundary of the minimum circumscribed rectangular frame to obtain a left boundary point sequence and a right boundary point sequence;
[0130] Specifically, the minimum bounding rectangle is a rectangular frame that completely surrounds the target object in the contour image, including four boundaries: upper, lower, left, and right. The upper and lower boundaries of the minimum bounding rectangle will be used in this step. Starting from the upper boundary of the minimum bounding rectangle, process the contour image row by row. For each row (from the upper boundary to the lower boundary), find the left boundary point and the right boundary point in the contour image. On each row, traverse from left to right, find the first non-background pixel point, and record its position as the left boundary point. On each row, traverse from right to left, find the first non-background pixel point, and record its position as the right boundary point. Store the left boundary point and right boundary point of each row in a sequence respectively. The left boundary point sequence represents the left boundary position of the target object on each row. The right boundary point sequence represents the right boundary position of the target object on each row.
[0131] By obtaining the left and right boundary point sequences, the horizontal width of the target object can be accurately determined, providing a reference for subsequent synthesis.
[0132] S53, calculating the width of each row of the contour image according to the left boundary point sequence and the right boundary point sequence to obtain a width sequence;
[0133] Specifically, to calculate the width of the target object on each line and provide information for subsequent synthesis and processing, for each line, the width is calculated using the left boundary points and the right boundary points to obtain a width sequence;
[0134] The width sequence can be calculated by the following formula:
[0135] Width(i)=||P_right(i)-P_left(i)||
[0136] Wherein, Width(i) represents the width sequence, P_right(i) and P_left(i) are the right boundary sequence and the left boundary sequence respectively, and “||||” represents the Euclidean distance between two points.
[0137] S54. Acquire a height sequence according to the width sequence, the rotation angle, and the scaling factor, wherein the height sequence is acquired based on the following formula:
[0138] Height(i)=R*Width(i)*α
[0139] Wherein, Height(i) is the height sequence, R is the aspect ratio of the standard template image, α is the rotation angle, and i is an integer greater than 0;
[0140] Specifically, the height of the target object in the contour area is calculated using the width sequence, the aspect ratio (R) of the standard template image, and the rotation angle (α). These height values help ensure that the size and shape of the target object match the contour area during subsequent processing and synthesis. Multiplying by R in the formula ensures that the target object maintains the correct height ratio while changing its width, while multiplying by α adjusts the height to match the rotation angle of the contour area. The height calculated for each row is stored in a sequence to form a height sequence. The height sequence provides important reference information for subsequent synthesis and processing, including the vertical size of the target object in the image. By combining the rotation angle and scaling factor to calculate the height sequence, the correct size of the target object can be ensured at different angles and scales.
[0141] S55. Acquire a rectangular region sequence according to the height sequence and the width sequence, wherein the rectangular region sequence includes a plurality of rectangular regions, the upper edges of the rectangular regions are determined according to the width sequence, and the heights of the rectangular regions are determined based on height values corresponding to the upper edges in the height sequence;
[0142] Specifically, multiple rectangular regions are generated by combining the height sequence and the width sequence. Each rectangular region represents a portion of the contour image at a different position and size. These regions provide the positioning information of the target object for subsequent processing and synthesis. For each row, the position of the upper boundary is obtained from the width sequence, and the corresponding height value is obtained from the height sequence. The width sequence and the height sequence are used to determine the boundary of the rectangular region, that is, the starting point, end point, and height of the rectangular region on each row. The rectangular region information of each row is stored in a sequence to form a rectangular region sequence. The rectangular region sequence includes multiple rectangular regions that cover different parts of the contour image.
[0143] S56. Acquire a designated area of the second image according to the rectangular area sequence;
[0144] Specifically, the sequence of rectangular regions is traversed to obtain each rectangular region. Each rectangular region represents a portion of the target object in the contour image. For each rectangular region, its corresponding position is located in the second image. The designated region is determined based on the position of the rectangular region in the second image. The coordinates, width, and height of the rectangular region are used to determine the position and range of the designated region in the second image. The position and range of the designated region correspond to the rectangular region.
[0145] Preferably, acquiring the designated area of the second image according to the rectangular area sequence includes:
[0146] S561: Obtain a rectangular region that meets a preset condition in the rectangular region sequence to obtain a plurality of candidate rectangular regions, wherein the preset condition includes: the color value of each pixel in the rectangular region is the second preset color value;
[0147] Specifically, the sequence of rectangular regions is traversed, and each rectangular region is checked one by one, with a preset condition that the color value of each pixel in each rectangular region is a second preset color value. For each rectangular region, if the color values of all pixels in the rectangular region meet the second preset color value, the region is added to the candidate rectangular region.
[0148] S562: Acquire the candidate rectangular region with the largest area to obtain a designated region of the second image.
[0149] Specifically, the largest rectangular region is selected from the candidate rectangular regions. The area is calculated by calculating the width and height of the rectangular region. The largest candidate rectangular region is usually the most suitable region, meeting the preset conditions and covering a large area. The largest candidate rectangular region is selected as the designated region of the second image. By screening and selecting the largest region that meets the conditions, the designated region is ensured to be more representative and accurate during the synthesis process.
[0150] S57 . Superimpose the standard template image on the designated area according to the camera parameter model to obtain a target composite image.
[0151] Specifically, in the final stage of image fusion, the morphology and pixel values of the standard template image are adjusted based on camera parameters and superimposed onto the designated area of the second image in a natural, continuous, and seamless manner, resulting in a target composite image with matching structure and uniform style. During the superposition process, the standard template image is overlaid onto the designated area. Synthesis algorithms, such as alpha blending or other transparency blending methods, can be used during the superposition process to achieve natural fusion. After the superposition is complete, the target composite image is generated. The target composite image is the result of fusing the standard template image with the second image in the designated area, and includes a combination of the target object and the target scene.
[0152] Preferably, the S57 includes:
[0153] S571, adjusting the size of the standard template image to be the same as the size of the designated area;
[0154] First, resize the standard template image to match the size of the designated area. Adjust the width and height of the standard template image to match the width and height of the designated area. This adjustment allows the standard template image to blend more accurately with the designated area, improving the synthesis effect.
[0155] S572: Superimpose the standard template image on the designated area according to a preset fusion rule to obtain a target composite image, wherein the preset fusion rule includes:
[0156] If the pixels of the designated area correspond to the pixels of the outline area in the standard template image, then the pixels of the designated area are replaced with the corresponding pixels of the outline area;
[0157] If the pixel points in the designated area correspond to the pixel points outside the contour area in the standard template image, the pixel points in the designated area are retained.
[0158] The preset fusion rule defines how the standard template image and the specified area are fused during the synthesis process. If the pixels in the specified area correspond to the pixels in the contour area of the standard template image, the pixels in the specified area are replaced with the pixels corresponding to the contour area in the standard template image. This rule ensures that the target object and the target scene are accurately matched in the image. If the pixels in the specified area correspond to the pixels outside the contour area of the standard template image, the pixels in the specified area are retained, and the original content in the second image is maintained. This rule ensures that the characteristics of the original scene are retained during the synthesis process without being interfered with by the standard template image. According to the preset fusion rule, the fusion of the target object and the target scene is ensured to be natural and realistic, and the pixels in the specified area corresponding to those outside the contour area are retained, maintaining the characteristics of the original scene.
[0159] By adjusting the standard template image to the same size as the designated area and overlaying it according to pre-set fusion rules, the target object and the target scene can be precisely integrated. These steps improve the naturalness and realism of the synthesis effect, ensuring that the target object is accurately presented in the target scene while preserving the characteristics of the original scene.
[0160] In one embodiment, superimposing the standard template image onto the designated area according to a preset fusion rule to obtain a target composite image includes:
[0161] Acquire calibration parameters and an imaging mode of an imaging device used to capture the second image according to the camera parameter model, wherein the calibration parameters include an intrinsic parameter matrix and a distortion coefficient, and the imaging mode includes a visible light mode and an infrared night vision mode;
[0162] Specifically, calibration parameters refer to a set of parameters reflecting the camera's internal geometry, obtained through calibration experiments. The intrinsic parameter matrix, which includes the focal length, principal point position, and pixel scale factor, describes the transformation from three-dimensional projection to the image plane. Distortion coefficients quantify radial or tangential image distortion caused by the lens, such as barrel distortion or pincushion distortion. Imaging mode refers to the camera's operating mode. For example, visible light mode typically captures three-channel RGB images, while infrared night vision mode outputs a single-channel grayscale image, often accompanied by significant noise or nonlinear brightness compression.
[0163] The purpose of this step is to provide basic information for the subsequent simulation of the physical imaging characteristics of the standard template image. Since the second image represents the final target scene, in order to make the overlay image (template image) consistent with it in terms of geometry and photosensitivity, it is necessary to first obtain the key parameters of its imaging device.
[0164] Correcting the standard target image according to the calibration parameters and the imaging mode to obtain a corrected standard template image;
[0165] The "standard target image" is a standard template image obtained by segmenting and expanding the first image. Its content is the target object itself, which needs to be simulated into "what the target camera sees" in the current step.
[0166] The purpose of this step is to convert the standard template image from the imaging characteristics of the original device to the imaging style of the target device (i.e., the camera that captured the second image), so that the two maintain consistency in terms of optical morphology and noise distribution, to avoid structural misalignment or style abruptness after fusion.
[0167] In practice, this step is typically divided into two types of operations: geometric correction and imaging style correction. Geometric correction utilizes the distortion coefficients and intrinsic matrix in the calibration parameters, calling the distortion correction function to map the standard template image, eliminating nonlinear distortion caused by the lens. Imaging mode simulation converts the image into grayscale depending on whether it is in infrared mode, and superimposes common noise under simulated low-light conditions, such as Poisson noise, Gaussian noise, or speckle noise, to complete the simulation of illumination and photosensitive structure.
[0168] This step essentially completes a physical alignment across device domains by optically restoring the template image, providing morphological coordination and noise structure consistency for the overall fused image.
[0169] extracting color histogram features and texture features of the second image to obtain a target style feature vector;
[0170] Specifically, color histogram features refer to the statistical histogram of the color distribution of each pixel in the image, which can reflect information such as hue and brightness tendency; "texture features" often use gray-level co-occurrence matrix, Laplacian edge frequency distribution or the intermediate layer output of pre-trained feature extraction networks (such as CLIP, VGG) to describe the texture roughness, edge density, etc. of the image.
[0171] The goal of this step is to quantify the visual style characteristics of the second image, thereby providing a style anchor for the template image. Understanding the style characteristics of the second image ensures that the template image, after normalization, has the same visual style, improving the naturalness of the fusion.
[0172] Calculating style normalization conversion parameters based on the target style feature vector and the corresponding feature vector of the corrected standard template image;
[0173] Style normalization conversion parameters refer to adjusting the color and texture features of the template image to the target Figure 1 The required transformation matrix or vector is converted, including brightness mean adjustment, contrast scaling factor, texture spectrum stretching ratio, etc. The purpose of this step is to establish a mapping relationship between the style features of the two images, thereby providing an accurate transformation basis for the next step of image content normalization. This operation ensures high-quality fusion even between source images with inconsistent styles. The calculation method can adopt the AdaIN (Adaptive Instance Normalization) algorithm commonly used in the field of style transfer. Its core logic is to adjust the feature mean and standard deviation of the template image to the corresponding values of the target image. If CLIP or VGG features are used, the transformation path can also be constructed through the style projection matrix or Gram matrix difference.
[0174] This step provides a controllable and explainable style synchronization method for image fusion, greatly reducing the visual fragmentation caused by the inconsistent style of the synthesized image.
[0175] Normalization mapping processing is performed on the standard template image corrected according to the style normalization conversion parameters to obtain a standard template image consistent with the style of the second image.
[0176] Normalization mapping applies the transformation parameters obtained in the previous step to the template image, transforming its color and texture to bring its overall style closer to that of the second image. The goal of this step is to achieve final style alignment, ensuring that the template image not only matches the target camera in terms of physical geometry but also closely resembles the target scene in terms of subjective visual style, thereby improving the naturalness of the overall composite image and its consistency with the training samples.
[0177] In the implementation, if AdaIN was used in the previous step, this step directly replaces the template image's feature channel means and variances; if a histogram matching algorithm was used, a channel mapping function is performed. This effectively maps template images, originally collected from diverse sources and with diverse styles, into the target image's style domain, providing a consistent, natural, and controllable image foundation for subsequent overlay and training.
[0178] Example 2
[0179] See also Figure 4 , an embodiment of the present invention provides a training data semantic synthesis device based on a camera parameter model, the device comprising:
[0180] an image acquisition module, configured to acquire a first image of a target object in a first scene and a second image including the target scene, wherein the target scene includes a plurality of scene objects, and the second image is acquired by an imaging device;
[0181] an image segmentation module, configured to segment the first image according to the target object to obtain an initial template image;
[0182] an image transformation module, configured to perform image change processing on the initial template image to obtain a first preset number of standard template images, wherein the image change processing includes image scaling and / or image rotation;
[0183] a semantic segmentation module, configured to input the second image into a pre-built semantic model to perform semantic segmentation on scene objects therein, and obtain a contour image of the scene objects associated with the target object;
[0184] An image synthesis module is used to superimpose the standard template image on a specified area of the second image based on the contour image, the rotation angle and scaling factor of the standard template image and the camera parameter model of the imaging device to obtain a target synthetic image.
[0185] It should be noted that the modules and units in the training data semantic synthesis device based on the camera parameter model in this embodiment correspond one-to-one to the steps in the training data semantic synthesis method based on the camera parameter model in the aforementioned embodiment. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned training data semantic synthesis method based on the camera parameter model, which will not be repeated here.
[0186] Example 3
[0187] In addition, combined Figure 1 The training data semantic synthesis method based on the camera parameter model described in the embodiment of the present invention can be implemented by an electronic device. Figure 5A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention is shown.
[0188] An electronic device may include a processor and a memory storing computer program instructions.
[0189] Specifically, the processor may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.
[0190] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include a removable or non-removable (or fixed) medium. Where appropriate, the memory may be inside or outside the data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0191] The processor implements any one of the camera parameter model-based training data semantic synthesis methods in the above embodiments by reading and executing computer program instructions stored in the memory.
[0192] In one example, the electronic device may further include a communication interface and a bus. Figure 5 As shown, the processor 401 , the memory 402 , and the communication interface 403 are connected via a bus 410 and communicate with each other.
[0193] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.
[0194] Bus comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.
[0195] Example 4
[0196] In addition, in conjunction with the camera parameter model-based training data semantic synthesis method in the above-mentioned embodiments, embodiments of the present invention may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when executed by a processor, the computer program instructions implement any of the camera parameter model-based training data semantic synthesis methods in the above-mentioned embodiments.
[0197] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0198] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0199] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0200] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.
Claims
1. A method for semantic synthesis of training data based on a camera parameter model, characterized in that: The method comprises: Acquire a first image of a target object in a first scene and a second image including the target scene, wherein the target scene includes a plurality of scene objects, and the second image is acquired by an imaging device; performing image segmentation on the first image according to the target object to obtain an initial template image; Performing image change processing on the initial template image to obtain a first preset number of standard template images, wherein the image change processing includes image scaling and / or image rotation; Inputting the second image into a pre-built semantic large model to perform semantic segmentation on the scene objects therein to obtain a contour image of the scene objects associated with the target object; According to the contour image, the rotation angle and the scaling factor of the standard template image and the camera parameter model of the imaging device, the standard template image is superimposed on a designated area of the second image to obtain a target composite image.
2. The method for semantic synthesis of training data based on a camera parameter model according to claim 1, characterized in that: The target object is located at the center of the first image, and segmenting the first image according to the target object to obtain an initial template image includes: Obtaining a seed point according to the image center of the first image; determining a preset range in the first image according to the seed point and the length or width of the first image, wherein the preset range is centered on the seed point; Randomly selecting a second preset number of random points within the preset range as first prompt points; Performing grayscale equalization processing according to the grayscale value of the first prompt point to obtain a plurality of second prompt points; Obtaining a segmented region of the first image according to the seed point, the second prompt point, and an image segmentation model; Pixels in an area outside the segmented area in the first image are set according to a preset transparency and a first preset color value to obtain an initial template image.
3. The method for semantic synthesis of training data based on a camera parameter model according to claim 2, characterized in that: Grayscale values of the first prompt points are uniformly processed to obtain a plurality of second prompt points, including: grayscale the first image to obtain a grayscale value of each of the first prompt points; Dividing the grayscale value range evenly to obtain a third preset number of grayscale value intervals, wherein the grayscale value range is 0 to 255; placing the first prompt point in the corresponding grayscale value interval according to the grayscale value of the first prompt point; If a grayscale value interval includes two or more first prompt points, remove the first prompt points that are farther away from the center point according to the coordinates of the first prompt points; A plurality of second prompt points are obtained according to the first prompt points in the grayscale value interval.
4. The method for semantic synthesis of training data based on a camera parameter model according to claim 2, wherein: Inputting the second image into a pre-built semantic model to perform semantic segmentation on the scene objects therein to obtain a contour image of the scene objects associated with the target object includes: Acquire scene items associated with the target object to obtain the target scene object; Performing semantic segmentation on the target scene object in the second image according to the semantic large model to obtain a segmented image; Performing contour extraction on the segmented image to obtain a contour area of the target scene object; Pixel points within the contour area of the second image are set according to a second preset color value to obtain a contour image.
5. The method for semantic synthesis of training data based on a camera parameter model according to claim 4, characterized in that: The step of superimposing the standard template image onto a designated area of the second image according to the contour image, the rotation angle and the scaling factor of the standard template image, and the camera parameter model of the imaging device to obtain a target composite image includes: Obtaining a minimum circumscribed rectangular frame of the contour image in the direction of the rotation angle; Sequentially acquiring left boundary points and right boundary points of the contour image according to the upper boundary and the lower boundary of the minimum circumscribed rectangular frame to obtain a left boundary point sequence and a right boundary point sequence; Calculating the width of each row of the contour image according to the left boundary point sequence and the right boundary point sequence to obtain a width sequence; A height sequence is acquired according to the width sequence, the rotation angle, and the scaling factor, wherein the height sequence is acquired based on the following formula: Height(i)=R*Width(i)*α Wherein, Height(i) is the height sequence, R is the aspect ratio of the standard template image, α is the rotation angle, and i is an integer greater than 0; Acquire a rectangular region sequence according to the height sequence and the width sequence, wherein the rectangular region sequence includes a plurality of rectangular regions, the upper edges of the rectangular regions are determined according to the width sequence, and the heights of the rectangular regions are determined based on height values corresponding to the upper edges in the height sequence; Acquire a designated area of the second image according to the rectangular area sequence; The standard template image is superimposed on the designated area according to the camera parameter model to obtain a target composite image.
6. The method for semantic synthesis of training data based on a camera parameter model according to claim 5, wherein acquiring the specified area of the second image according to the rectangular area sequence comprises: Obtaining a rectangular region that meets a preset condition in the rectangular region sequence to obtain a plurality of candidate rectangular regions, wherein the preset condition includes: the color value of each pixel point in the rectangular region is the second preset color value; The candidate rectangular region with the largest area is obtained to obtain the designated region of the second image.
7. The method for semantic synthesis of training data based on a camera parameter model according to claim 5, characterized in that: The step of superimposing the standard template image onto the designated area according to the camera parameter model to obtain a target composite image includes: Adjusting the size of the standard template image to be the same as the size of the designated area; The standard template image is superimposed on the designated area according to a preset fusion rule and the camera parameter model to obtain a target composite image, wherein the preset fusion rule includes: If the pixels of the designated area correspond to the pixels of the outline area in the standard template image, then the pixels of the designated area are replaced with the corresponding pixels of the outline area; If the pixel points in the designated area correspond to the pixel points outside the contour area in the standard template image, the pixel points in the designated area are retained.
8. The method for semantic synthesis of training data based on a camera parameter model according to claim 7, characterized in that: The method comprises: superimposing the standard template image onto the designated area according to a preset fusion rule and the camera parameter model to obtain a target composite image, including: Acquire calibration parameters and an imaging mode of an imaging device used to capture the second image according to the camera parameter model, wherein the calibration parameters include an intrinsic parameter matrix and a distortion coefficient, and the imaging mode includes a visible light mode and an infrared night vision mode; Correcting the standard target image according to the calibration parameters and the imaging mode to obtain a corrected standard template image; extracting color histogram features and texture features of the second image to obtain a target style feature vector; Calculating style normalization conversion parameters based on the target style feature vector and the corresponding feature vector of the corrected standard template image; Normalization mapping processing is performed on the standard template image corrected according to the style normalization conversion parameters to obtain a standard template image consistent with the style of the second image.
9. A training data semantic synthesis device based on a camera parameter model, characterized in that: The device comprises: an image acquisition module, configured to acquire a first image of a target object in a first scene and a second image including the target scene, wherein the target scene includes a plurality of scene objects, and the second image is acquired by an imaging device; an image segmentation module, configured to segment the first image according to the target object to obtain an initial template image; an image transformation module, configured to perform image change processing on the initial template image to obtain a first preset number of standard template images, wherein the image change processing includes image scaling and / or image rotation; a semantic segmentation module, configured to input the second image into a pre-built semantic model to perform semantic segmentation on scene objects therein, and obtain a contour image of the scene objects associated with the target object; An image synthesis module is used to superimpose the standard template image on a specified area of the second image based on the contour image, the rotation angle and scaling factor of the standard template image and the camera parameter model of the imaging device to obtain a target synthetic image.
10. An electronic device, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1 to 7 when the computer program instructions are executed by the processor.