A real scene image synthesis method and system
By adopting the real scene image synthesis method in the field of object detection and using semantic graphs and 3D models to generate real scene images, the problems of difficulty in building data sets and lack of universality in the prior art are solved, high-quality image synthesis is achieved, and the application scope is expanded.
Patent Information
- Application Number
- CN202111530511.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Existing deep learning-based object detection methods require a large amount of image data and manual annotation. Especially in complex visual tasks, data set construction is difficult, and synthetic data set construction methods lack universality.
By obtaining multiple semantic maps to be synthesized, input them into the trained real scene image synthesis network model, and a real scene synthesis map corresponding to the semantic map to be synthesized is generated. The method includes establishing a 3D model based on the detection object and building a virtual scene, adding random parameters, rendering images using computer graphics technology, outputting training images and labeling information, building a training data set, and training the network model through a gradient descent algorithm.
It realizes the rapid, effective and reliable synthesis of photographic scene images with stronger sense of reality, improves the reality and visual quality of the composite images, and expands the application range and application scenarios.
Smart Images

Figure CN114372940B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular relates to a real scene image synthesis method and system. Background Art
[0002] At present, the target detection technology based on deep learning has developed to a relatively mature level, and many applications based on target detection have been born, such as face recognition, vehicle detection, and autonomous driving. However, there is still a long way to go before the target detection method can be applied to all walks of life, and data is an inevitable problem. The target detection method based on deep learning is data-driven, which requires a large amount of image data and manual annotation. When encountering complex visual tasks, such as the detection objects involving multiple categories, the complex detection environment is not conducive to data collection, and the detection objects involve expert knowledge, it will bring many difficulties to the construction of the data set.
[0003] As a new dataset construction method, synthetic data has been widely used in recent years to evaluate and train deep neural network models, such as MPI-Sintel, SceneNet, GTA-V, Flying Chairs, etc. These synthetic datasets have proved that synthetic data is as powerful as real data, but their methods of constructing datasets are not universal. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for synthesizing real scene images, so that the construction of synthetic images has certain versatility and ease of use.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] The present invention provides a method for synthesizing a real scene image, comprising:
[0007] Obtain multiple semantic graphs to be synthesized;
[0008] Inputting the semantic graph to be synthesized into the trained real scene image synthesis network model to obtain a real scene synthesis image corresponding to the semantic graph to be synthesized;
[0009] The training process of the real scene image synthesis network model includes:
[0010] Build a 3D model and a virtual scene based on the detection object, add random parameters, use computer graphics technology to render the image, output the training image and the corresponding annotation information, and build a training data set;
[0011] The real scene image synthesis network model is trained through the training data set, the loss function of the real scene image synthesis network model is calculated, and the network weight of the loss function when the loss takes an approximate minimum value is obtained through the gradient descent algorithm to obtain the trained real scene image synthesis network model.
[0012] Preferably, the types of random parameters include: the number and position of point light sources, the illumination intensity of ambient light; the position of the virtual camera relative to the detection target; the texture and background of the 3D model; and the number, shape, texture and size of interference objects added to the virtual scene.
[0013] Preferably, the image rendering is performed using computer graphics technology, and the calculation process includes:
[0014] (I, L) = Render (M, R, W, H) # (1)
[0015] Among them, Render is the image rendering function, (I, L) is the output training image and the corresponding annotation information, M is the set of 3D models in the three-dimensional virtual scene, R is the randomized component set of 3D models in the virtual scene, and W and H represent the width and height of the output image respectively.
[0016] Preferably, the real scene image synthesis network model is trained by a training data set, and the process includes:
[0017] Get the top left vertex P of the training image bounding box from the annotation information 1 (X 1 , Y 1 ), the lower right vertex P 2 (X 2 , Y 2 ) composition and category C;
[0018] The width, height, and center point coordinates of the training image bounding box are calculated as the training data set in COCO format. The formula is:
[0019]
[0020]
[0021] W c =X 2 -X 1
[0022] H c =Y 2 -Y 1
[0023] In the formula, X c Represented as the horizontal coordinate of the center point of the training image bounding box; Y cW is represented as the vertical coordinate of the center point of the bounding box of the training image; c Represented as the width of the bounding box of the training image; H c Represented as the height of the bounding box of the training image;
[0024] The real scene image synthesis network model is trained using a COCO format training dataset.
[0025] Preferably, the real scene image synthesis network model is calculated to obtain a loss function, and the process includes:
[0026] The prediction function of the real scene image synthesis network model is:
[0027]
[0028] In the formula, is the prediction function of the real scene image synthesis network model, w is the neural network weight, x is the training image input to the real scene image synthesis network model, and l is the prediction result of the real scene image synthesis network model;
[0029] Then the total error function of the real scene image synthesis network model is:
[0030]
[0031] In the formula, Loss(·) represents the error function between the prediction result and the labeled value of the labeled information, Li represents the labeled value of the labeled information, and D represents the labeled information of the training data set.
[0032] Preferably, when the loss is approximately minimized by the gradient descent algorithm, the network weight of the loss function is obtained, and the process includes:
[0033] The training data set is randomly shuffled and then divided into n batches of m capacity;
[0034] Calculate the average error function gradient of the i-th batch, the calculation formula is:
[0035]
[0036] In the formula, grad i is the average error function gradient of the i-th batch of data, x ij is the jth training image in the i-th batch of data, L ij For the training image x ij The annotation value corresponding to the annotation information;
[0037] Using grad i Update the network weights w of the loss function.
[0038] Preferably, use gradi Update the network weight w of the loss function. The process includes:
[0039] The iterative formula of the network weight w is:
[0040] w i+1 =w i -lr·grad i
[0041] In the formula, lr is the learning rate, w is i is the network weight w corresponding to the i-th batch of data;
[0042] When i>n, the network weight w converges, the network weight w of the loss function is obtained;
[0043] When i>n and the network weight w diverges, the training data set is re-divided to calculate the network weight w of the loss function.
[0044] Another aspect of the present invention provides a real scene image synthesis system, comprising:
[0045] The virtual model building module is used to build a 3D model and a virtual scene according to the detection object, add random parameters, use computer graphics technology to render images, and output training images and corresponding annotation information;
[0046] A first acquisition module is used to acquire the training images and corresponding annotation information output by the virtual model building module to construct a training data set;
[0047] A training module trains the real scene image synthesis network model through a training data set, calculates the real scene image synthesis network model to obtain a loss function, obtains the network weight of the loss function when the loss takes an approximate minimum value through a gradient descent algorithm, and obtains a trained real scene image synthesis network model;
[0048] The second acquisition module is used to acquire multiple semantic graphs to be synthesized;
[0049] A synthesis module is used to input the semantic graph to be synthesized into the trained real scene image synthesis network model to obtain a real scene synthesis image corresponding to the semantic graph to be synthesized.
[0050] Preferably, the virtual model building module includes:
[0051] 3D model building module, used to make 3D models and textures of detection objects;
[0052] Virtual scene building module, which builds virtual scenes for detection objects through the Unreal game engine;
[0053] Random component, used to add random parameters to 3D models.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] In the present invention, a 3D model is established and a virtual scene is constructed according to the detection object, random parameters are added to the 3D model, image rendering is performed using computer graphics technology, training images and corresponding annotation information are output, and a training data set is constructed; the method of constructing a data set in the present invention is faster and more accurate than the traditional method of using a camera to acquire images and annotate them with annotation tools, and can greatly reduce time and labor costs.
[0056] In the present invention, a real scene image synthesis network model is trained by a training data set, the real scene image synthesis network model is calculated to obtain a loss function, and the network weight of the loss function when the loss takes an approximate minimum value is obtained by a gradient descent algorithm to obtain a trained real scene image synthesis network model; the semantic graph to be synthesized is input into the trained real scene image synthesis network model to obtain a real scene synthesis graph corresponding to the semantic graph to be synthesized; the method or system of the present invention can quickly, effectively and reliably synthesize photographic-level scene images with stronger sense of reality, improve the sense of reality and visual quality of the synthesized images, and expand the scope of application and application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A flow chart of a real scene image synthesis method provided by the present invention; DETAILED DESCRIPTION
[0058] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0059] Embodiment 1
[0060] like Figure 1 As shown, a real scene image synthesis method includes: obtaining a plurality of semantic graphs to be synthesized; inputting the semantic graphs to be synthesized into the trained real scene image synthesis network model to obtain a real scene synthesis image corresponding to the semantic graphs to be synthesized;
[0061] The training process of the real scene image synthesis network model includes:
[0062] A 3D model is established and a virtual scene is built according to the detection object, and random parameters are added. The types of random parameters include: the number and position of point light sources, the illumination intensity of ambient light; the position of the virtual camera relative to the detection target; the texture and background of the 3D model; and the number, shape, texture and size of interference objects added to the virtual scene.
[0063] Image rendering is performed using computer graphics technology, training images and corresponding annotation information are output, and a training data set is constructed; wherein, image rendering is performed using computer graphics technology, and the calculation process includes:
[0064] (I, L) = Render (M, R, W, H) # (1)
[0065] Among them, Render is the image rendering function, (I, L) is the output training image and the corresponding annotation information, M is the set of 3D models in the three-dimensional virtual scene, R is the randomized component set of 3D models in the virtual scene, and W and H represent the width and height of the output image respectively.
[0066] The real scene image synthesis network model is trained through the training data set, the loss function of the real scene image synthesis network model is calculated, and the network weight of the loss function when the loss takes an approximate minimum value is obtained through the gradient descent algorithm to obtain the trained real scene image synthesis network model.
[0067] The real scene image synthesis network model is trained through the training data set. The process includes:
[0068] Get the top left vertex P of the training image bounding box from the annotation information 1 (X 1 , Y 1 ), the lower right vertex P 2 (X 2 , Y 2 ) composition and category C;
[0069] The width, height, and center point coordinates of the training image bounding box are calculated as the training data set in COCO format. The formula is:
[0070]
[0071]
[0072] W c =X 2 -X 1
[0073] H c =Y 2 -Y 1
[0074] In the formula, X c Represented as the horizontal coordinate of the center point of the training image bounding box; Y c W is represented as the vertical coordinate of the center point of the bounding box of the training image; c Represented as the width of the bounding box of the training image; H cRepresented as the height of the bounding box of the training image;
[0075] The real scene image synthesis network model is trained using a COCO format training dataset.
[0076] Calculate the real scene image synthesis network model to obtain the loss function. The process includes:
[0077] The prediction function of the real scene image synthesis network model is:
[0078]
[0079] In the formula, is the prediction function of the real scene image synthesis network model, w is the neural network weight, x is the training image input to the real scene image synthesis network model, and l is the prediction result of the real scene image synthesis network model;
[0080] Then the total error function of the real scene image synthesis network model is:
[0081]
[0082] In the formula, Loss(·) represents the error function between the predicted result and the labeled value of the labeled information, and L i It represents the label value of the label information; D represents the label information of the training data set.
[0083] When the loss is approximated to the minimum value through the gradient descent algorithm, the network weight of the loss function is obtained. The process includes:
[0084] The training data set is randomly shuffled and then divided into n batches of m capacity;
[0085] Calculate the average error function gradient of the i-th batch, the calculation formula is:
[0086]
[0087] In the formula, grad i is the average error function gradient of the i-th batch of data, x ij is the jth training image in the i-th batch of data, L ij For the training image x ij The annotation value corresponding to the annotation information;
[0088] Using grad i Update the network weights w of the loss function.
[0089] Using grad i Update the network weight w of the loss function. The process includes:
[0090] The iterative formula of the network weight w is:
[0091] w i+1 =w i -lr·grad i
[0092] In the formula, lr is the learning rate, w is i is the network weight w corresponding to the i-th batch of data;
[0093] When i>n, the network weight w converges, the network weight w of the loss function is obtained;
[0094] When i>n and the network weight w diverges, the training data set is re-divided to calculate the network weight w of the loss function.
[0095] Embodiment 2
[0096] A real scene image synthesis system, comprising:
[0097] The virtual model building module is used to build a 3D model and a virtual scene according to the detection object, add random parameters, use computer graphics technology to render images, and output training images and corresponding annotation information;
[0098] A first acquisition module is used to acquire the training images and corresponding annotation information output by the virtual model building module to construct a training data set;
[0099] A training module trains the real scene image synthesis network model through a training data set, calculates the real scene image synthesis network model to obtain a loss function, obtains the network weight of the loss function when the loss takes an approximate minimum value through a gradient descent algorithm, and obtains a trained real scene image synthesis network model;
[0100] The second acquisition module is used to acquire multiple semantic graphs to be synthesized;
[0101] A synthesis module is used to input the semantic graph to be synthesized into the trained real scene image synthesis network model to obtain a real scene synthesis image corresponding to the semantic graph to be synthesized.
[0102] The virtual model building module includes:
[0103] 3D model building module, used to make 3D models and textures of detection objects;
[0104] Virtual scene building module, which builds virtual scenes for detection objects through the Unreal game engine;
[0105] Random component, used to add random parameters to 3D models.
[0106] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0107] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0108] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0110] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A real scene image synthesis method, It is characterized in that include: Obtain multiple semantic graphs to be synthesized; Inputting the semantic graph to be synthesized into a trained real scene image synthesis network model to obtain a real scene synthesis image corresponding to the semantic graph to be synthesized; The training process of the real scene image synthesis network model includes: Build a 3D model and a virtual scene based on the detection object, add random parameters, use computer graphics technology to render the image, output the training image and the corresponding annotation information, and build a training data set; The real scene image synthesis network model is trained through the training data set. The process includes: Get the top left vertex P of the training image bounding box from the annotation information 1 (X 1 , Y 1 ), the lower right vertex P 2 (X 2 , Y 2 ) composition and category C; The width, height, and center point coordinates of the training image bounding box are calculated as the training data set in COCO format. The formula is: ; ; ; ; In the formula, X c Represented as the horizontal coordinate of the center point of the training image bounding box; Y c W is represented as the vertical coordinate of the center point of the training image bounding box; c Represented as the width of the bounding box of the training image; H c Represented as the height of the bounding box of the training image; The real scene image synthesis network model is trained using the COCO format training dataset; The loss function is obtained by calculating the real scene image synthesis network model, and the network weight of the loss function when the loss takes an approximate minimum value is obtained through the gradient descent algorithm to obtain the trained real scene image synthesis network model.
2. A real scene image synthesis method according to claim 1, It is characterized in that The types of random parameters include: the number and position of point light sources, the illumination intensity of ambient light; the position of the virtual camera relative to the detection target; the texture and background of the 3D model; and the number, shape, texture and size of interference objects added to the virtual scene.
3. A real scene image synthesis method according to claim 1, It is characterized in that Image rendering is performed using computer graphics technology. The calculation process includes: ; Among them, Render is the image rendering function, It is the output training image and the corresponding annotation information, M is the set of 3D models in the three-dimensional virtual scene, R is the randomized component set of 3D models in the virtual scene, and W and H represent the width and height of the output image respectively.
4. The method for synthesizing real scene images according to claim 1, It is characterized in that Calculate the real scene image synthesis network model to obtain the loss function. The process includes: The prediction function of the real scene image synthesis network model is: ; In the formula, is the prediction function of the real scene image synthesis network model, is the neural network weight, To input the real scene image to synthesize the training image of the network model, Synthesize the predictions of the network model for real scene images; Then the total error function of the real scene image synthesis network model is: ; In the formula, Represents the error function between the predicted result and the labeled value of the labeled information, L i It represents the label value of the label information; D represents the label information of the training data set.
5. A real scene image synthesis method according to claim 4, It is characterized in that When the loss is approximated to the minimum value through the gradient descent algorithm, the network weight of the loss function is obtained. The process includes: The training data set is randomly shuffled and then divided into The capacity is Batch data of Calculate the The average error function gradient is calculated as: ; In the formula, It is The average error function gradient of the batch data, For the The first training images, For training images The annotation value corresponding to the annotation information; use Update the network weights of the loss function .
6. A real scene image synthesis method according to claim 5, It is characterized in that use Update the network weights of the loss function The process includes: The network weights have the following iterative formula: ; In the formula, is the learning rate, For the Network weights corresponding to batch data ; when , network weight When the convergence of ; when , network weight When the network weights of the training data set for the loss function diverge, Perform calculations.
7. A real scene image synthesis system, It is characterized in that include: The virtual model building module is used to build a 3D model and a virtual scene according to the detection object, add random parameters, use computer graphics technology to render images, and output training images and corresponding annotation information; A first acquisition module is used to acquire the training images and corresponding annotation information output by the virtual model building module to construct a training data set; A training module trains the real scene image synthesis network model through a training data set, calculates the real scene image synthesis network model to obtain a loss function, obtains the network weight of the loss function when the loss takes an approximate minimum value through a gradient descent algorithm, and obtains a trained real scene image synthesis network model; The second acquisition module is used to acquire multiple semantic graphs to be synthesized; A synthesis module for inputting the semantic graph to be synthesized into the trained real-scene image synthesis network model to obtain a real-scene synthesis graph corresponding to the semantic graph to be synthesized; The training module trains the real-scene image synthesis network model through a training data set, and the process includes: Get the top left vertex P of the training image bounding box from the annotation information 1 (X 1 , Y 1 ), the lower right vertex P 2 (X 2 , Y 2 ) composition and category C; Calculating the width, height and center point coordinates of the training image bounding box as the training data set in COCO format, and the formula is: ; ; ; ; In the formula, X c Represented as the horizontal coordinate of the center point of the training image bounding box; Y c W is represented as the vertical coordinate of the center point of the bounding box of the training image; c Represented as the width of the bounding box of the training image; H c Represented as the height of the bounding box of the training image; Training the real-scene image synthesis network model through the training data set in COCO format.
8. A real-scene image synthesis system according to claim 7, wherein, The virtual model building module includes: A 3D model building module for making a 3D model and texture of the detection object; A virtual scene building module for building a virtual scene for the detection object through the Unreal game engine; A random component for adding random parameters to the 3D model.
Citation Information
Patent Citations
A real scene image synthesis method and system
CN109447897A
Article recognition method and device, vending system, and storage medium
WO2020134102A1