Scene synthesis model training method, scene synthesis method and electronic equipment

By using a training data set and corpus with real images and text description information of preset scenes, the scene synthesis model is trained, which solves the problem of difficult to guarantee the correctness and quality of scene synthesis results in the prior art, and achieves higher training accuracy and generation quality.

CN120032200APending Publication Date: 2025-05-23ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410370530.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the existing intelligent scene synthesis system generates a specific scenario, the accuracy and quality of the synthesis results are difficult to guarantee, and there are problems with distorted information.

Method used

By obtaining the training data set and corpus with the real image of the preset scene and its text description information, the text description information is input to the scene synthesis model to be trained, and the training is stopped until the comparison result between the generated image and the real image meets the preset requirements.

Benefits of technology

The training accuracy of the scene synthesis model is improved, making the generated results more refined and realistic, and avoiding the generation of distorted information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032200A_ABST
    Figure CN120032200A_ABST
Patent Text Reader

Abstract

The invention relates to a scene synthesis model training method, a scene synthesis method and electronic equipment, and the method comprises the steps: obtaining a training data set and a corpus; wherein the training data set comprises a real image with a preset scene, and the corpus comprises text description information of the real image with the preset scene; respectively inputting the text description information of each image in the corpus into a to-be-trained scene synthesis model to obtain a plurality of generated images, and stopping training the to-be-trained scene synthesis model until a comparison result of each generated image and a corresponding real image meets a first preset requirement to obtain a trained scene synthesis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to intelligent network model technology, and more specifically, to a scene synthesis model training method, a scene synthesis method and an electronic device. Background Art

[0002] In recent years, large models such as pre-trained language models have made great progress in the field of natural language processing. Some companies have begun to apply large model technology to realize intelligent scene synthesis systems that automatically generate images based on text descriptions. These systems can synthesize various specific scenes, such as indoor scenes, urban scenes, natural scenes, etc. However, when faced with certain specific scenes, the correctness of the synthesis results still faces challenges.

[0003] Currently, there is a need to provide a training method for a scene synthesis model to improve the quality of generated content and avoid distorted information. Summary of the invention

[0004] An object of the present invention is to provide a new technical solution for a training method of a scene synthesis model.

[0005] According to a first aspect of the present invention, a method for training a scene synthesis model is provided, comprising:

[0006] Acquire a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of the real images with preset scenes;

[0007] The text description information of each image in the corpus is respectively input into the scene synthesis model to be trained to obtain multiple generated images, until the comparison result of each generated image with the corresponding real image meets the first preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0008] Optionally, the method further includes:

[0009] Input the generated image and the corresponding text description information into the language model to obtain the matching degree value between the generated image and the corresponding text description;

[0010] When the matching degree value satisfies the second preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0011] Optionally, the method further includes:

[0012] Build a common sense knowledge graph;

[0013] According to the common sense knowledge graph, detecting whether the relationship between the elements of the generated image conforms to common sense;

[0014] When the relationship between the elements of the generated image conforms to common sense, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0015] Optionally, the method further includes:

[0016] Building a physics engine;

[0017] According to the physical engine, checking whether the motion information of the moving object in the generated image conforms to the laws of physics;

[0018] When the motion information of the moving object in the generated image complies with the physical law, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0019] Optionally, the method further includes:

[0020] Using an image quality evaluation method, checking whether the texture of the generated image meets a third preset requirement;

[0021] When the texture of the generated image meets the third preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0022] Optionally, the method further includes:

[0023] Construct a structured knowledge base of prior knowledge;

[0024] According to the prior knowledge structured knowledge base, checking whether each element in the generated image meets the prior knowledge constraints;

[0025] When each element in the generated image meets the prior knowledge constraint, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0026] Optionally, the scene synthesis model to be trained is a generative adversarial network model.

[0027] According to a second aspect of the present invention, there is provided a scene synthesis method, comprising:

[0028] Obtain text description information of the scene to be synthesized;

[0029] The text description information of the scene to be synthesized is input into a trained scene synthesis model obtained by any method described in the first aspect to obtain an image of the scene to be synthesized.

[0030] According to a third aspect of the present invention, there is provided a training device for a scene synthesis model, comprising:

[0031] An acquisition module, used to acquire a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of the real images with preset scenes;

[0032] A training module is used to input the text description information of each image in the corpus into the scene synthesis model to be trained, respectively, to obtain multiple generated images, until the comparison result of each generated image with the corresponding real image meets the first preset requirement, then stop training the scene synthesis model to be trained, and obtain a trained scene synthesis model.

[0033] According to a fourth aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program for controlling the processor to operate to execute the method according to any one of the first aspect or the second aspect of the present invention.

[0034] The training method of the scene synthesis model provided by the present invention obtains a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of real images with preset scenes, and the text description information of each image in the corpus is respectively input into the scene synthesis model to be trained to obtain multiple generated images, until the comparison result of each generated image with the corresponding real image meets the first preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model, thereby improving the training accuracy of the scene synthesis model and making the generation result of the trained scene synthesis model more refined and realistic.

[0035] Features and advantages of the embodiments of the present specification will become apparent from the following detailed description of exemplary embodiments of the present specification with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the specification and, together with the description, serve to explain the principles of the embodiments of the specification.

[0037] Figure 1 It is a schematic diagram of the processing flow of a method for training a scene synthesis model according to an embodiment of the present invention.

[0038] Figure 2 4 is a schematic diagram of a processing flow of a scene synthesis method according to an embodiment of the present invention.

[0039] Figure 3 It is a principle block diagram of a training device for a scene synthesis model according to an embodiment of the present invention.

[0040] Figure 4FIG. 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings.

[0042] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0043] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0044] <Method Example>

[0045] In this embodiment, a method for training a scene synthesis model is provided. Figure 1 As shown, the training method of the scene synthesis model of this embodiment may include the following steps S110 to S120.

[0046] Step S110, obtaining a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of real images with preset scenes.

[0047] The real image with the preset scene is an image taken by the user according to his / her own needs. The preset scene can be a science and education live broadcast scene or a sports event live broadcast scene.

[0048] The text description information of the real image with the preset scene is information determined by the user based on the image taken according to the user's own needs.

[0049] Step S120, respectively input the text description information of each image in the corpus into the scene synthesis model to be trained to obtain multiple generated images, until the comparison result of each generated image with the corresponding real image meets the first preset requirement, stop training the scene synthesis model to be trained, and obtain a trained scene synthesis model.

[0050] When the comparison result between each generated image and the corresponding real image does not meet the first preset requirement, the parameters of the scene synthesis model to be trained are adjusted, and the text description information of each image in the corpus continues to be input into the scene synthesis model to be trained to obtain multiple generated images. When the comparison result between each generated image and the corresponding real image meets the first preset requirement, the training of the scene synthesis model to be trained is stopped.

[0051] In one embodiment, the scene synthesis model to be trained is a generative adversarial network model. The generative adversarial network model includes a generator and a discriminator. The generator is used to generate a corresponding generated image according to the text description information of each image. The discriminator is used to compare each generated image with the corresponding real image to obtain a comparison result, and determine whether the comparison result of each generated image with the corresponding real image meets the first preset requirement according to the comparison result.

[0052] Using the text description information of the real image with the preset scene to train the generative adversarial network model is a process of adjusting the parameters of the generator. When the comparison result between the generated image generated by the generator and the corresponding real image meets the first preset requirement, the training of the generator is stopped.

[0053] It should be noted that the scene synthesis model to be trained can also be any published open source model.

[0054] The scene synthesis model training method provided by the present invention improves the training accuracy of the scene synthesis model, making the generation result of the trained scene synthesis model more precise and realistic.

[0055] In one embodiment of the present invention, before the text description information of each image in the corpus is respectively input into the scene synthesis model to be trained to obtain multiple generated images, the text description information of each image is input into a pre-trained language model to obtain the elements in the text description information of each image and the attribute information corresponding to the elements. The elements in the text description information of each image and the attribute information corresponding to the elements are input into the scene synthesis model to be trained. Through the pre-trained language model, key information in the text description information, such as element category, position relationship, and scene layout, can be accurately captured, so that images that meet semantic requirements can be generated, making the generated images more refined and realistic.

[0056] In one embodiment of the present invention, when the comparison results between each generated image and the corresponding real image meet the first preset requirement, the method further includes: inputting the generated image and the corresponding text description information into a language model to obtain a matching degree value between the generated image and the corresponding text description; when the matching degree value meets the second preset requirement, stopping the training of the scene synthesis model to be trained to obtain a trained scene synthesis model.

[0057] Specifically, the language model is used to extract the first semantic information from the generated image and the second semantic information from the text description information. A vector corresponding to the first semantic information is obtained based on the first semantic information, and a vector corresponding to the second semantic information is obtained based on the second semantic information. The cosine value of the vector corresponding to the first semantic information and the vector corresponding to the second semantic information is calculated, and the cosine value is used as the matching degree value between the generated image and the corresponding text description. It is determined whether the matching degree value meets the second preset requirement. When the matching degree value meets the second preset requirement, the training of the scene synthesis model to be trained is stopped to obtain the trained scene synthesis model.

[0058] In this embodiment, a language model is used to obtain a matching degree value between a generated image and a corresponding text description. When the matching degree value meets a second preset requirement, the training of the scene synthesis model to be trained is stopped, and the trained scene synthesis model is obtained and tested from the aspect of semantic consistency, thereby improving the training accuracy of the scene synthesis model.

[0059] In one embodiment of the present invention, when the comparison result of each generated image and the corresponding real image meets the first preset requirement, the target detection algorithm is used to identify the attribute information of each element in the generated image. The attribute information includes the name, position, quantity, and color of the object. The attribute information of each element is compared with the text description information to obtain the degree of coincidence between the two. When the degree of coincidence meets the second preset requirement, the training of the scene synthesis model to be trained is stopped to obtain the trained scene synthesis model.

[0060] In this embodiment, a target detection algorithm is used to identify the attribute information of each element in a generated image, and the attribute information of each element is compared with the text description information to obtain a matching degree value between the two. When the matching degree value meets the second preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model, which is tested from the aspect of semantic consistency, thereby improving the training accuracy of the scene synthesis model.

[0061] In one embodiment of the present invention, when the comparison results between each generated image and the corresponding real image meet the first preset requirement, the method also includes: constructing a common sense knowledge graph; based on the common sense knowledge graph, detecting whether the relationship between the elements of the generated image conforms to common sense; when the relationship between the elements of the generated image conforms to common sense, stopping the training of the scene synthesis model to be trained, and obtaining a trained scene synthesis model.

[0062] The common sense knowledge graph is an implicit logical relationship graph determined by learning common sense based on a large model.

[0063] The common sense knowledge graph is used to determine whether the relationship between elements in the generated image conforms to common sense. The elements in the generated image include people and objects.

[0064] For example, based on the common sense knowledge graph, it is determined whether the identity and posture of the person in the generated image are consistent. For another example, based on the common sense knowledge graph, it is determined whether the position of the object in the generated image is consistent with the common sense of use.

[0065] In this embodiment, the common sense knowledge graph is used to detect whether the relationship between the elements of the generated image conforms to common sense. The naturalness and authenticity of the details of the generated image are tested, thereby improving the training accuracy of the scene synthesis model.

[0066] In one embodiment of the present invention, when the comparison results between each generated image and the corresponding real image meet the first preset requirement, the method also includes: building a physical engine; based on the physical engine, checking whether the motion information of the moving object in the generated image conforms to the physical laws; when the motion information of the moving object in the generated image conforms to the physical laws, stopping the training of the scene synthesis model to be trained, and obtaining the trained scene synthesis model.

[0067] The physics engine is specifically used for the inspection of generated images of moving objects involving physical motion.

[0068] For example, when the generated image includes a moving sphere, the physical engine is used to check whether the moving trajectory of the moving sphere complies with the laws of physics.

[0069] Specifically, the motion information of the moving object is extracted from the generated image, and the motion information of the moving object in the generated image is checked according to the physical engine to see whether it conforms to the physical law.

[0070] In this embodiment, based on the physical engine, it is checked whether the motion information of the moving object in the generated image conforms to the physical laws. When the motion information of the moving object in the generated image conforms to the physical laws, the training of the scene synthesis model to be trained is stopped to obtain the trained scene synthesis model. The test is carried out from the aspect of the physical laws of motion, thereby improving the training accuracy of the scene synthesis model.

[0071] In one embodiment of the present invention, when the comparison results between each generated image and the corresponding real image meet the first preset requirement, the method also includes: using an image quality evaluation method to verify whether the texture of the generated image meets the third preset requirement; when the texture of the generated image meets the third preset requirement, stopping the training of the scene synthesis model to be trained to obtain a trained scene synthesis model.

[0072] In this embodiment, an image quality evaluation method is used to check whether the texture of the generated image meets the third preset requirement. The naturalness and authenticity of the details of the generated image are checked to improve the training accuracy of the scene synthesis model.

[0073] In one embodiment of the present invention, a physical engine is used to check whether the global layout of the generated image is reasonable. When it is determined that the global layout of the generated image is reasonable, the training of the scene synthesis model to be trained is stopped to obtain a trained synthesis model.

[0074] In one embodiment of the present invention, when the comparison results between each generated image and the corresponding real image meet the first preset requirement, the method also includes: constructing a priori knowledge structured knowledge base; based on the priori knowledge structured knowledge base, checking whether each element in the generated image meets the prior knowledge constraints; when each element in the generated image meets the prior knowledge constraints, stopping the training of the scene synthesis model to be trained, and obtaining a trained scene synthesis model.

[0075] A priori knowledge structured knowledge base is a knowledge base established based on empirical knowledge (for example, social common sense).

[0076] For example, based on the prior knowledge structured knowledge base, check whether the clothing of the characters in the generated image matches the era background information in the generated image. For another example, based on the prior knowledge structured knowledge base, check whether the buildings in the generated image match the geographical information in the generated image.

[0077] In this embodiment, based on the structured knowledge base of prior knowledge, it is checked whether each element in the generated image meets the prior knowledge constraints. When each element in the generated image meets the prior knowledge constraints, the training of the scene synthesis model to be trained is stopped to obtain the trained scene synthesis model. The test is carried out from the aspect of prior knowledge constraints, thereby improving the training accuracy of the scene synthesis model.

[0078] An embodiment of the present invention provides a scene synthesis method, according to Figure 2 As shown, the scene synthesis method includes steps S210 to S220.

[0079] Step S210: obtaining text description information of the scene to be synthesized.

[0080] Step S220: input the text description information of the scene to be synthesized into a trained scene synthesis model obtained by any of the above methods to obtain an image of the scene to be synthesized.

[0081] The training process of the trained scene synthesis model can refer to the training steps provided in any of the above embodiments, and will not be described in detail here.

[0082] <Device Example>

[0083] One embodiment of the present invention provides a training device for a scene synthesis model. Figure 3 As shown, the scene synthesis model training device 300 includes an acquisition module 310 and a training module 320.

[0084] The acquisition module 310 is used to acquire a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of real images with preset scenes.

[0085] The training module 320 is used to input the text description information of each image in the corpus into the scene synthesis model to be trained, respectively, to obtain multiple generated images, until the comparison result of each generated image with the corresponding real image meets the first preset requirement, then stop training the scene synthesis model to be trained, and obtain a trained scene synthesis model.

[0086] In one embodiment of the present invention, the training device 300 of the scene synthesis model also includes an extraction module. The extraction module is used to input the text description information of each image into the pre-trained language model to obtain the elements in the text description information of each image and the attribute information corresponding to the elements. The elements in the text description information of each image and the attribute information corresponding to the elements are input into the scene synthesis model to be trained. Through the pre-trained language model, key information in the text description information, such as element category, position relationship, and scene layout, can be accurately captured, so that images that meet semantic requirements can be generated, making the generated images more refined and realistic.

[0087] In one embodiment of the present invention, the training module 320 is also used to generate an image and corresponding text description information and input it into a language model to obtain a matching degree value between the generated image and the corresponding text description; when the matching degree value meets the second preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0088] In one embodiment of the present invention, the training module 320 is also used to construct a common sense knowledge graph; based on the common sense knowledge graph, it is detected whether the relationship between the elements of the generated image conforms to common sense; when the relationship between the elements of the generated image conforms to common sense, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0089] In one embodiment of the present invention, the training module 320 is also used to build a physical engine; based on the physical engine, check whether the motion information of the moving objects in the generated image conforms to the physical laws; when the motion information of the moving objects in the generated image conforms to the physical laws, stop training the scene synthesis model to be trained to obtain a trained scene synthesis model.

[0090] In one embodiment of the present invention, the training module 320 is also used to use an image quality evaluation method to verify whether the texture of the generated image meets the third preset requirement; when the texture of the generated image meets the third preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0091] In one embodiment of the present invention, the training module 320 is also used to construct a structured knowledge base of prior knowledge; based on the structured knowledge base of prior knowledge, it is checked whether each element in the generated image meets the prior knowledge constraints; when each element in the generated image meets the prior knowledge constraints, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

[0092] In one embodiment of the present invention, the scene synthesis model to be trained is a generative adversarial network model.

[0093] Figure 4 4 is a schematic diagram of an electronic device 400 provided by an embodiment of the present disclosure. The electronic device 400 includes a processor 410 and a memory 420. The memory 420 stores computer instructions. When the computer instructions are executed by the processor 410, the method disclosed in any of the above embodiments is implemented.

[0094] The embodiments of the present disclosure further provide a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in any of the aforementioned embodiments is implemented.

[0095] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For the electric vehicle embodiment, its related parts can be referred to the partial description of the method embodiment.

[0096] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0097] The embodiments of the present specification may be systems, methods and / or computer program products. The computer program product may include a computer-readable storage medium carrying computer instructions for causing a processor to implement various aspects of the embodiments of the present specification.

[0098] A computer-readable storage medium can be a tangible device that can hold and store computer instructions for use by a computer instruction execution device. The computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having computer instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0099] The computer instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer instructions from the network and forwards the computer instructions for storage in a computer-readable storage medium in each computing / processing device.

[0100] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to the multiple embodiments of this specification. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a computer instruction, and a part of a module, a program segment or a computer instruction contains one or more executable computer instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.

[0101] The embodiments of the present specification have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for training a scene synthesis model, characterized in that: include: Acquire a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of the real images with preset scenes; The text description information of each real image in the corpus is respectively input into the scene synthesis model to be trained to obtain multiple generated images, until the comparison result between each generated image and the corresponding real image meets the first preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

2. The method according to claim 1, characterized in that The method further comprises: Input the generated image and the corresponding text description information into the language model to obtain the matching degree value between the generated image and the corresponding text description; When the matching degree value satisfies the second preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

3. The method according to claim 1 or 2, characterized in that: The method further comprises: Build a common sense knowledge graph; According to the common sense knowledge graph, detecting whether the relationship between the elements of the generated image conforms to common sense; When the relationship between the elements of the generated image conforms to common sense, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

4. The method according to claim 1 or 2, characterized in that: The method further comprises: Building a physics engine; According to the physical engine, checking whether the motion information of the moving object in the generated image conforms to the laws of physics; When the motion information of the moving object in the generated image complies with the physical law, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

5. The method according to claim 1 or 2, characterized in that: The method further comprises: Using an image quality evaluation method, checking whether the texture of the generated image meets a third preset requirement; When the texture of the generated image meets the third preset requirement, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

6. The method according to claim 1 or 2, characterized in that: The method further comprises: Construct a structured knowledge base of prior knowledge; According to the prior knowledge structured knowledge base, checking whether each element in the generated image meets the prior knowledge constraints; When each element in the generated image meets the prior knowledge constraint, the training of the scene synthesis model to be trained is stopped to obtain a trained scene synthesis model.

7. The method according to claim 1 or 2, characterized in that: The scene synthesis model to be trained is a generative adversarial network model.

8. A scene synthesis method, characterized in that: include: Obtain text description information of the scene to be synthesized; The text description information of the scene to be synthesized is input into the trained scene synthesis model obtained by the method as described in any one of claims 1 to 7 to obtain an image of the scene to be synthesized.

9. A training device for a scene synthesis model, characterized in that: include: An acquisition module, used to acquire a training data set and a corpus; wherein the training data set includes real images with preset scenes, and the corpus includes text description information of the real images with preset scenes; A training module is used to input the text description information of each real image in the corpus into the scene synthesis model to be trained, respectively, to obtain multiple generated images, until the comparison result of each generated image with the corresponding real image meets the first preset requirement, then stop training the scene synthesis model to be trained, and obtain a trained scene synthesis model.

10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer instructions, and the computer instructions implement the method according to any one of claims 1 to 8 when executed by the processor.