Method and apparatus for generating picture, computing device and computer program product

By using multiple neural network models to identify and adjust the risks in prompt words during the image generation process, the problem of difficulty in ensuring the security of large models when generating pictures is solved, and higher risk identification accuracy and image security are achieved.

CN120032002APending Publication Date: 2025-05-23ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510123020.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When using large models to generate pictures, it is difficult for the prior art to effectively identify and avoid risks such as improper content, privacy leakage and false information, making it difficult to guarantee the security of the generated pictures.

Method used

The first neural network model is used to determine whether there is a risk in the prompt words entered by the user. If there is, the prompt words are adjusted through the second neural network model to remove the risk. Finally, the third neural network model generates a picture based on the adjusted prompt words.

Benefits of technology

It improves the accuracy of risk identification, enhances the security of generated images, and avoids the occurrence of improper content, privacy leakage and false information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032002A_ABST
    Figure CN120032002A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for generating pictures, computing equipment and a computer program product. According to the method, firstly, whether risks exist in cue words input by a user or not is judged through a text risk discriminator, and if the risks exist, the cue words are modified through a cue word rewriter, so that the risks in the cue words are removed. And then, generating a corresponding picture based on the adjusted prompt word through a picture generator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification generally relate to the field of computers, and more particularly to methods, apparatuses, computing devices, and computer program products for generating images. Background Art

[0002] Image generation tasks refer to the process of automatically creating image content from existing data through computer models and algorithms. Taking the text-generated image task as an example, it usually uses large model technology to generate corresponding images based on input text descriptions, images or other content.

[0003] With the development and application of artificial intelligence, models with large parameters, such as diffusion models, are increasingly used in image generation tasks. However, corresponding to the development of artificial intelligence technology, risks such as inappropriate content, privacy leakage, and false information are also increasing. Therefore, how to prevent and avoid risks is an important issue in the field of artificial intelligence. Summary of the invention

[0004] The embodiments of the present specification provide a method, an apparatus, a computing device, and a computer program product for generating a picture.

[0005] In a first aspect of the present specification, a method for generating an image is provided, comprising: determining, through a first neural network model, whether there is a risk in a prompt word input by a user; in response to the presence of a risk in the prompt word, adjusting the prompt word through a second neural network model to remove the existing risk to obtain a target prompt word; and generating a corresponding image based on the target prompt word through a third neural network model.

[0006] In a second aspect of the present specification, a device for generating a picture is provided, comprising: a risk determination module, configured to determine whether there is a risk in a prompt word input by a user through a first neural network model; a prompt word adjustment module, configured to adjust the prompt word through a second neural network model to remove the existing risk in response to the existence of a risk in the prompt word to obtain a target prompt word; and a picture generation module, configured to generate a corresponding picture based on the target prompt word through a third neural network model.

[0007] In a third aspect of the present specification, a computing device is provided, comprising: a processor; and a memory coupled to the processor, the memory having instructions stored therein, and when the instructions are executed by the processor, the computing device executes the method according to the first aspect of the present specification.

[0008] In a fourth aspect of the present specification, a computer program product is provided, comprising a computer program, wherein the computer program is executed by a processor to implement the method according to the first aspect of the present specification.

[0009] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of this specification, nor are they intended to limit the scope of this disclosure. Other features of this specification will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present specification will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram illustrating an example environment in which some embodiments of the present specification may be implemented;

[0012] Figure 2 A flowchart of a method for generating a picture according to some embodiments of the present specification is shown;

[0013] Figure 3 A schematic diagram showing a first process for generating a picture according to some embodiments of the present specification;

[0014] Figure 4 A schematic diagram showing a second process for generating a picture in some embodiments of the present specification;

[0015] Figure 5 A schematic diagram showing a third process for generating a picture according to some embodiments of the present specification; and

[0016] Figure 6 A schematic block diagram of a computing device is shown that illustrates some embodiments of the present description.

[0017] Throughout the drawings, the same or similar reference numbers refer to the same or similar elements. DETAILED DESCRIPTION

[0018] The embodiments of the present specification will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present specification are shown in the accompanying drawings, it should be understood that the present specification can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present specification. It should be understood that the drawings and embodiments of the present specification are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of the embodiments of this specification, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0020] As mentioned above, risk prevention is an important issue in the field of artificial intelligence. With the development of deep learning and large model technology, especially the widespread application of neural network models with Transformer architecture, the ability to generate multimedia content such as images, text, audio, and video through large models has made great progress. Among them, represented by text-to-image technology, corresponding pictures can be automatically generated through text descriptions provided by users, greatly expanding the application scenarios of creative expression. As the ability of large models to generate content improves, the probability of risks such as inappropriate content, privacy leakage, and false information also increases. Therefore, how to control the security of generated content has become a difficult problem in the field of artificial intelligence.

[0021] In the related art, a predefined set of rules is used to predict the risk of the text entered by the user, and the text entered by the user is dynamically adjusted to make it conform to the rules. For example, multiple risk words can be predefined and matched in the text entered by the user to confirm whether there are risk words. However, the rules in the related art are all customized manually and cannot cover all risk scenarios, and the rule base needs to be updated frequently. In addition, due to the fixed rules, the handling of complex and emerging risks is not flexible enough, and there are cases of risk misjudgment. Therefore, the accuracy of risk identification in the related art is low.

[0022] To this end, a method for generating a picture is provided in an embodiment of the present specification. First, a text risk discriminator is used to determine whether there is a risk in the prompt word input by the user. If there is a risk, the prompt word is modified by a prompt word rewriter to remove the risk in the prompt word. Then, a picture generator is used to generate a corresponding picture based on the adjusted prompt word. In this way, the prompt word risk can be identified and the prompt word can be rewritten by the text risk discriminator and the prompt word rewriter, without the need to manually set various risk identification rules, thereby improving the accuracy of risk identification and thereby improving the safety of the generated picture.

[0023] Figure 1 A schematic diagram of an example environment 100 is shown in which some embodiments of the present specification may be implemented. Figure 1The example environment 100 includes a terminal 102 and a server 104. The terminal 102 includes but is not limited to a mobile phone, a tablet computer, a personal computer and other devices, and the server 104 includes but is not limited to a single server, a distributed server, or a cloud-based server. The terminal 102 can communicate with the server 104 in a wired or wireless manner.

[0024] In some embodiments, a text risk discriminator 110 (which may be referred to as a first neural network model), a prompt word rewriter 112 (which may be referred to as a second neural network model), an image generator 114 (which may be referred to as a third neural network model), and an image risk discriminator 116 (which may be referred to as a fourth neural network model) are deployed in the server 104. The text risk discriminator 110, the prompt word rewriter 112, the image generator 114, and the image risk discriminator 116 may be independent neural network models or different parts of an integrated neural network model. In addition, any one of the text risk discriminator 110, the prompt word rewriter 112, the image generator 114, and the image risk discriminator 116 may run in the same service device or in a distributed manner on multiple service devices.

[0025] In some embodiments, the text risk discriminator 110 is a neural network model based on natural language processing and machine learning, which is used to identify and evaluate risk factors in the prompt word 106. The text risk discriminator 110 may include multiple components to identify and analyze different types of risks. In addition, the text risk discriminator 110 in this embodiment can also combine multimodal data (such as text, audio, image, video) to identify risks. The prompt word rewriter 112 is a text optimization model based on natural language processing, which can rewrite, expand and optimize the text in the prompt word 106 so that the rewritten prompt word 106 can meet security requirements. The prompt word rewriter 112 can be implemented based on recurrent neural networks, long short-term memory networks, transformers, etc.

[0026] In some embodiments, the image generator 114 is a neural network model that generates images based on text descriptions, which is trained with a large amount of text and image data so as to learn the potential relationship between text descriptions and images. The image generator 114 can be implemented based on generative adversarial networks, variational autoencoders, and diffusion models. The image risk discriminator 116 is a neural network model for processing image data and identifying image risks, which can analyze potential risks through features such as texture and edges in the image. Similar to the text risk discriminator 110, the image risk discriminator 116 may also include multiple components to identify different types of risks.

[0027] In some embodiments, the user may input text through the terminal 102 and form a prompt word 106, and the terminal 102 sends the prompt word 106 to the server 104. After receiving the prompt word 106, the server 104 may input the prompt word 106 into a text risk discriminator 110, so as to identify whether there is a risk in the prompt word 106 through the text risk discriminator 110. If there is a risk in the prompt word 106, the server 104 inputs the prompt word 106 into a prompt word rewriter 112, so that the prompt word rewriter 112 rewrites the prompt word 106 to remove the risk therein.

[0028] Further, the server 104 inputs the rewritten prompt word into the picture generator 114 to generate the picture 108. After the server 104 generates the picture 108, the picture 108 can be sent to the terminal 102, and the terminal 102 can display it to the user on the screen after receiving the picture 108. Alternatively or additionally, the server 104 inputs the rewritten prompt word into the picture generator 114 to generate an initial picture, and the server 104 inputs the initial picture into the picture risk discriminator 116 to determine whether the initial picture has a risk. If there is no risk, the server 104 uses the initial picture as the final picture 108 to be generated and sends it to the terminal 102, and the terminal 102 can display it to the user on the screen after receiving the picture 108.

[0029] In this embodiment, the text risk discriminator 110 is used to determine whether there is a risk in the prompt word 106 input by the user. If there is a risk, the prompt word 106 is modified by the prompt word rewriter 112 to remove the risk in the prompt word 106. Then, the corresponding image 108 is generated based on the adjusted prompt word by the image generator 114. In this way, the risk of the prompt word 106 can be identified and the prompt word 106 can be rewritten by the text risk discriminator 110 and the prompt word rewriter 112, without manually setting various risk identification rules, thereby improving the accuracy of risk identification and further improving the safety of the generated image 108.

[0030] It should be understood that the architecture and functions in the example environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present description. The embodiments of the present description may also be applied to other environments with different structures and / or functions.

[0031] The following will combine Figures 2 to 6 The process according to the embodiment of this specification is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It should be understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0032] Figure 2 FIG. 2 is a flowchart of a method 200 for generating a picture according to some embodiments of the present specification. Figure 1 In the example environment 100 shown in FIG. 1 , the method 200 may be performed by the server 104. It should be understood that although the following content is described with the server 104 as the execution subject, the method 200 may also be performed by other devices. The method 200 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0033] At 202, a first neural network model is used to determine whether there is a risk in the prompt word input by the user. In some embodiments, the first neural network model may be a text risk discriminator 110. After receiving the prompt word 106 from the terminal 102, the server 110 may input the prompt word 106 into the text risk discriminator 110, so as to identify whether there is a risk in the prompt word 106 through the text risk discriminator 110.

[0034] In some embodiments, the server 104 obtains a prompt word input by a user; obtains context information related to the prompt word from a database, where the database is located in a local or remote server; and determines, through a first neural network model, whether there is a risk in the prompt word based on the prompt word and the context information.

[0035] In some embodiments, the server 104 obtains a first data node related to a prompt word input by a user in an ontology database; determines a second data node associated with the first data based on the relationship between the data in the ontology database; and generates a knowledge graph based on the first data node and the second data node.

[0036] At 204, if there is a risk in the prompt word, the prompt word is adjusted by the second neural network model to remove the risk to obtain a target prompt word. In some embodiments, the second neural network model may be a prompt word rewriter 112. If the text risk discriminator 110 determines that there is a risk in the prompt word 106, the server 104 inputs the prompt word 106 into the prompt word rewriter 112, so that the prompt word 106 is rewritten by the prompt word rewriter 112 to remove the risk therein, thereby obtaining a modified prompt word (which may be referred to as a target prompt word). If the text risk discriminator 110 determines that there is no risk in the prompt word 106, the server may determine that the prompt word 106 at this time does not need to be rewritten by the prompt word rewriter 112.

[0037] In some embodiments, in the current iteration, server 104 determines, through a first neural network model, whether there is a risk in the first prompt word, and the first prompt word is related to the prompt word; and in response to the presence of a risk in the first prompt word, adjusts the first prompt word through a second neural network model to obtain a second prompt word for the next iteration, and the second prompt word is related to the target prompt word.

[0038] At 206, a corresponding image is generated based on the target prompt word through the third neural network model. In some embodiments, the third neural network model can be the image generator 114. After obtaining the rewritten prompt word through the prompt word rewriter 112, the server 104 can input the rewritten prompt word into the image generator 114 to generate the image 108.

[0039] In some embodiments, the server 104 determines whether there is a risk in the picture corresponding to the target prompt word through a fourth neural network model; in response to the existence of a risk in the picture corresponding to the target prompt word, adjusts the target prompt word through a second neural network model; and generates a corresponding picture based on the adjusted target prompt word through a third neural network model.

[0040] In some embodiments, in the current iteration, server 104 determines, through a fourth neural network model, whether there is a risk in the picture corresponding to the third prompt word, and the third prompt word is related to the target prompt word; in response to the existence of a risk in the picture corresponding to the third prompt word, the third prompt word is adjusted through the second neural network model to obtain a fourth prompt word; and a corresponding picture is generated based on the fourth prompt word through the third neural network model, and the picture corresponding to the fourth prompt word is used for the next iteration.

[0041] In some embodiments, the server 104 determines a penalty coefficient for the target prompt word in response to the presence of risk in the picture corresponding to the target prompt word; determines a reward coefficient for the target prompt word in response to the absence of risk in the picture corresponding to the target prompt word; and adjusts the target prompt word based on the penalty coefficient or the reward coefficient through a second neural network model.

[0042] In some embodiments, server 104 obtains a user's rating of a picture corresponding to a target prompt word, where the rating indicates a risk level of the picture corresponding to the target prompt word; in response to the rating being higher than a rating threshold, adjusts the target prompt word through a second neural network model; and generates a corresponding picture based on the adjusted target prompt word through a third neural network model.

[0043] In the embodiments of this specification, a text risk discriminator is used to determine whether there is a risk in the prompt word input by the user. If there is a risk, the prompt word is modified by a prompt word rewriter to remove the risk in the prompt word. Then, a corresponding image is generated based on the adjusted prompt word by an image generator. In this way, the prompt word risk can be identified and the prompt word can be rewritten by a text risk discriminator and a prompt word rewriter, without the need to manually set various risk identification rules, thereby improving the accuracy of risk identification and thereby improving the safety of the generated image.

[0044] Figure 3 FIG. 3 is a schematic diagram showing a first process 300 for generating a picture in some embodiments of the present specification. Figure 1 In the example environment 100 shown in FIG. 1 , the process 300 may be performed by the server 104. It should be understood that although the following content is described with the server 104 as the execution subject, the process 300 may also be performed by other devices. The process 300 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0045] In some embodiments, the server 104 obtains the prompt word 106 input by the user, and evaluates the prompt word 106 through the text risk discriminator 110 to determine whether there is a risk, thereby deciding whether to modify the prompt word 106. If there is a risk in the prompt word 106, the server 104 inputs the prompt word 106 into the prompt word rewriter 112, and modifies the prompt word 106 through the prompt word rewriter 112 to remove the risk factor therein, so as to obtain a modified prompt word (which may be called a target prompt word). Then, the server 104 inputs the modified prompt word into the image generator 114. If there is no risk in the prompt word 106, the server 104 may not need to modify the prompt word 106, and directly input the prompt word 106 into the image generator 114.

[0046] Alternatively or additionally, if the server 104 determines that there is no risk in the prompt word 106 through the text risk discriminator 110, the prompt word text may be further optimized through the prompt word rewriter 112. Then, the server 104 may input the optimized prompt word text into the image generator 114 to generate an image that better meets the user's expectations.

[0047] Then, the server 104 generates a corresponding image based on the adjusted prompt word (when there is a risk in the prompt word 106) or the prompt word 106 (when there is no risk in the prompt word 106) through the image generator 114. Further, the server 104 inputs the generated image into the image risk discriminator 116, so as to determine whether there is a risk in the generated image through the image risk discriminator 116. If there is a risk in the generated image, the server 104 further adjusts the previously adjusted prompt word (when there is a risk in the prompt word 106) or adjusts the prompt word 106 (when there is no risk in the prompt word 106) through the prompt word rewriter 112, and inputs the adjusted prompt word into the image generator 114 again, so as to generate the image 108 again through the image generator 114. If there is no risk in the generated image, the server 104 can directly use the generated image as the final image 108.

[0048] In this way, the generated picture 108 can be subjected to dual risk identification based on the text risk discriminator 110 and the image risk discriminator 116, thereby improving the accuracy of risk identification and further ensuring the safety of the picture 108. In addition, through the prompt word rewriter 112, the prompt word can be modified when there is a risk in the prompt word or the generated picture, thereby improving the efficiency of removing risks and optimizing the prompt word, thereby improving the efficiency of generating safe pictures.

[0049] Figure 4 FIG. 4 is a schematic diagram showing a second process 400 for generating a picture according to some embodiments of the present specification. Figure 1 In the example environment 100 shown in FIG. 1 , process 400 may be performed by server 104. It should be understood that although the following content is described with server 104 as the execution subject, process 400 may also be performed by other devices. Process 400 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0050] In some embodiments, the server 104 obtains the prompt word 106 input by the user, and evaluates the prompt word 106 through the text risk discriminator 110 to determine whether there is a risk, thereby deciding whether to modify the prompt word 106. If there is a risk in the prompt word 106, the server 104 inputs the prompt word 106 into the prompt word rewriter 112, and modifies the prompt word 106 through the prompt word rewriter 112 to remove the risk factor therein to obtain a modified prompt word. Then, the server 104 inputs the modified prompt word into the text risk discriminator 110 again, thereby determining again whether there is a risk in the modified prompt word. If there is still a risk in the modified prompt word, the above process is repeated until there is no risk in the modified prompt word, and the server 104 inputs the prompt word after multiple iterations (which can be called the target prompt word) into the image generator 114.

[0051] Note that after N iterations, in the i-th iteration of the N iterations, the text risk discriminator 110 is used to evaluate whether there is a risk in the prompt word i (which can be called the first prompt word). If there is a risk in the prompt word i, the server 104 can input the prompt word i into the prompt word rewriter 112, and the prompt word rewriter 112 can modify the prompt word i to remove the risk factor therein to obtain the prompt word (i+1). Then, the server 104 assigns (i+1) to i, and continues to the next iteration until a prompt word N (i.e., the target prompt word) without risk is obtained.

[0052] Then, the server 104 generates a corresponding image based on the prompt word N (i.e., the target prompt word) modified through multiple iterations through the image generator 114. Further, the server 104 inputs the generated image into the image risk discriminator 116, so as to determine whether there is a risk in the generated image through the image risk discriminator 116. If there is a risk in the generated image, the server 104 further adjusts the prompt word N modified through multiple iterations through the prompt word rewriter 112, and inputs the adjusted prompt word into the image generator 114 again, so as to generate the image again through the image generator 114. If there is still a risk in the regenerated image, the server 104 repeats the above process until a risk-free image is generated after multiple iterations.

[0053] Note that after M iterations, in the jth iteration of the M iterations, the server 104 determines whether there is a risk in the picture corresponding to the prompt word j (which can be called the third prompt word) through the picture risk discriminator 116, where the prompt word j is obtained by the prompt word N (i.e., the target prompt word) after the above multiple iterations. If there is a risk in the picture corresponding to the prompt word j, the server 104 modifies the prompt word j through the prompt word rewriter 112 to obtain the prompt word (j+1) (which can be called the fourth prompt word). Then, the server 104 generates the corresponding picture based on the prompt word (j+1) through the picture generator 114. After completing j iterations, the server 104 assigns (j+1) to j, thereby performing the next round of iterations until a picture corresponding to the prompt word M without risk is obtained.

[0054] In this way, the generated picture 108 can be subjected to dual risk identification based on the text risk discriminator 110 and the image risk discriminator 116, and the prompt words and pictures can be continuously optimized through a dual loop mechanism, thereby further improving the accuracy of risk identification and thus ensuring the safety of the picture 108. Moreover, through the prompt word rewriter 112, when there is a risk in the prompt word or the generated picture, the prompt word can be modified in combination with the loop mechanism, thereby improving the efficiency of removing risks and optimizing the prompt word, and reducing labor costs.

[0055] In some embodiments, the server 104 may also combine a reward mechanism and a penalty mechanism to reward and punish in the process of adjusting the prompt word through the prompt word rewriter 112. Specifically, if the server 104 detects that there is a risk in the generated picture through the picture risk discriminator 116, a certain reward coefficient may be set for the prompt word at this time. If the server 104 detects that there is no risk in the generated picture through the picture risk discriminator 116, a certain penalty coefficient may be set for the prompt word at this time. In the process of adjusting the prompt word, the server 104 may adjust the prompt word through the reward coefficient and the penalty coefficient by the prompt word rewriter 112. In this way, the number of iterations of adjusting the prompt word can be reduced, the modification speed of the prompt word can be accelerated, and the efficiency of generating pictures can be improved.

[0056] In some embodiments, the server 104 can adjust the prompt word in combination with the user's feedback on the generated image. Specifically, after the server 104 generates the image 108 through the image risk discriminator 116, the user can manually determine whether there is a degree of risk in the image 108, and score the image 108 according to the degree of risk. Among them, the higher the user's score, the higher the degree of risk in the image 108. When the user's score is higher than the score threshold, the server 104 can continue to modify the current prompt word through the prompt word rewriter 112, and generate a new image based on the modified prompt word through the image generator 114. Then, the server 104 receives the user's score for the generated new image again, thereby repeating the above iterative process. In this way, the risk in the image can be further judged in combination with user feedback, thereby improving the accuracy of risk identification.

[0057] In some embodiments, the server 104 can pre-train neural network models such as the text risk identifier 110, the prompt word rewriter 112, the image generator 114, and the image risk discriminator 116 in a large amount of security data, so that they tend to generate safe content in the task of generating images. At the same time, the server 104 can also fine-tune the above models in combination with security data in specific fields, thereby enhancing the security and reliability of the above models in specific fields.

[0058] Figure 5 FIG. 5 is a schematic diagram showing a third process 500 for generating a picture in some embodiments of the present specification. Figure 1 In the example environment 100 shown in FIG. 1 , process 500 may be performed by server 104. It should be understood that although the following content is described with server 104 as the execution subject, process 500 may also be performed by other devices. Process 500 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0059] In some embodiments, the server 104 may obtain the prompt word 106 input by the user, and then retrieve the context information 506 related to the prompt word 106 from the knowledge base 504 according to the prompt word 106. In the process of retrieving the context information 506, the server 104 may also combine the pre-set retrieval rules 502. For example, when the user requires to generate a certain picture, the corresponding knowledge base may be searched according to the current time, and the knowledge base corresponding to different times may be different. Among them, the knowledge base is a database, which can be set in the server 104 or in other remote servers. The context information 506 includes but is not limited to multimedia information such as text, audio, image, video, etc. related to the prompt word 106 input by the user.

[0060] After obtaining the context information 506, the server 104 can input the prompt word 106 and the context information 506 into the text risk discriminator 110 together, so as to determine whether there is a risk in the prompt word 106 through the text risk discriminator 110. It should be understood that in the process of the server 104 processing the prompt word 106 through the text risk discriminator 110, the prompt word rewriter 112, the image generator 114, and the image risk discriminator 116, the context information 506 can be combined to generate a more accurate processing result. In addition, in some embodiments, after obtaining the context information 506, the server 104 can update the retrieval rule 502 according to the user's feedback on the context information 506, so as to ensure the timeliness of the retrieval rule 502.

[0061] In some embodiments, if the server 104, based on the prompt word 106 and the context information 506, identifies a risk in the prompt word 106 by the text risk discriminator 110, the server 104 inputs the prompt word 106 and the context information 506 into the prompt word rewriter 112, and the prompt word rewriter 112 combines the context information 506 to modify the prompt word 106 to remove the risk factors therein, so as to obtain the modified prompt word. Then, the server 104 inputs the modified prompt word and the context information 506 into the image generator 114. If there is no risk in the prompt word 106, the server 104 does not need to modify the prompt word 106 and directly inputs the prompt word 106 and the context information 506 into the image generator 114.

[0062] Then, the server 104, through the image generator 114, generates a corresponding image based on the adjusted prompt word (when there is a risk in the prompt word 106) or the prompt word 106 (when there is no risk in the prompt word 106), and the context information 506. Further, the server 104 inputs the generated image and the context information 506 into the image risk discriminator 116, so as to determine whether there is a risk in the generated image through the image risk discriminator 116. If there is a risk in the generated image, the server 104 further adjusts the previously adjusted prompt word (when there is a risk in the prompt word 106) or the prompt word 106 (when there is no risk in the prompt word 106) through the prompt word rewriter 112, and inputs the adjusted prompt word into the image generator 114 again, so as to generate the image 108 again through the image generator 114. If there is no risk in the generated image, the server 104 can directly use the generated image as the final image 108.

[0063] In this embodiment, a retrieval enhancement mechanism is combined to combine the context information of the prompt word in the process of processing the text image task to perform risk identification and image generation tasks. In this way, multiple neural network models such as the text risk discriminator 110 can more accurately understand the user's intention and the semantics of the prompt word, which not only improves the accuracy of risk identification and the security of generated images, but also ensures that the generated image content is highly relevant to the text input by the user, reducing misunderstandings and thus meeting the user's expectations.

[0064] In some embodiments, the knowledge base 504 may be an ontology-based database, in which data nodes and the relationships between data nodes are stored in the knowledge base 504 in the form of triples or the like. After obtaining the prompt word 106 input by the user, the server 104 may determine the data node (which may be referred to as the first data) related to the prompt word 106 in the ontology database, and determine other data nodes (which may be referred to as the second data node) adjacent to the data node of the prompt word 106 based on the relationship between the data nodes. Further, the server 104 may extract the data node of the prompt word 106 and other adjacent data nodes from the ontology database, thereby forming context information 506 in the form of a knowledge graph. In this way, context information can be expressed in the form of a knowledge graph, thereby improving the accuracy and richness of the semantics of the context information. Combined with the knowledge graph, the ability of neural network models such as the risk text discriminator 110 to understand semantics can be improved, thereby improving the accuracy of risk identification and the accuracy of the generated image.

[0065] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, computing device embodiment, and computer program product embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0066] Figure 6 Schematically shows a block diagram of a computing device 600 suitable for implementing an embodiment of the present invention. The computing device 600 can be used to implement the server 104. Figure 6 As shown, the computing device 600 includes a processing unit (CPU) 601, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or computer program instructions loaded from a storage unit 608 to a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computing device 500 can also be stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 505 is also connected to the bus 604.

[0067] Multiple components in the computing device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, and a storage unit 608. The processing unit 601 performs the various methods and processes described above, such as executing the method 200. For example, in some embodiments, the various processes or operations described above can be implemented as a computer software program, which is stored in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the computing device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the CPU 601, the various methods and processes described above can be executed, such as executing one or more operations of the method 200. Alternatively, in other embodiments, the CPU 601 can be configured to execute the various methods and processes described above, such as executing one or more actions of the method 200, by any other appropriate means (for example, by means of firmware).

[0068] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0069] The computer program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present invention.

[0070] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of other programmable data processing devices, thereby producing a machine, so that when these instructions are executed by the processing unit of a computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other device to work in a specific manner.

[0071] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0072] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

[0073] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for generating a picture, comprising: Determining whether there is risk in the prompt word input by the user through the first neural network model; In response to the presence of risk in the prompt word, adjusting the prompt word through a second neural network model to remove the risk, so as to obtain a target prompt word; as well as A corresponding picture is generated based on the target prompt word through a third neural network model.

2. The method according to claim 1, wherein adjusting the prompt word comprises: In a current iteration, determining, by means of the first neural network model, whether there is a risk in a first prompt word, the first prompt word being related to the prompt word; as well as In response to the presence of risk in the first prompt word, the first prompt word is adjusted by the second neural network model to obtain a second prompt word for a next iteration, where the second prompt word is related to the target prompt word.

3. The method according to claim 1 or 2, further comprising: Determine whether there is risk in the picture corresponding to the target prompt word through the fourth neural network model; In response to the existence of risk in the picture corresponding to the target prompt word, adjusting the target prompt word through the second neural network model; as well as The third neural network model is used to generate a corresponding picture based on the adjusted target prompt word.

4. The method according to claim 3, further comprising: In the current iteration, determining whether there is a risk in the picture corresponding to the third prompt word through the fourth neural network model, wherein the third prompt word is related to the target prompt word; In response to the presence of risk in the picture corresponding to the third prompt word, adjusting the third prompt word by the second neural network model to obtain a fourth prompt word; A corresponding picture is generated based on the fourth prompt word through the third neural network model, and the picture corresponding to the fourth prompt word is used for the next iteration.

5. The method according to claim 3, wherein adjusting the target prompt word comprises: In response to the existence of risk in the picture corresponding to the target prompt word, determining a penalty coefficient for the target prompt word; In response to the absence of risk in the picture corresponding to the target prompt word, determining a reward coefficient for the target prompt word; as well as The target prompt word is adjusted based on the penalty coefficient or the reward coefficient through the second neural network model.

6. The method according to claim 1, wherein determining whether the prompt word has a risk comprises: Acquire the prompt word input by the user; Acquire context information related to the prompt word from a database, wherein the database is set in a local or remote server; as well as By using the first neural network model, based on the prompt word and the context information, it is determined whether there is a risk in the prompt word.

7. The method according to claim 6, wherein the database comprises an ontology database, the context information comprises a knowledge graph, and obtaining the knowledge graph comprises: Acquire a first data node related to the prompt word input by the user in the ontology database; Based on the relationship between the data in the benefit database, determining a second data node associated with the first data; as well as The knowledge graph is generated based on the first data node and the second data node.

8. The method according to claim 1, further comprising: Obtaining a score given by the user to the picture corresponding to the target prompt word, wherein the score indicates a risk level of the picture corresponding to the target prompt word; In response to the score being higher than a score threshold, adjusting the target prompt word by the second neural network model; as well as The third neural network model is used to generate a corresponding picture based on the adjusted target prompt word.

9. A device for generating a picture, comprising: A risk determination module is configured to determine whether there is a risk in the prompt word input by the user through a first neural network model; a prompt word adjustment module, configured to, in response to the risk existing in the prompt word, adjust the prompt word through a second neural network model to remove the risk, so as to obtain a target prompt word; as well as The picture generation module is configured to generate a corresponding picture based on the target prompt word through a third neural network model.

10. A computing device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the computing device to perform the method according to any one of claims 1 to 8.

11. A computer program product, comprising a computer program, the computer program being executed by a processor to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Sample generation method and device

    CN116663547A

  • Task processing method and device based on large language model and electronic equipment

    CN117131540A

  • Painting processing method and device, storage medium and electronic equipment

    CN118918199A

  • Security image generation method and device, electronic equipment and storage medium

    CN118967879A