Method and apparatus for generating image, computing device and program product
By identifying and adjusting the modal prompt words in the generative model, the problem of inappropriate content generation of generative models is solved, and risk prevention and control is achieved before content generation is generated, ensuring the legality, compliance and improvement of user experience of generated content.
Patent Information
- Application Number
- CN202510127714.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-08
AI Technical Summary
The personalized content generated by the generative model is unpredictable, prone to improper or harmful content, and it is difficult to ensure legality and compliance.
By obtaining the first and second modal prompt words, using the corresponding adjustment model to identify and adjust the hazard information, generate prompt words that meet the requirements, and finally generate legal and compliant images.
Identify and remove potentially dangerous information before content generation, significantly reduce the possibility of inappropriate content generation, improve user experience, and enhance the controllability and legality of generated results.
Smart Images

Figure CN120451295A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and more particularly, to methods, apparatuses, computing devices, and program products for generating images. Background Art
[0002] With the rapid development of artificial intelligence (AI) and advancements in deep learning, AI is increasingly being applied to creating personalized content, such as generating personalized images or text. For example, generative models can be used to complete various natural language processing tasks, such as text generation, video generation, and image generation.
[0003] With the explosive growth in demand for content, the applications of artificial intelligence are becoming increasingly widespread. Due to the inherent characteristics of generative models, the personalized content they generate is unpredictable. This can easily lead to risks such as inappropriate or harmful content and copyright infringement. Therefore, ensuring the legality and compliance of content generated by generative models has become an inevitable requirement for industry development. Summary of the Invention
[0004] The embodiments of this specification provide a method, an apparatus, a computing device, and a program product for generating an image.
[0005] In a first aspect of the present specification, a method for generating an image is provided. The method includes obtaining a first modal prompt word and a second modal prompt word. The method also includes adjusting the danger information in the first modal prompt word using a first modal adjustment model to determine an adjusted first modal prompt word. The method also includes adjusting the danger information in the second modal prompt word using a second modal adjustment model to determine an adjusted second modal prompt word. In addition, the method also includes generating a corresponding image based on the adjusted first modal prompt word and the adjusted second modal prompt word.
[0006] In the second aspect of the present specification, a device for generating an image is provided. The device includes a prompt word acquisition unit, which is configured to acquire a first modal prompt word and a second modal prompt word. The device also includes a first modal prompt word adjustment unit, which is configured to adjust the danger information in the first modal prompt word through a first modal adjustment model and determine the adjusted first modal prompt word. The device also includes a second modal prompt word adjustment unit, which is configured to adjust the danger information in the second modal prompt word through a second modal adjustment model and determine the adjusted second modal prompt word. In addition, the device also includes an image generation unit, which is configured to generate a corresponding image based on the adjusted first modal prompt word and the adjusted second modal prompt word.
[0007] In a third aspect of this specification, a computing device is provided. The computing device includes one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, causes the one or more processors to implement the method provided in accordance with the first aspect of this specification.
[0008] In a fourth aspect of the present specification, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method provided according to the first aspect of the present specification.
[0009] In a fifth aspect of the present specification, a computer program product is provided, comprising a computer program, wherein the computer program is executed by a processor to implement the method provided in the first aspect of the present specification.
[0010] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of this specification, nor are they intended to limit the scope of this specification. Other features of this specification will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages and aspects of the embodiments of this specification will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0012] Figure 1 A schematic diagram illustrating an example environment in which various embodiments of the present specification may be implemented;
[0013] Figure 2 A flowchart illustrating a method for generating an image according to some embodiments of the present specification is shown;
[0014] Figure 3 A schematic diagram showing multi-layer risk assessment of text prompt words according to some embodiments of this specification is shown;
[0015] Figure 4 Schematic diagram of determining a hazard warning word according to some embodiments of this specification;
[0016] Figure 5 A schematic diagram showing an interactive interface according to some embodiments of this specification; and
[0017] Figure 6 A block diagram of a computing device 600 suitable for implementing embodiments of the present invention is schematically shown. DETAILED DESCRIPTION
[0018] The following describes embodiments of the present specification in more detail with reference to the accompanying drawings. Although certain embodiments of the present specification are shown in the accompanying drawings, it should be understood that the present specification can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present specification. It should be understood that the drawings and embodiments of the present specification are for illustrative purposes only and are not intended to limit the scope of protection of the present specification.
[0019] In the description of the embodiments of this specification, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. can refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0020] As mentioned above, generative models have a high degree of freedom, and the generated content is often unpredictable. Generative models typically generate content based on the introductory text. In other words, the legality and compliance of the content generated by generative models largely depends on whether the introductory text contains dangerous information. Therefore, identifying and filtering illegal content in the introductory text is key to ensuring the legality and ethics of the generated content.
[0021] To this end, an embodiment of this specification proposes a method for generating an image. The method includes obtaining a first modal prompt word and a second modal prompt word. The method also includes adjusting the danger information in the first modal prompt word using a first modal adjustment model to determine an adjusted first modal prompt word. The method also includes adjusting the danger information in the second modal prompt word using a second modal adjustment model to determine an adjusted second modal prompt word. In addition, the method also includes generating a corresponding image based on the adjusted first modal prompt word and the adjusted second modal prompt word.
[0022] This approach removes dangerous information from prompts across different modalities, ensuring that the prompts fed into the generative model are safe, compliant, and meet requirements. This allows risk prevention and control before content is generated, significantly reducing the likelihood of inappropriate content output from the generative model and improving the user experience.
[0023] Figure 1 1 is a schematic diagram of an example environment 100 in which various embodiments of the present specification may be implemented. Figure 1As shown, in the environment 100, a terminal device 102 and a computing device 104 that communicates with the terminal device 102 through a network are included. A first modality adjustment model 106, a second modality adjustment model 108, and a generative model 110 are deployed in the computing device 104. The terminal device 102 can be a smart phone, a desktop computer, a laptop computer, a notebook computer, a tablet computer, a personal digital assistant (PDA), a smart watch, or any combination thereof. The computing device 104 can be an electronic device capable of data transmission and data processing, or a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. The terminal device 102 and the computing device can be directly or indirectly connected via wired or wireless communication.
[0024] In the example environment 100, the user can input information through the terminal device 102, for example, by voice input, text input, touch input to enter their own request ( Figure 1 (not shown in the figure). For example, the input information can be text content or image content. In one example, the input information can be "a cat playing on the grass" and an image of a cat on the grass. It can be understood that the text content and the image content can correspond to each other, for example, the text content can be text content used to describe the image. Based on the user's input information, a prompt word (Prompt) can be generated. Among them, the prompt word refers to an instruction or prompt input into a language model or a generative model, usually in the form of text. The guide text or prompt text can be a text description or a parameter description in a certain format. For example, a first modal prompt word 112 (text prompt word) and a second modal prompt word 114 (image prompt word) can be generated based on the user's input information.
[0025] In actual applications, the information input by the user may contain inappropriate content. For example, it may contain information that does not comply with laws and regulations or contain sensitive words. This content can directly or indirectly lead the generative model to generate content that does not meet the requirements. Therefore, it is necessary to remove or adjust the dangerous information contained in the prompt words so that the content generated by the generative model 110, such as images or text, meets the requirements. Dangerous information refers to information that does not comply with laws and regulations, is not true, or does not meet preset requirements.
[0026] like Figure 1As shown, the first modal adjustment model 106 can be used to adjust the danger information in the first modal prompt word 112 so that the adjusted danger information meets the requirements. Among them, the first modal adjustment model 106 can have the functions of detection and adjustment, wherein detection refers to determining whether the first modal prompt word 112 contains danger information. Adjustment refers to removing the danger information in the first modal prompt word 112 by rewriting or regenerating. On this basis, the first modal adjustment model 106 can be a pre-trained neural network model. Of course, the first modal adjustment model 106 can only have the adjustment function, and the danger information in the first modal prompt word 112 can be identified by other pre-trained models. Similarly, the second modal adjustment model 106 can also be a model with detection and adjustment functions.
[0027] In the example environment 100, the adjusted first modal prompt word and the adjusted second modal prompt word are input into the generative model 110, and the generative model 110 generates corresponding content based on the prompt word, such as generating a corresponding image 116. A generative model is a deep learning model with a large number of parameters that can be trained with available data. The generative model can be fine-tuned to enable it to complete various generation tasks. Depending on the generation task, the generative model can be divided into a language model, a visual model, a speech model, and a multimodal model. For example, the generative model can be a text-image model that can generate a corresponding image based on the adjusted prompt word. In this way, multimodal input information enables the generative model to absorb a wider range of knowledge and representation methods. In this way, the content generated by the generative model guided by the prompt words of multiple modalities is richer and more diverse, with a higher level of creativity.
[0028] Combined with the above Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present specification can be implemented is depicted. It should be understood that the environment 100 is merely illustrative and does not limit the scope of the present specification. The environment 100 may include Figure 1 More components are shown, and various components in environment 100 may also be implemented in different ways.
[0029] Figure 2 FIG. 2 shows a flow chart of a method 200 for generating an image according to some embodiments of the present specification. The method 200 may be performed by Figure 1 The computing device 106 shown in FIG. Figure 2As shown, in box 202, method 200 may include obtaining a first modal prompt word and a second modal prompt word. Among them, modality refers to the description method (perspective or field) adopted for the same object. For example, modality can be divided into visual modality (such as image, video), text modality and auditory modality (such as sound). The first modality is different from the second modality. For example, the first modality can be a text modality and the second modality can be an image modality. On this basis, the first modal prompt word can be a text prompt word and the second modal prompt word can be an image prompt word. In one example, the text prompt word can be the text content of "a kitten playing on the grass" input by the user through the interactive interface. The image prompt word can be an image of a cat lying on the grass. It should be noted that the first modality and the second modality are not limited to text modality and image modality, but can also be other types of modalities. This specification does not limit the modality of the prompt word.
[0030] At block 204, method 200 may include adjusting the dangerous information in the first modal prompt word using the first modal adjustment model to determine an adjusted first modal prompt word. Dangerous information can refer to information that violates regulations or is otherwise potentially dangerous under specific circumstances or legal conditions, or it can also be sensitive information. For example, it can include information that violates platform regulations or legal provisions, or disseminates false content.
[0031] On this basis, it is possible to identify whether the first modal prompt word contains dangerous information, and when it is identified that the first modal prompt word contains dangerous information, the first modal prompt word is adjusted using the first modal adjustment model (for example, the dangerous information in the first modal prompt word is removed) to determine the first modal prompt word that meets the requirements. Among them, the first modal adjustment model can be a speech model, and the language model can be a deep learning model trained using text data, which can process natural language such as generating natural language text or understanding natural language text. The language model can adjust the text prompt word to determine the text prompt word that meets the requirements. Among them, the text prompt word that meets the requirements means that the prompt word does not contain dangerous information, and the content contained is legal and compliant.
[0032] In some embodiments, a trained machine learning model (such as a neural network model, a deep learning network model, etc.) can be used to identify whether the text prompt word contains dangerous information. Among them, adjustment can refer to modifying or redefining the text prompt word. For example, false information in the text prompt word can be modified to accurately verified information, or discriminatory information in the text prompt word can be modified to neutral information. Of course, in the case where the text prompt word contains many dangerous information elements and modification is more difficult, a new text prompt word can be re-determined.
[0033] At block 204, method 200 may include adjusting the danger information in the second modal prompt word using the second modality adjustment model to determine an adjusted second modal prompt word. Similarly, prompt words in other modalities may also contain danger information. For example, in the image modality, the image input by the user may contain danger information (e.g., danger elements, danger words).
[0034] Specifically, prompts in different modalities often carry unique information and expression conventions, such as language style in text modalities, visual elements in image modalities, and speech characteristics in audio modalities. Therefore, different adjustment or detection models are required for prompts in different modalities to ensure that the adjusted prompts conform to the expression standards of the specific modality and do not contain any dangerous or inappropriate content.
[0035] It is understood that the method for adjusting the danger information of the second modal prompt word is similar to the above. The difference is that the adjustment model or detection model used to adjust or detect the second modal prompt word is different. The adjustment models of different modalities are used to adjust the prompt words of different modalities. In this way, it can ensure that the prompt word can better adapt to the expression method of the specific modality. In addition, using a specific model to adjust the prompt word of a specific modality can improve the efficiency and accuracy of the prompt word adjustment.
[0036] At block 206, method 200 may include generating a corresponding image based on the adjusted first modal prompt word and the adjusted second modal prompt word. After obtaining the first modal prompt word and the second modal prompt word that meet the requirements, they may be input into a generative model, which then generates corresponding content such as an image, text, video, or audio.
[0037] In this way, dangerous information can be removed from prompts across different modalities, ensuring that the prompts fed into the generative model are safe, compliant, and meet requirements. This allows risk prevention and control before content is generated, significantly reducing the likelihood of the generative model outputting inappropriate content and improving the user experience. Furthermore, by removing dangerous information from prompts, the generative model can be guided to learn from a wider and more balanced set of data, reducing bias in the generated content and improving the controllability of the generated results.
[0038] In real-world applications, some dangerous information contained in text or image prompts is potentially subtle and difficult to detect. Furthermore, over time or as the environment changes, some prompts that were once considered safe may no longer be safe. To ensure that the prompts fed into the model are safe and compliant, multiple layers of risk assessment can be performed on the prompts to minimize the risk of dangerous information contained in them.
[0039] Figure 3 Schematic diagram showing multi-layer risk assessment of text prompt words in some embodiments of this specification. Figure 3 As shown, the text prompt word 302 input by the user through the interactive interface of the terminal device can be input into the text risk discriminator 304, and the text risk discriminator 304 identifies the dangerous information in the text prompt word 302. The text risk discriminator 304 can be a pre-trained neural network model or a deep learning model, for example, including but not limited to a decision tree, a random forest, a support vector machine (SVM), a convolutional neural network, a recurrent neural network, and a long short-term memory network. For example, the initial text risk discriminator can be trained using text samples containing dangerous labels, and the network parameters of the initial text risk discriminator can be iteratively adjusted to determine the trained text risk discriminator 304.
[0040] like Figure 3 As shown, after a text prompt word 302 is input into a text risk discriminator 304, the text risk discriminator 304 can output a risk label corresponding to the text prompt word 302. The risk label may include a first label, a second label, a third label, a fourth label, and so on. It is understood that a text prompt word 302 may correspond to multiple risk labels, for example, including both a first label and a second label. Of course, the text prompt word 302 may not include risk information. In this case, the text prompt word can be directly input into the generative model, and the generative model 302 can generate a corresponding image based on the text prompt word.
[0041] In actual applications, different users have different generation requirements. For example, user A wants the generative model to generate more rigorous and regular content, such as a more regular image rather than a jumpy image. The prompt word plays a guiding role in the generation of content by the generative model. On this basis, the user can pre-set the prompt word requirements according to their own needs. For example, preset requirement 1, preset requirement 2, and preset requirement 3 can be set. In some examples, preset requirement 1 can be "unified style", that is, emphasizing the consistency of the required style in the prompt word, such as "maintaining traditional aesthetics, following design specifications", etc., to ensure that the generated image or content remains unified in style. Preset requirement 2 can be "specified color range, shape type or pattern element". In this way, the danger label of the text prompt word can also be not meeting the preset requirements.
[0042] In other embodiments, different users may have different emotions or feelings towards the same content. For example, user A and user B have different cultural backgrounds and values. User A does not have any special emotions towards a certain cultural information contained in the text prompt word s, while user B feels that the text prompt word S contains unfriendly information. On this basis, the user's emotions can be used as a basis for judging dangerous information. In other words, dangerous information can include emotional risk information. Among them, emotional risk information can be information that triggers negative emotions in users. For example, the sentiment analysis technology of natural language processing can be used to determine the emotional tendency in the text prompt word, and combined with information such as the user's cultural background and values, it can be determined whether the text prompt word contains emotional risk information.
[0043] like Figure 3 As shown, after determining the danger label of the text prompt word 302, the text risk discriminator 304 can input the danger label into the prompt word rewriting model 306, and the prompt word rewriting model 306 adjusts the text prompt word 306 containing the danger label to obtain the adjusted text prompt word 308. Among them, the prompt word rewriting model 306 can be another type of generative model (not the same model as the generative model used to generate the corresponding image). For example, a danger prompt word can be generated based on the danger label. The generative model can generate compliant and safe text prompt words that meet the requirements based on the danger prompt words and the original text prompt words. The generated text prompt words can be used as the adjusted text prompt words 308 as the guiding text for generating the image. The danger prompt words can be used to describe the category and danger level of the danger information contained in the text prompt words.
[0044] In actual applications, the rewritten text prompt may still contain potential risks. Based on this, the rewritten text prompt can be re-entered into the text risk discriminator 304, which then determines whether the rewritten text prompt contains any dangerous information. If not, the rewritten text prompt is used as the final text prompt and input into the generative model. If the rewritten text prompt still contains dangerous information, the rewritten text prompt can be rewritten or a new text prompt can be generated.
[0045] This approach allows for a more detailed and comprehensive assessment of prompts using a multi-layered risk analysis approach, preventing potential risks from being missed. Furthermore, the detection of hazardous information in text prompts is similar to a closed-loop system, minimizing the risk of hazardous information in prompts through iterative risk identification and adjustment.
[0046] In some embodiments, the danger information contained in the text prompt word may be of multiple types. On this basis, two types of prompt words can be generated according to the two types of danger information, thereby helping the generative model to regenerate new text prompt words that meet the requirements. Figure 4 FIG2 shows a schematic diagram of determining a danger warning word according to some embodiments of this specification. Figure 4 As shown, each hazard type corresponds to a prompt word. For example, the first hazard type corresponds to prompt word 1, the second hazard type corresponds to prompt word 2, the third hazard type corresponds to prompt word 3, and the fourth hazard type corresponds to prompt word 4. The prompt word indicates the processing requirements for the corresponding hazard content, that is, the requirements for the generated image.
[0047] Based on the risk labels output by the text risk discriminator, it is possible to determine what dangerous information is contained in the text prompt word. If the text prompt word does not contain sensitive content, prompt word 1 can be empty or a special identifier. After determining multiple dangerous prompt words, the generative model can generate new text prompt words based on the dangerous prompt words and text prompt words. In this way, the regenerated text prompt words meet all requirements, thereby reducing the risk of generating illegal images.
[0048] In practical applications, historical danger signs often reflect past problems and potential risks encountered during text processing or communication. These problems may involve inappropriate word choice, sensitive expressions, or information that could potentially lead to misunderstandings. Based on this, historical danger signs can, to a certain extent, serve as a basis for adjusting text signs. For example, by analyzing historical danger signs, we can better understand the sensitive points and potential risks in the current context. Historical danger signs can be the danger signs corresponding to text signs entered by users at historical times. By analyzing historical danger signs, we can identify words that are prone to problems or expressions that are prone to misunderstandings. Thus, a problematic vocabulary list can be formed based on the frequency of occurrence of problematic words. This vocabulary list can then be referenced when determining whether text signs contain dangerous information or when adjusting text signs. This allows for faster determination of whether text signs contain dangerous information and quicker identification of text content that requires adjustment, thereby improving the efficiency and accuracy of adjusting text signs.
[0049] In some embodiments, context information refers to information involved in the process of user interaction with the generative model, such as multiple rounds of user input information and multiple rounds of output information of the generative model. It is understood that context information can be stored with a computing device (e.g. Figure 1 ) or stored in a database of a terminal device (e.g., a computing device 104 shown in FIG. Figure 1In the local storage space of the terminal device 102 shown in FIG. In combination with the context information, the user's preferences and habits can be analyzed. Similarly, the user's preferences and habits can also be used as a basis for identifying whether there is dangerous information in the text prompt words to a certain extent.
[0050] For example, if users frequently enter words or expressions involving sensitive topics, this may indicate that they have a higher tolerance or interest in such topics. Of course, this may also mean that they are more susceptible to certain types of bad information or misleading. Therefore, when identifying dangerous information in text prompt words, the user's preferences and habits can be used as a basis for judgment, so as to more accurately judge which content is more likely to mislead users. Specifically, the dangerous information screening range can be personalized according to the user's specific needs and preferences. Combined with this screening range, it is possible to more quickly adjust the dangerous information in the text prompt words or more quickly identify whether the text prompt words contain dangerous information.
[0051] In some embodiments, the user may be allowed to provide real-time feedback and adjustments during the adjustment process of the prompt words. For example, the text risk discriminator may be used to provide feedback to the user regarding uncertain information in the text prompt words. The user may determine whether the information is dangerous information based on their own professional knowledge and experience. In this way, a more detailed and comprehensive assessment of the prompt words may be ensured to avoid missing any potential risks. In addition, during the adjustment process of the text prompt words, the user's feedback information may also be used as a basis for adjusting the text prompt words. For example, the user may be emotionally dissatisfied or opposed to the text prompt words, for example, the user may think that certain words are unfriendly or offensive. On this basis, the dangerous information in the text prompt words may be adjusted based on the user's feedback. Of course, the text risk discriminator may also be adjusted, for example, by adjusting the network parameters of the text risk discriminator, so that the adjusted text risk discriminator is more likely to identify the dangerous information in the text prompt words, thereby improving the accuracy and efficiency of recognition.
[0052] Similarly, the identification of dangerous information in image prompt words and the adjustment of image prompt words are similar to the identification of dangerous information in text prompt words and the adjustment of text prompt words. For example, an image risk discriminator can be used to identify whether an image prompt word contains dangerous information. The image prompt word is in the form of an image. Identifying whether an image prompt word contains dangerous information can be identifying whether the various elements, layout, color, etc. in the image are legal and compliant and meet the set requirements. For example, if an image contains the identity information of a user, there may be a risk of violating the user's privacy, and it can be determined that there is dangerous information in the image.
[0053] Unlike text-based prompt word adjustments, images contain more information, such as color, shape, texture, and background imagery, making image adjustments more complex. Therefore, if an image is identified as containing dangerous information, the text generation model can generate corresponding text to replace the original image information. This approach not only ensures the legitimacy and safety of prompt words, but also improves the efficiency of image prompt word adjustments.
[0054] For example, after determining the dangerous information contained in the image, the dangerous information and the features of the image can be input as prompt words into the copy generation model, and the corresponding copy content can be generated by the copy generation model. Among them, the copy generation model is a type of generative model, for example, it can be a language model, which is used to generate text content. Among them, the image features can be extracted by the feature extraction module. In addition, the key elements describing the image and the descriptive words of the image can also be input into the copy generation model to assist the copy generation model in generating more accurate copy content, which can both describe the image and does not contain dangerous information. It can be understood that the copy content can be used as an adjusted image prompt word and input into the generative model as a guiding text for the generative model to generate images. Of course, in other embodiments, the dangerous information contained in the image can also be filtered by the generative model to obtain a legal and compliant image as the basis for generating the image.
[0055] In this way, using text content to replace images containing dangers can not only ensure the legality and safety of image prompts, but also improve the efficiency and flexibility of image processing.
[0056] In some embodiments, in order to improve the rewriting efficiency of text prompt words and ensure that the output of the generative language model is both in line with the specifications and can accurately convey information, multiple rewriting strategies or generation strategies can be pre-set. Among them, the rewriting strategy can be set in advance by the user based on the characteristics of the text. Among them, the rewriting strategy may include but is not limited to replacing ambiguous words with synonyms or extended words, deleting words that do not meet the requirements, providing text content in another way of expression based on the semantics of the text prompt words, adjusting the word order, etc. For example, the sentence structure can be reorganized and different grammar or expressions can be used to convey the same information.
[0057] It's understandable that the images generated by the generative model may not meet user needs or preferences. Based on this, the images generated by the generative model can be modified based on user feedback. For example, certain elements in the image (such as icons, text, etc.) can be modified, or the image layout can be altered. For example, users can provide feedback on what they would like to adjust through voice or text input. For example, a user's feedback could be, "Please add some beautiful colors to the image."
[0058] In some embodiments, if the user is highly dissatisfied with the generated image, the generative model can be adjusted, and new image content can be regenerated using the adjusted generative model. If the user is less satisfied with the generated image, the generative model or other neural network model can be used to adjust the generated image, such as adjusting the image filter or style. This approach provides users with relatively free editing permissions, expanding the scope of user adjustments. Furthermore, users only need to describe the image at the semantic level, without having to directly adjust the image, which can improve the user experience.
[0059] It is understandable that since the generative model has a high degree of freedom when generating personalized content, even if the prompt words input into the generative model meet the requirements and are legal and compliant, the generative model may still output content that does not meet the requirements. On this basis, a new image can be regenerated by adjusting the generative model (for example, adjusting the weight value or network parameters of the generative model). Or a new image can be regenerated by adjusting the image prompt words or text prompt words. Specifically, after determining the generated image, risk identification can be performed on the image to determine whether there is a risk in the image. Among them, the presence of risk means that there is inappropriate information in the image. After determining the risk in the image, the text prompt words or image prompt words can be adjusted according to the risk information, so that the generative model generates an image without risk based on the adjusted text prompt words or adjusted image prompt words. In this way, the legality and security of the generated content are ensured, reducing legal and ethical risks for enterprises and individual users.
[0060] As mentioned above, users can set the requirements of the guide text or set the style or characteristics of the content they want to generate according to their own preferences. For example, users can use a terminal device (such as Figure 1 The user can input information through an interactive interface provided by the terminal device 102 shown in FIG. 1 or other device. The interactive interface may include a command interface, a menu interface, a graphical user interface, etc. Thus, by selecting a preset style option or entering a custom style description on the interactive interface, the user can ensure that the generated content meets their expectations in terms of style.
[0061] Figure 5 Schematic diagram of the interactive interface according to some embodiments of this specification is shown. Figure 5As shown, interactive interface 502 may include display component 504 and confirmation control 506. Display component 504 may include multiple sub-display components for displaying selectable styles or selectable sensitivity levels. Users can define the style of generated content by selecting a single or combined selection based on their needs. For example, style 1 may be "humorous" and style 2 may be "creative."
[0062] In some embodiments, users can choose different sensitivity levels according to their needs. Among them, the sensitivity level refers to the tolerance or screening level for dangerous information contained in text prompt words or image prompt words. Adult content may allow a wider range of expressions (but the expression should also be within the scope of social norms and laws). On this basis, if the generated content, such as images, is aimed at children, the user can choose a high sensitivity level; if the generated content is aimed at adults, the user can choose a low sensitivity level. After the user selects the corresponding style and sensitivity level, he can choose whether to trigger the confirmation control 506 according to his own needs. The confirmation control 506 can be triggered to run and complete the response function through a trigger event. Among them, the trigger event may include a mouse click, a user's finger click, etc. The confirmation control 506 can be in the form of a button control, and the button control may include a text box button, a pure icon button, etc.
[0063] It is understandable that the interactive interface 502 can be provided with clear and explicit definitions (instructions) and examples for different sensitivity levels to help users understand how their choices will affect the generated content. For example, when a user clicks or double-clicks the "Advanced Sensitivity" display component, the definition and examples corresponding to the sensitivity level can be displayed in a pop-up window or other form. Similarly, the interactive interface can also provide simple instructions for different styles, which will not be elaborated in this manual. Of course, the interactive interface can also provide other settings for users, such as keyword filtering, specific emotional tendencies, image layout, image color scheme, etc.
[0064] In this way, users can customize generated content to their preferences and needs by inputting guidance text and requirements into the interactive interface of their terminal or other smart devices. This personalized setting not only improves user experience satisfaction but also provides broader space and possibilities for the application of generative models.
[0065] Figure 6 The block diagram of a computing device 600 suitable for implementing an embodiment of the present invention is schematically shown. The computing device 600 may be a computer system for implementing Figure 2 The apparatus of the method 200 is shown. Figure 6As shown, computing device 600 includes a processing unit (CPU) 601, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 602 or loaded from storage unit 608 into random access memory (RAM) 603. Various programs and data required for the operation of computing device 600 can also be stored in RAM 603. CPU 601, ROM 602, and RAM 603 are connected to each other via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0066] Multiple components in computing device 600 are connected to I / O interface 605, including input unit 606, output unit 607, and storage unit 608. Processing unit 601 executes the various methods and processes described above, such as method 200. For example, in some embodiments, the various processes or operations described above may be implemented as a computer software program stored on a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on computing device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by CPU 601, the various methods and processes described above may be executed, such as one or more operations of method 200. Alternatively, in other embodiments, CPU 601 may be configured to execute the various methods and processes described above, such as one or more actions of method 200, by any other suitable means (e.g., via firmware).
[0067] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0068] The computer program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The computer readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can be personalized by utilizing the state information of the computer readable program instructions, and the electronic circuit can execute the computer readable program instructions, thereby realizing various aspects of the present invention.
[0069] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of other programmable data processing devices, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner.
[0070] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0071] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0072] The above are merely optional embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for generating an image, comprising: Obtaining a first modal prompt word and a second modal prompt word; adjusting the danger information in the first modal prompt word using a first modal adjustment model to determine an adjusted first modal prompt word; adjusting the danger information in the second modal prompt word using a second modal adjustment model to determine an adjusted second modal prompt word; as well as Based on the adjusted first modal prompt word and the adjusted second modal prompt word, a corresponding image is generated.
2. The method according to claim 1, wherein adjusting the danger information in the first modal prompt word using the first modal adjustment model to determine the adjusted first modal prompt word comprises: Determining whether the first modal prompt word contains dangerous information, where the dangerous information includes at least one or more of sensitive information, emotional risk information, and information that does not meet preset requirements; as well as In response to the first modal prompt word containing the danger information, the first modal prompt word is adjusted using the first modal adjustment model to determine an adjusted first modal prompt word.
3. The method according to claim 2, further comprising: determining whether the adjusted first modal prompt word includes the danger information; as well as In response to the adjusted first modal prompt word containing the danger information, the first modal prompt word is iteratively adjusted until the adjusted first modal prompt word does not contain the danger information.
4. The method according to claim 2, wherein adjusting the first modal prompt word comprises regenerating the first modal prompt word, and comprises: generating a danger prompt word based on the danger information contained in the first modal prompt word; as well as Based on the danger prompt word, a first modal prompt word is regenerated.
5. The method according to claim 2, wherein adjusting the first modal prompt word comprises: Based on the danger information contained in the first modal prompt word, obtaining a target rewriting strategy corresponding to the danger information from a plurality of preset rewriting strategies; as well as Based on the target rewriting strategy, the first modal prompt word is adjusted.
6. The method according to claim 4, wherein regenerating the first modal prompt word further comprises: Obtain historical hazard warning words and contextual information; as well as A first modal prompt word is regenerated based on the historical danger prompt word, the context information and the danger prompt word.
7. The method according to claim 2, wherein adjusting the first modal prompt word comprises: Obtaining user input information for the first modal prompt word; In response to the input information, determining the user's adjustment intention with respect to the first modal prompt word; as well as Based on the adjustment intention, the first modal prompt word is adjusted.
8. The method according to claim 7, wherein adjusting the first modal prompt word comprises: determining, based on the input information, an adjustment instruction of the user for the first modality adjustment model; In response to the adjustment instruction, adjusting the first modal adjustment model; as well as Based on the adjusted first modal adjustment model, the first modal prompt word is adjusted.
9. The method according to claim 2, wherein generating a corresponding image based on the adjusted first modal prompt word and the adjusted second modal prompt word comprises: Generate prompt words based on user requests; as well as Based on the adjusted first modal prompt word, the adjusted second modal prompt word, and the prompt word, a corresponding image is generated through a generative model.
10. The method according to claim 1 or 9, further comprising: determining whether the generated image poses a risk; as well as In response to the generated image being risky, the corresponding image is regenerated by adjusting the first modal prompt word and / or the second modal prompt word.
11. The method according to claim 10, wherein regenerating the corresponding image by adjusting the first modal prompt word and / or the second modal prompt word comprises: Obtaining user feedback on the generated image; Regenerate a corresponding image based on the feedback information, the adjusted first modal prompt word, and the adjusted second modal prompt word.
12. An apparatus for generating an image, comprising: a prompt word acquisition unit, configured to acquire a first modal prompt word and a second modal prompt word; a first modal prompt word adjustment unit configured to adjust the danger information in the first modal prompt word using a first modal adjustment model and determine an adjusted first modal prompt word; a second modal prompt word adjustment unit, configured to adjust the danger information in the second modal prompt word using a second modal adjustment model and determine an adjusted second modal prompt word; as well as The image generating unit is configured to generate a corresponding image based on the adjusted first modal prompt word and the adjusted second modal prompt word.
13. A computing device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the computing device to perform the method according to any one of claims 1 to 11.
14. A computer program product comprising a computer program, the computer program being executed by a processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Image cue word generation method and device, electronic equipment and storage medium
CN116580283A
Method and device for generating image, equipment and medium
CN117671067A
Content generation method and device, equipment and storage medium
CN117874239A
End-to-end multi-mode text image desensitization method and device and storage medium
CN117912045A
Data desensitization optimization method based on multi-modal algorithm network
CN118070324A