Image-text interaction method and device and storage medium

By displaying iterative denoising images and obtaining user input information, and displaying interactive information based on similarity, the problem of single interaction process of existing graphic and text interaction applications is solved, and more interesting and interactive graphic and text interaction is achieved.

CN120523375APending Publication Date: 2025-08-22BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410190218.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The interactive process of existing graphic and text interaction applications is single and has limited interest. The user input content must be completely consistent with the target answer to give the correct interaction result and affect the user experience.

Method used

By displaying multiple images obtained during the iterative denoising process, users are obtained input information, and interactive information is displayed based on similarity, and image collection is generated using the diffusion model to enhance the randomness and fun of the interaction process.

Benefits of technology

It enhances the interactivity and fun of the graphics and text interaction process, reduces the storage space requirement, and improves the user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523375A_ABST
    Figure CN120523375A_ABST
Patent Text Reader

Abstract

The invention relates to an image-text interaction method and device and a storage medium. The image-text interaction method comprises the steps that a first image in an image set is displayed, the image set comprises a plurality of noise images obtained by conducting iteration denoising processing on random noise images corresponding to different iteration times and a target image obtained after denoising processing is completed, and the first image is one of the noise images. Obtaining first input information which is input by a user based on the first image and represents image content, and determining the similarity between the image content represented by the first input information and the image content of the target image. Based on the similarity, interaction information is displayed, and the interaction information is used for representing whether the image content represented by the first input information is matched with the image content of the target image or not. Through the image-text interaction method and device, interactivity and interestingness of an image-text interaction process can be enhanced, and interaction experience of a user is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method, device, and storage medium for graphic-text interaction. Background Art

[0002] With the rapid development of information technology, human-computer interaction technology has made significant progress in various fields. Among them, game applications that utilize human-computer interaction are highly popular among users for their fun and interactivity. Existing image-text interaction applications primarily focus on directly generating an image based on user input, or guessing images based on contour images or blurred images. These applications suffer from a single interactive process and limited fun. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides a graphic-text interaction method, device and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, a method for image-text interaction is provided, comprising: displaying a first image in an image set, the image set comprising a plurality of noise images obtained by iterative denoising of a random noise image corresponding to different numbers of iterations and a target image after the denoising process, the first image being one of the plurality of noise images; obtaining first input information representing image content input by a user based on the first image, and determining the similarity between the image content represented by the first input information and the image content of the target image; and displaying interaction information based on the similarity, the interaction information being used to represent whether the image content represented by the first input information matches the image content of the target image.

[0005] In one embodiment, the image set is determined in the following manner: the first text content and the random noise image are input into a diffusion model, and the random noise image is subjected to m iterative denoising processes, where m is a positive integer greater than 1; starting from the i-th iterative process, n denoised images are selected in sequence according to a set step size until the image subjected to the m-th iterative denoising process is selected, where i and n are positive integers less than or equal to m; the image subjected to the m-th iterative process among the n images is used as the target image, and the remaining n-1 images among the n images are used as the multiple noise images.

[0006] In one embodiment, the first text content is determined by: obtaining the image style type and image content category information input by the user; and generating the first text content based on a text generation model, the image style type and the image content category information.

[0007] In one embodiment, the displaying of interaction information based on the similarity includes: in response to the similarity being less than a first threshold, displaying first interaction information, the first interaction information indicating that the image content represented by the input information does not match the image content of the target image; in response to the similarity being greater than the first threshold and less than a second threshold, displaying second interaction information, the second interaction information indicating that the image content represented by the input information does not completely match the image content of the target image; and in response to the similarity being greater than the second threshold, displaying third interaction information, the third interaction information indicating that the image content matches the image content of the target image.

[0008] In one embodiment, the method further includes: in response to displaying the first interaction information or the second interaction information and the user clicking a prompt operation, displaying a second image, wherein the second image is an image among the multiple noise images that is different from the first image and has a greater number of iterations than the first image; obtaining second input information representing image content input by the user based on the second image, and determining the similarity between the image content represented by the second input information and the image content of the target image; based on the similarity, displaying interaction information, wherein the interaction information is used to represent whether the image content represented by the second input information matches the image content of the target image; and repeating the above process until the displayed second image is the target image.

[0009] In one embodiment, before determining the similarity between the image content represented by the first input information and the image content of the target image, the method further includes: determining that a matching degree between second text content corresponding to the first input information and the first text content is lower than a threshold.

[0010] In one embodiment, the method further includes: in response to a matching degree between the second text content corresponding to the first input information and the first text content being greater than a threshold, displaying third interaction information, wherein the third interaction information indicates that the image content matches the image content of the target image.

[0011] According to a second aspect of an embodiment of the present disclosure, there is provided a graphic-text interaction device, comprising: a display unit for displaying a first image in an image set, wherein the image set includes a plurality of noise images obtained by iteratively denoising a random noise image for different numbers of iterations and a target image after the denoising process, wherein the first image is one of the plurality of noise images. A processing unit for obtaining first input information representing image content input by a user based on the first image, and determining the similarity between the image content represented by the first input information and the image content of the target image. The display unit is further configured to display interaction information based on the similarity, wherein the interaction information is configured to represent whether the image content represented by the first input information matches the image content of the target image.

[0012] In one embodiment, the image set is determined in the following manner: the first text content and the random noise image are input into a diffusion model, and the random noise image is subjected to m iterative denoising processes, where m is a positive integer greater than 1; starting from the i-th iterative process, n denoised images are selected in sequence according to a set step size until the image subjected to the m-th iterative denoising process is selected, where i and n are positive integers less than or equal to m; the image subjected to the m-th iterative process among the n images is used as the target image, and the remaining n-1 images among the n images are used as the multiple noise images.

[0013] In one embodiment, the first text content is determined by: obtaining the image style type and image content category information input by the user; and generating the first text content based on a text generation model, the image style type and the image content category information.

[0014] In one embodiment, the display unit displays interaction information based on the similarity in the following manner: in response to the similarity being less than a first threshold, first interaction information is displayed, the first interaction information indicating that the image content represented by the input information does not match the image content of the target image; in response to the similarity being greater than the first threshold and less than a second threshold, second interaction information is displayed, the second interaction information indicating that the image content represented by the input information does not completely match the image content of the target image; in response to the similarity being greater than the second threshold, third interaction information is displayed, the third interaction information indicating that the image content matches the image content of the target image.

[0015] In one embodiment, the display unit is further used to: in response to displaying the first interaction information or the second interaction information and the user clicking a prompt operation, display a second image, where the second image is an image among the multiple noise images that is different from the first image and has a greater number of iterations than the first image; obtain second input information representing the image content input by the user based on the second image, and determine the similarity between the image content represented by the second input information and the image content of the target image; based on the similarity, display interaction information, where the interaction information is used to represent whether the image content represented by the second input information matches the image content of the target image; repeat the above process until the displayed second image is the target image.

[0016] In one embodiment, before determining the similarity between the image content represented by the first input information and the image content of the target image, the processing unit is further used to: determine whether the matching degree between the second text content corresponding to the first input information and the first text content is lower than a threshold.

[0017] In one embodiment, the processing unit is further used to: in response to the matching degree between the second text content corresponding to the first input information and the first text content being greater than a threshold, display third interaction information, wherein the third interaction information indicates that the image content matches the image content of the target image.

[0018] According to a third aspect of an embodiment of the present disclosure, a graphic-text interaction device is provided, including:

[0019] A processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute the graphic-text interaction method described in the first aspect or any one of the embodiments of the first aspect.

[0020] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor of a terminal, the terminal can perform the method described in the first aspect or any one of the implementations of the first aspect.

[0021] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: by displaying the first image obtained during the iterative denoising process and obtaining the first input information input by the user based on the first image, interactive information is displayed based on the similarity between the image content represented by the first input information and the image content of the target image, the intelligence level of the displayed interactive information is enhanced, the interactivity of the graphic and text interaction process is enhanced, and the image set is obtained during the iterative denoising process, and there is no need to pre-store the image library in advance, which saves storage space, improves the randomness of the interaction, and enhances the user's interactive experience.

[0022] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0024] Figure 1 The figure is a flowchart of a method for interacting with images and texts according to an exemplary embodiment.

[0025] Figure 2 The figure is a flowchart of a method for acquiring an image set according to an exemplary embodiment.

[0026] Figures 3a to 3e FIG. 1 is a schematic diagram of obtaining an iterative denoising-processed image according to an exemplary embodiment of the present disclosure.

[0027] Figure 4 The figure is a flowchart of a method for obtaining first text content according to an exemplary embodiment.

[0028] Figure 5 The figure is a flowchart of a method for displaying interactive information according to an exemplary embodiment.

[0029] Figure 6 The figure is a flowchart of a method for displaying interactive information according to an exemplary embodiment.

[0030] Figure 7 The figure is a flowchart of a method for interacting with images and texts according to an exemplary embodiment.

[0031] Figure 8 The figure is a flowchart of a method for interacting with images and texts according to an exemplary embodiment.

[0032] Figure 9 The figure is a schematic diagram of a process of a graphic-text interactive application according to an exemplary embodiment of the present disclosure.

[0033] Figure 10 The figure is a block diagram showing a device for image-text interaction according to an exemplary embodiment.

[0034] Figure 11 It is a block diagram of a device for image-text interaction according to an exemplary embodiment. DETAILED DESCRIPTION

[0035] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present disclosure.

[0036] In the accompanying drawings, the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be understood as limiting the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure. The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0037] The image-text interaction method provided in the embodiments of the present disclosure is applied to the field of artificial intelligence generated content (AIGC). The image-text interaction method provided in the embodiments of the present disclosure is mainly used to generate multiple images containing a target image from a random noise image using a noise reduction model, determine the similarity between the user input information and the target image, and display corresponding interactive information based on the similarity.

[0038] In related technologies, large-scale image generation models are primarily used to directly generate an image based on user input, resulting in a simple interaction process. Furthermore, most related image-text interaction applications process an image, displaying it as an outline or blurring it. The user's input must completely match the target answer to produce a correct result. Furthermore, to ensure that users can guess correctly, text must be kept very short, and image content is limited, impacting the user's interactive experience.

[0039] In light of this, the disclosed embodiments provide a method for image-text interaction that decodes multiple latent space feature maps generated during the diffusion model inference process, allowing users to guess images. This overcomes the shortcomings of existing image guessing applications, which can only guess images based on a single feature, such as an image outline. By using the model to generate an image, the model determines the similarity between the image and the user input, and provides interaction information based on this similarity, this method solves the problem of previously relying solely on a perfect match with the set text. This method increases the randomness and fun of the interaction process, reduces the limitations of the user interaction process, and enhances the user's interactive experience.

[0040] Figure 1 is a flow chart showing a method for interacting with text and images according to an exemplary embodiment. Figure 1As shown, the following steps are included.

[0041] In step S11, a first image in an image set is displayed. The image set includes multiple noise images corresponding to different numbers of iterations of iterative denoising of a random noise image and a target image after the denoising process. The first image is one of the multiple noise images.

[0042] In the disclosed embodiments, a random noise image may be an image randomly generated by a model based on initial values ​​for initializing model parameters or generating random numbers, with the parameters or random numbers used to generate the random noise image varying each time. It should be understood that if different models are used to generate the random noise image, the parameters or random numbers used to generate the random noise image may be the same.

[0043] In the disclosed embodiment, by performing multiple iterative denoising processes on a random noise image, the target image finally obtained after the denoising process meets the preset requirements, and in the process of multiple iterative denoising, each denoising image is related to the preset required content, but the similarity between the content contained in the images obtained with different numbers of iterations and the preset required content is different, and the similarity between the image and the preset required content increases with the increase in the number of iterations. That is, by performing iterative denoising on a random noise image, not only can the target image after multiple iterative denoising operations be obtained, but also multiple intermediate noise images during the denoising operation can be obtained, one of which is selected as the first image, the first image contains more noise information than the target image, and the similarity between the first image and the preset requirements is lower than the similarity between the target image and the preset requirements. By displaying the first image in the image set, the electronic device can display image content that is similar to the target image but different from the target image.

[0044] In step S12, first input information representing image content input by the user based on the first image is obtained, and the similarity between the image content represented by the first input information and the image content of the target image is determined.

[0045] In an embodiment of the present disclosure, the user can input content for representing the image content based on the first image displayed by the electronic device, that is, the first input information content. It should be understood that the first input information content can be text content, or voice content, etc. In an embodiment of the present disclosure, if the first input content is text content, the first input content can be set as the first input information representation image content; if the first input information content is non-text content, the first input information needs to be converted into text content, and the obtained text content is set as the first input information representation image content. For example, if the first input information is voice information, the voice information can be converted into text content by using a model with a voice-to-text function, and then the similarity is determined. The reception of multimodal information input by the user and the extraction of content are realized.

[0046] In the disclosed embodiments, a pre-trained similarity matching model can be used to determine the similarity between the image content represented by the first input information and the image content of the target image. For example, a Chinese text-to-image matching model based on the Bert structure can be used to match the image content represented by the first input information with the image content of the target image to obtain a matching value, i.e., similarity information, between the image content represented by the first input information and the image content of the target image.

[0047] In step S13 , based on the similarity, interaction information is displayed, where the interaction information is used to indicate whether the image content of the first input information matches the image content of the target image.

[0048] In the embodiment of the present disclosure, after obtaining the similarity between the image content represented by the first input information and the image content of the target image, the interval to which the similarity belongs can be obtained by comparing with a preset similarity threshold, and the interactive information that needs to be displayed can be determined. For example, a first similarity threshold can be set. If the similarity between the image content represented by the first input information and the image content of the target image is greater than or equal to the first similarity threshold, the interactive information "the input information matches the target image" is displayed; if the similarity between the image content represented by the first input information and the image content of the target image is less than the first similarity threshold, the interactive information "the input information does not match the target image" is displayed. By displaying the corresponding interactive information based on the first input information input by the user, intelligent interaction between the user and the electronic device is achieved.

[0049] In the disclosed embodiment, a denoising model is used to iteratively denoise a random noise image to obtain a target image that has undergone denoising and a first image that has not undergone denoising but contains some characteristic image content of the target image. The first image is displayed and the user's first input information is obtained. Based on the similarity between the image content represented by the first input information and the image content of the target image, interactive information is displayed. The user does not need to completely guess the text information corresponding to the image. The similarity between the image content represented by the first input information and the image content of the target image is used to intelligently display the interactive information. This allows the user to interact based on the first image and the target image. Both the first image and the target are obtained by denoising the random noise image, eliminating the need to prepare an image library, saving storage resources, and enhancing the randomness and fun of the image-text interaction due to the low repetition rate of the image obtained by denoising.

[0050] In the embodiment of the present disclosure, the inverse process of the diffusion model can be used to remove noise from the image, and the target image can be acquired through multiple iterative denoising processes.

[0051] Figure 2 FIG. 1 is a flow chart of a method for acquiring an image set according to an exemplary embodiment. Figure 2 As shown, the following steps are included.

[0052] In step S21 , the first text content and the random noise image are input into the diffusion model, and the random noise image is subjected to m iterative denoising processes.

[0053] In the disclosed embodiment, m is a positive integer greater than 1. Diffusion Models are a type of machine learning model, particularly latent variable models for image processing and computer vision tasks. Diffusion models can implement a forward process (diffusion process) of gradually adding noise to an image until the image becomes completely random, and a reverse process (inverse diffusion process) of gradually removing noise from the noisy image to restore the original image. In the disclosed embodiment, the reverse process of the diffusion model is mainly used to perform m iterative denoising on the random noise image. Through m iterative denoising, the noise in the image after each iterative denoising is less than that in the image after the previous denoising.

[0054] It should be understood that the impact of noise information on the image in the embodiments of the present disclosure can be not only the clarity, but also the degree of similarity between the image content and the features of the first text content. The more times the denoising process is iterated, the higher the degree of similarity between the denoised image and the first text content.

[0055] In step S22 , starting from the i-th iteration, n denoised images are selected in sequence according to a set step size until the image denoised by the m-th iteration is selected.

[0056] In the disclosed embodiment, i and n are positive integers less than or equal to m. By selecting n denoised images after i denoising iterations during m denoising iterations, and then selecting n denoised images according to a set step size, n denoised images are acquired. For example, if m is 20, i is 1, and n is 5, the step size can be determined to be 4, and 5 denoised images, the 5th, 9th, 13th, 17th, and 20th denoising iterations, are selected. By selecting n denoised images according to a set step size, the number of acquired images and the degree of denoising of the acquired images can be controlled.

[0057] It should be understood that for iterative noise reduction processing that is not selected, the model can only generate feature information after denoising, and does not decode the feature information to generate the corresponding image, thereby reducing the pressure on electronic equipment to decode and store.

[0058] In step S23, the image processed in the mth iteration among the n images is used as the target image, and the remaining n-1 images among the n images are used as a plurality of noise images.

[0059] In the embodiment of the present disclosure, the image set is composed of n acquired images. The image most similar to the first text content, that is, the image after the mth iteration of denoising, is used as the target image, and the remaining images are used as multiple noise images to determine the first image.

[0060] In an exemplary embodiment, assume that the number of iterative denoising processes m performed on a random noise image is 20, the starting point i of the iterative process is 0, and the number of denoised images obtained is 5. Then, the images obtained after the 4th, 8th, 12th, 16th, and 20th denoising processes are obtained and stored in the image set, as shown in FIG. Figures 3a to 3e shown. Figures 3a to 3e is a schematic diagram of obtaining an iterative denoising image according to an exemplary embodiment of the present disclosure, Figures 3a to 3e Different images contain varying amounts of noise, have varying image content, and exhibit varying degrees of similarity to the target image. Due to the random nature of the noise, each iteration yields a fresh and dynamic image, potentially drastically changing the image content and enhancing the user experience. Furthermore, the images in the image collection are generated in real time by the iterative denoising model, eliminating the need for a pre-prepared image library. This results in a lower repetition rate than images selected from a conventional image library.

[0061] In the embodiment of the present disclosure, the step size is set by iterative denoising times and the number n of selected images, and n images are acquired based on the set step size, and an image set is formed to determine the target image and multiple noise images.

[0062] In the embodiment of the present disclosure, the first text content is used to instruct the diffusion model to perform denoising on the random noise image, so as to control the diffusion model to output an image consistent with the first text content.

[0063] Figure 4 is a flow chart showing a method for obtaining first text content according to an exemplary embodiment. Figure 4 As shown, the following steps are included.

[0064] In step S31, the image style type and image content category information input by the user are obtained.

[0065] In the disclosed embodiment, the image style type is the artistic style or characteristics displayed by the image, such as realism, abstraction, oil painting, watercolor, comics, etc. The image content category information is information that classifies or labels the content displayed by the image, such as natural scenery, animals, human sports activities, etc. Through the image style type and image content category information, the style type and content category of the first text content and the generated target image can be determined, which enables users to customize the interactive experience according to their preferences. In addition, different image style types and image content categories may have different difficulties for users. By providing users with the choice of image style type and image content category information, it is convenient for users to adjust the difficulty level according to their own needs, thereby improving the user's interactive experience.

[0066] In step S32, first text content is generated based on the text generation model, the image style type and the image content category information.

[0067] In the disclosed embodiment, the image style type, image content category information, and preset template information are combined into input text and input into a text generation model to realize the generation of the first text content. For example, the text generation model can be a large language model (LLM), the image style type is animation, the image content category information is animals, and the input template is "Please write a short sentence containing (), and the image style is []", the small brackets in the input template are used to fill in the content of the image content category information, and the square brackets are used to fill in the image style type, so the input text content of "Please write a short sentence containing animals, and the image style is animation" can be obtained, and the input text content is input into the large language model to obtain the output result of "an animated image of an elephant dancing ballet", and the output result is the first text content.

[0068] It should be understood that the image style type can also be used in the process of generating images. For example, the Lora style model is used to perform style processing on images in the image collection so that the image style in the image collection meets the image style type selected by the user.

[0069] In the embodiment of the present disclosure, multiple similarity thresholds can be set to determine the interactive information to be displayed by comparing the similarity between the image content represented by the first input information and the image content of the target image with the magnitude relationship of the multiple similarity thresholds.

[0070] Figure 5 is a flow chart showing a method for displaying interactive information according to an exemplary embodiment. Figure 5 As shown, the following steps are included.

[0071] In step S41 , in response to the similarity being less than a first threshold, the first interaction information is displayed.

[0072] In the embodiment of the present disclosure, two similarity thresholds can be set, and based on the magnitude relationship between the similarity and the two similarity thresholds, three different types of interaction information are determined to be displayed. In the embodiment of the present disclosure, the two similarity thresholds are respectively a first threshold and a second threshold, and the first threshold is smaller than the second threshold.

[0073] In the disclosed embodiment, the first interaction information indicates that the image content represented by the input information does not match the image content of the target image. When the detected similarity is less than a first threshold, it indicates that the similarity between the image content represented by the first input information and the image content of the target image is less than the first and second thresholds, i.e., the current similarity does not meet the minimum requirement. By displaying the first interaction information, the electronic device is controlled to display interaction information indicating that the image content represented by the input information does not match the image content of the target image. For example, the first interaction information may be "Sorry, the content you entered does not match the target image."

[0074] In step S42 , in response to the similarity being greater than the first threshold and less than the second threshold, the second interaction information is displayed.

[0075] In the disclosed embodiment, if the detected similarity is greater than a first threshold and less than a second threshold, a second interaction message is displayed. The second interaction message indicates that the image content represented by the input information does not completely match the image content of the target image. In other words, there is a certain tendency for the image content represented by the input information to match the image content of the target image, but it has not yet reached the standard of a complete match. For example, the second interaction message may read, "The content you entered is close to the content of the target image, but it does not completely match."

[0076] In step S43 , in response to the similarity being greater than the second threshold, the third interaction information is displayed.

[0077] In this disclosed embodiment, the second threshold is used to determine whether the image content represented by the input information and the image content of the target image meet the matching requirements. If the detected similarity is greater than the second threshold, it indicates that the image content represented by the input information matches the image content of the target image, and a third interaction message is displayed. The third interaction message indicates that the image content matches the image content of the target image. For example, the third interaction message may be "Congratulations, your input content matches the target image."

[0078] In an embodiment of the present disclosure, if the similarity detected is less than the second threshold, a prompt operation button can be displayed while displaying the first interaction information or the second interaction information to provide the user with a noise image that is more similar to the target image content, thereby reducing the difficulty of interaction.

[0079] Figure 6is a flow chart showing a method for displaying interactive information according to an exemplary embodiment. Figure 6 As shown, the following steps are included.

[0080] In step S51, in response to displaying the first interaction information or the second interaction information and the user clicking the prompt operation, a second image is displayed, which is an image among multiple noise images that is different from the first image and has a greater number of iterations than the first image.

[0081] In an embodiment of the present disclosure, when it is detected that the similarity between the image content represented by the first input information input by the user based on the first image and the image content of the target image is less than a second threshold, a prompt operation button can be displayed while displaying the interactive information. When the user clicks the prompt operation, the second image in the image set is used to replace the first image for display, thereby providing the user with image content with less noise and reducing the difficulty of interaction.

[0082] In step S52, second input information representing image content input by the user based on the second image is obtained, and the similarity between the image content represented by the second input information and the image content of the target image is determined.

[0083] In the embodiment of the present disclosure, the second image is any image in the image set having less noise than the first image. After the second image is displayed, the user can input a second input language for representing the content of the second image based on their understanding of the second image. The electronic device determines the similarity between the image content represented by the second input language and the image content of the target image based on the acquired second input language. It should be understood that the determination of the similarity between the image content represented by the second input information and the image content of the target image adopts the same method as the determination of the similarity between the image content represented by the first input information and the image content of the target image. Again, no further details will be given, and reference may be made to the relevant description of determining the similarity between the image content represented by the first input information and the image content of the target image.

[0084] In step S53, the interaction information is displayed based on the similarity.

[0085] In the disclosed embodiments, the interaction information is used to indicate whether the image content represented by the second input information matches the image content of the target image. It should be understood that the process of displaying the interaction information based on the similarity between the image content represented by the second input information and the image content of the target image is the same as the process of displaying the interaction information based on the similarity between the image content represented by the first input information and the image content of the target image. This description will not be repeated here; reference may be made to the relevant description regarding determining the similarity between the image content represented by the first input information and the image content of the target image.

[0086] In step S54, the above process is repeated until the displayed second image is the target image.

[0087] In the disclosed embodiment, if a user clicks a prompt when interactive information about a second image is displayed, a noise image with less noise than the current second image or a greater number of iterations than the current second image is selected from the image set and used to replace the current second image. In response to the user clicking the prompt, an image from the image set is selected to replace the second image.

[0088] It should be understood that in response to the second image being the target image, when the user clicks the prompt operation, there is no noise image in the image set whose noise is less than the current second image or the number of iterations is greater than the current second image. At this time, the electronic device displays the first text content.

[0089] It should be understood that in the embodiments of the present disclosure, while displaying the interactive information, in addition to displaying the content used to trigger the prompt operation, other display content can also be set based on the needs of the developer. For example, while displaying the interactive information and the prompt button, a button to end the interaction can also be displayed to meet the different interactive needs of users.

[0090] In the embodiment of the present disclosure, before determining the similarity between the image content represented by the user's input information and the image content of the target image, the user's input information may be matched with the first text content.

[0091] Figure 7 is a flow chart showing a method for interacting with text and images according to an exemplary embodiment. Figure 7 As shown, the following steps are included.

[0092] Figure 7 The steps S61 and S63 in Figure 1 Steps S11 and S13 are the same and will not be described again here. Please refer to the relevant description of the above embodiment. Only the differences will be described below.

[0093] In step S62, first input information representing image content input by the user based on the first image is obtained, it is determined that the matching degree between the second text content corresponding to the first input information and the first text content is lower than a threshold, and the similarity between the image content represented by the first input information and the image content of the target image is determined.

[0094] In the disclosed embodiment, before determining the similarity between the image content represented by the first input information and the image content of the target image, the corresponding second text content obtained using the first input information input by the user can be matched with the first text content. When the matching degree between the second text content and the first text content is lower than a preset threshold, the similarity between the image content represented by the first input information and the image content of the target image can be further determined. This avoids interactive information display errors caused by the generated target image not matching the first text content.

[0095] In the embodiment of the present disclosure, if it is detected that the matching degree between the second text content and the first text content is greater than a threshold, there is no need to determine the similarity between the first input information and the target image.

[0096] Figure 8 is a flow chart showing a method for interacting with text and images according to an exemplary embodiment. Figure 8 As shown, the following steps are included.

[0097] Figure 8 The steps in step S71 in Figure 7 The same as step S61 in , which will not be repeated here. Please refer to the relevant description of the above embodiment. Only the differences are described below.

[0098] In step S72, first input information representing image content input by the user based on the first image is obtained, and in response to a matching degree between the second text content corresponding to the first input information and the first text content being greater than a threshold, third interaction information is displayed.

[0099] In the disclosed embodiment, the third interaction information indicates that the image content matches the image content of the target image. If the degree of match between the second text content and the first text content is detected to be greater than a preset threshold, it indicates that the second text content matches the first text content. Since the target image is generated based on the first text content, there is no need to further determine the similarity between the image content represented by the first input information and the image content of the target image. The third interaction information can be displayed and the interaction ends. This avoids the waste of resources caused by additionally determining the similarity between the image content represented by the first input information and the image content of the target image.

[0100] In an exemplary embodiment, the image-text interaction method can be applied on a terminal device in the form of a guessing image app, an interactive game, etc. In this exemplary embodiment, the guessing image software is used as an example for explanation. Figure 9 As shown, Figure 9 The figure is a schematic diagram of a process of a graphic-text interactive application according to an exemplary embodiment of the present disclosure.

[0101] exist Figure 9In the example, the first interactive information is "error", the second interactive information is "close", and the third interactive information is "error". The first input information or the second input information is specifically a guessing text. The similarity between the image content represented by the first input information or the second input information and the image content of the target image is denoted as λ, the first threshold is denoted as λ2, the second threshold is denoted as λ1, and the n images in the image set obtained during the m-iterative denoising process are denoted as m1, m2...m respectively. n , where m1, m2...m n-1 are multiple noise images, m n is the target image.

[0102] exist Figure 9 In the image guessing software, the user triggers the guessing function of the interactive application or directly selects the image style and image category in the terminal's guessing software. The terminal randomly generates prompt words that match the image category through the text generation model, such as "an elephant is dancing ballet". The terminal performs image denoising based on the prompt words through the Wenshengtu model, where the Wenshengtu model includes at least a diffusion model similar to Stablediffusion and a Lora style model. After encoding the prompt words using the text encoder, the encoded data is input into the diffusion model in the Wenshengtu model. The input of the diffusion model also includes a random noise map, and it iterates m steps of denoising. Select m1, m2...m from the m steps. n The denoised latent space graph of the step is decoded by the decoder to form images m1, m2...m n , where picture m n As the final picture. i (i=1,2……n) is displayed to the user, and the user is instructed to guess the picture. The user guesses the content of the picture based on the picture, and inputs the guessed content into the interactive application in the form of voice or text. The user inputs the second text content (or inputs the voice and converts it into the second text content) to match the final picture with the image-text similarity. The matching model used is a Chinese text and picture matching model based on the Bert structure. The similarity thresholds λ1 and λ2 are set. When the model outputs the similarity matching value λ>λ1, the interactive application outputs "Guess the picture correctly" to the user, and the guessing ends. It should be understood that the degree of image-text matching can also be output based on the image-text similarity value, for example, "The user's answer is 60% close to the image content." When λ2<λ<λ1, Xiao Ai outputs "close" to the user, and when the model outputs the similarity matching value λ<λ2, Xiao Ai outputs "Guess the picture incorrectly" to the user. The user can choose to continue entering the guessing text and enter step 8; or choose not to continue entering the guessing text and click the "Hint" button to change the picture m i+1 Display it to the user, and so on, until the value of i reaches n, ending the interaction.

[0103] In the disclosed embodiment, by displaying a first image from among multiple noisy images obtained at different iterations during an iterative denoising process in an image collection, first input information representing image content input by a user based on the first image is obtained, and interaction information is displayed based on the similarity between the image content represented by the first input information and the image content of the target image. This makes the images displayed during the interaction more creative, and intelligently determines interaction information based on the similarity between the image and text, thereby enhancing the user's interactive experience.

[0104] Based on the same concept, an embodiment of the present disclosure also provides a graphic-text interaction device.

[0105] It is understandable that the graphic device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. In combination with the units and algorithm steps of each example disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.

[0106] Figure 10 FIG is a block diagram of a graphic-text interaction device 100 according to an exemplary embodiment. Figure 10 , the device includes a display unit 101 and a processing unit 102.

[0107] The display unit 101 is used to display the first image in the image set, which includes multiple noise images obtained by iterative denoising of the random noise image corresponding to different numbers of iterations and a target image after the denoising process, and the first image is one of the multiple noise images.

[0108] The processing unit 102 is configured to obtain first input information representing image content input by a user based on a first image, and determine a similarity between the image content represented by the first input information and image content of a target image.

[0109] The display unit 101 is further configured to display interaction information based on the similarity, where the interaction information is used to indicate whether the image content represented by the first input information matches the image content of the target image.

[0110] In one embodiment, the image set is determined as follows:

[0111] The first text content and the random noise image are input into the diffusion model, and the random noise image is subjected to m iterative denoising processes, where m is a positive integer greater than 1. Starting from the i-th iterative process, n denoised images are selected in sequence according to the set step size until the image subjected to the m-th iterative denoising process is selected, where i and n are positive integers less than or equal to m. The image subjected to the m-th iterative process among the n images is used as the target image, and the remaining n-1 images among the n images are used as multiple noise images.

[0112] In one embodiment, the first text content is determined in the following manner:

[0113] Get the image style type and image content category information input by the user;

[0114] First text content is generated based on the text generation model, the image style type, and the image content category information.

[0115] In one embodiment, the display unit 101 displays the interaction information based on the similarity in the following manner:

[0116] In response to the similarity being less than a first threshold, first interaction information is displayed, the first interaction information indicating that the image content represented by the input information does not match the image content of the target image; in response to the similarity being greater than the first threshold and less than a second threshold, second interaction information is displayed, the second interaction information indicating that the image content represented by the input information does not completely match the image content of the target image; in response to the similarity being greater than the second threshold, third interaction information is displayed, the third interaction information indicating that the image content matches the image content of the target image.

[0117] In one embodiment, the display unit 101 is further configured to:

[0118] In response to displaying the first interaction information or the second interaction information and the user clicking on the prompt operation, a second image is displayed, where the second image is an image among multiple noise images that is different from the first image and has a greater number of iterations than the first image; second input information representing the image content input by the user based on the second image is obtained, and the similarity between the image content represented by the second input information and the image content of the target image is determined; based on the similarity, interaction information is displayed, where the interaction information is used to represent whether the image content represented by the second input information matches the image content of the target image; the above process is repeated until the displayed second image is the target image.

[0119] In one embodiment, before determining the similarity between the image content represented by the first input information and the image content of the target image, the processing unit 102 is further configured to:

[0120] It is determined that a degree of matching between the second text content corresponding to the first input information and the first text content is lower than a threshold.

[0121] In one embodiment, the processing unit 102 is further configured to:

[0122] In response to a matching degree between the second text content corresponding to the first input information and the first text content being greater than a threshold, third interaction information is displayed, where the third interaction information indicates that the image content matches the image content of the target image.

[0123] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0124] Figure 11 2 is a block diagram illustrating an apparatus 200 for text-graphic interaction according to an exemplary embodiment. For example, the apparatus 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0125] Reference Figure 11 , apparatus 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .

[0126] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.

[0127] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0128] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 200.

[0129] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0130] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.

[0131] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0132] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect changes in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and temperature changes of the device 200. The sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0133] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0134] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0135] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, which can be executed by the processor 220 of the apparatus 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0136] It is understood that in this disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of related objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0137] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.

[0138] It is further understood that, unless otherwise specified, “connection” includes a direct connection where there are no other components between the two elements, and also includes an indirect connection where there are other elements between the two elements.

[0139] It is further understood that although operations are described in a particular order in the drawings in the embodiments of the present disclosure, this should not be construed as requiring that the operations be performed in the particular order shown or in a serial order, or that all of the operations shown be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.

[0140] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0141] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.

Claims

1. A graphic-text interaction method, characterized in that: include: Displaying a first image in an image set, the image set including a plurality of noise images obtained by performing iterative denoising on a random noise image for different numbers of iterations and a target image after the denoising process, the first image being one of the plurality of noise images; obtaining first input information representing image content input by a user based on the first image, and determining a similarity between the image content represented by the first input information and image content of the target image; Based on the similarity, interaction information is displayed, where the interaction information is used to indicate whether image content represented by the first input information matches image content of the target image.

2. The image-text interaction method according to claim 1, characterized in that: The image set is determined in the following manner: Inputting the first text content and the random noise image into the diffusion model, and performing m-iterative denoising processes on the random noise image, where m is a positive integer greater than 1; Starting from the i-th iteration, n denoised images are selected in sequence according to a set step size until the m-th denoised image is selected, where i and n are positive integers less than or equal to m; The image processed in the mth iteration among the n images is used as the target image, and the remaining n-1 images among the n images are used as the multiple noise images.

3. The image-text interaction method according to claim 2, characterized in that: The first text content is determined in the following manner: Get the image style type and image content category information input by the user; First text content is generated based on a text generation model, the image style type, and the image content category information.

4. The image-text interaction method according to claim 1, characterized in that: The displaying of interaction information based on the similarity includes: In response to the similarity being less than a first threshold, displaying first interaction information, wherein the first interaction information indicates that image content represented by the input information does not match image content of the target image; In response to the similarity being greater than a first threshold and less than a second threshold, displaying second interaction information, wherein the second interaction information indicates that content of the image represented by the input information is not completely matched with content of the image of the target image; In response to the similarity being greater than a second threshold, third interaction information is displayed, where the third interaction information indicates that image content matches image content of the target image.

5. The image-text interaction method according to claim 4, characterized in that: The method further comprises: In response to displaying the first interaction information or the second interaction information and the user clicking a prompt operation, displaying a second image, where the second image is an image among the multiple noise images that is different from the first image and has a greater number of iterations than the first image; obtaining second input information representing image content input by a user based on the second image, and determining a similarity between the image content represented by the second input information and the image content of the target image; Based on the similarity, displaying interaction information, where the interaction information is used to indicate whether image content represented by the second input information matches image content of the target image; The above process is repeated until the displayed second image is the target image.

6. The image-text interaction method according to claim 1 or 2, characterized in that: Before determining the similarity between the image content represented by the first input information and the image content of the target image, the method further includes: It is determined that a degree of matching between the second text content corresponding to the first input information and the first text content is lower than a threshold.

7. The image-text interaction method according to claim 4 or 6, characterized in that: The method further comprises: In response to a matching degree between the second text content corresponding to the first input information and the first text content being greater than a threshold, third interaction information is displayed, where the third interaction information indicates that image content matches image content of the target image.

8. A graphic-text interaction device, characterized in that: include: a display unit configured to display a first image in an image set, the image set including a plurality of noise images obtained by performing iterative denoising on a random noise image for different numbers of iterations and a target image after the denoising process, the first image being one of the plurality of noise images; a processing unit, configured to obtain first input information representing image content input by a user based on the first image, and determine a similarity between the image content represented by the first input information and image content of the target image; The display unit is further configured to display interaction information based on the similarity, where the interaction information is configured to indicate whether the image content represented by the first input information matches the image content of the target image.

9. A graphic-text interaction device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores instructions, and when the instructions in the storage medium are executed by a processor of the terminal, the terminal is enabled to execute the method according to any one of claims 1 to 7.