An image generation method, apparatus, electronic device, and storage medium

CN117932181BActive Publication Date: 2026-08-07BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING QIYI CENTURY SCI & TECH CO LTD
Filing Date
2024-02-05
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]然而,网络资源的主题图像为固定的图像,而固定的图像可能并不是用户感兴趣的图像,则电子设备基于固定的图像向用户展示网络资源的有效性不高

Benefits of technology

[0054]Based on the above processing, since the second text is used to describe the target online resource and the historical online resources matched with the target account, the thematic image generated based on the second text can represent both the target and historical online resources. That is, the thematic image can effectively represent the content of the target online resource and the content of the historical online resources that the user is interested in. When thematic images are subsequently recommended to the user, the user can browse thematic images that are of interest to them and can represent the content of the target online resource. This increases the probability that the user will choose the thematic image that interests them, thereby improving the effectiveness of electronic devices in displaying online resources to the user based on thematic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117932181B_ABST
    Figure CN117932181B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image generation method and device, electronic equipment and storage medium, and relate to the technical field of image processing. The method comprises: obtaining a first text corresponding to a target network resource, and obtaining an interest label corresponding to a target account; wherein the first text is used to describe the target network resource; the interest label is used to describe a historical network resource matched with the target account; based on the first text and the interest label, a second text corresponding to the target account for the target network resource is generated; wherein the second text is used to describe the target network resource and the historical network resource; a guide text is generated based on the second text, and the guide text is input into a preset image generation model to obtain an image output by the image generation model as a theme image of the target network resource matched with the target account, which can improve the effectiveness of the electronic equipment in showing the network resource to the user based on the theme image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of computer technology, client applications can provide users with more and more functions. For example, users can browse online resources through client applications, such as movies, TV series, novels, comics, etc.

[0003] In related technologies, the client can display an identifier for a network resource on the display interface. For example, the identifier for a network resource can be a theme image of the network resource, such as a video frame from the network resource; or, the theme image can be a poster. When the client detects that a user has selected a theme image of a network resource, it displays the selected network resource.

[0004] However, since the main images of online resources are fixed images, and fixed images may not be images that users are interested in, the effectiveness of electronic devices in displaying online resources to users based on fixed images is not high. Summary of the Invention

[0005] The purpose of this invention is to provide an image generation method, apparatus, electronic device, and storage medium to improve the effectiveness of electronic devices in displaying network resources to users based on subject images. The specific technical solution is as follows:

[0006] In a first aspect of the present invention, an image generation method is provided, the method comprising:

[0007] Obtain the first text corresponding to the target network resource and obtain the interest tags corresponding to the target account; wherein, the first text is used to describe the target network resource; the interest tags are used to describe historical network resources that match the target account;

[0008] Based on the first text and the interest tags, a second text corresponding to the target account and the target network resource is generated; wherein, the second text is used to describe the target network resource and the historical network resource;

[0009] Based on the second text, guide text is generated and input into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0010] Optionally, the target network resource is a film and television resource;

[0011] The acquisition of the first text corresponding to the target network resource includes:

[0012] Obtain the third text corresponding to the target network resource, and the type tag of the target network resource; wherein, the third text is the script of the target network resource; the third text includes at least one of the following: the scene in the target network resource, the characters in the target network resource, and the lines of each character in the target network resource; based on the first preset keyword of the type tag, extract script fragments including at least one keyword from the third text to obtain a fourth text; perform text analysis on the fourth text to obtain a content summary of the target network resource, which serves as the first text.

[0013] Optionally, based on the first preset keywords of the type tags, the fourth text is obtained by extracting script fragments including at least one keyword from the third text, including:

[0014] The third text is divided according to the second preset keywords to obtain multiple script fragments;

[0015] For each script fragment, if the script fragment includes at least one first preset keyword corresponding to the type tag, the script fragment is determined to be the fourth text.

[0016] Optionally, the step of performing text analysis on the fourth text to obtain a summary of the target network resource, which serves as the first text, includes:

[0017] The fourth text is segmented using a word segmenter to obtain the words to be analyzed.

[0018] The words to be analyzed are input into the text analysis model to obtain a brief description of the target network resource output by the text analysis model, which serves as the first text.

[0019] Optionally, obtaining the interest tags corresponding to the target account includes:

[0020] Obtain interaction information corresponding to the target account; wherein, the interaction information is obtained based on the interaction instructions for the historical network resources; the historical network resources are network resources in the historical access records of the target account;

[0021] For each feature tag of the historical network resource, the matching degree between the target account and the feature tag is determined based on the interaction information; wherein, the feature tag of the historical network resource is used to describe the historical network resource;

[0022] Among the feature tags of the historical network resources, the feature tags with a matching degree greater than a preset threshold with the target account are identified as the interest tags corresponding to the target account.

[0023] Optionally, generating second text for the target network resource corresponding to the target account based on the first text and the interest tag includes:

[0024] The first text and the interest tag are concatenated to obtain the second text corresponding to the target account for the target network resource; or, the first text and the interest tag are input into a text fusion model to obtain the second text corresponding to the target account for the target network resource output by the text fusion model.

[0025] Optionally, the step of generating guiding text based on the second text and inputting the guiding text into a preset image generation model to obtain an image output by the image generation model as the theme image of the target network resource matching the target account includes:

[0026] Based on the second text, guide text is generated. The guide text and the preset theme image of the target network resource are input into a preset image generation model to obtain the image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0027] In a second aspect of the invention, an image generation apparatus is provided, the apparatus comprising:

[0028] The first acquisition module is used to acquire first text corresponding to the target network resource and interest tags corresponding to the target account; wherein, the first text is used to describe the target network resource; and the interest tags are used to describe historical network resources that match the target account.

[0029] The fusion module is used to generate a second text corresponding to the target account for the target network resource based on the first text and the interest tag; wherein the second text is used to describe the target network resource and the historical network resource;

[0030] The second acquisition module is used to generate guiding text based on the second text, and input the guiding text into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0031] Optionally, the target network resource is a film and television resource;

[0032] The first acquisition module is specifically used for:

[0033] Obtain the third text corresponding to the target network resource, and the type tag of the target network resource; wherein, the third text is the script of the target network resource; the third text includes at least one of the following: the scene in the target network resource, the characters in the target network resource, and the lines of each character in the target network resource; based on the first preset keyword of the type tag, extract script fragments including at least one keyword from the third text to obtain a fourth text; perform text analysis on the fourth text to obtain a content summary of the target network resource, which serves as the first text.

[0034] Optionally, the first acquisition module is specifically used for:

[0035] The third text is divided according to the second preset keywords to obtain multiple script fragments;

[0036] For each script fragment, if the script fragment includes at least one first preset keyword corresponding to the type tag, the script fragment is determined to be the fourth text.

[0037] Optionally, the first acquisition module is specifically used for:

[0038] The fourth text is segmented using a word segmenter to obtain the words to be analyzed.

[0039] The words to be analyzed are input into the text analysis model to obtain a brief description of the target network resource output by the text analysis model, which serves as the first text.

[0040] Optionally, the first acquisition module is specifically used for:

[0041] Obtain interaction information corresponding to the target account; wherein, the interaction information is obtained based on the interaction instructions for the historical network resources; the historical network resources are network resources in the historical access records of the target account;

[0042] For each feature tag of the historical network resource, the matching degree between the target account and the feature tag is determined based on the interaction information; wherein, the feature tag of the historical network resource is used to describe the historical network resource;

[0043] Among the feature tags of the historical network resources, the feature tags with a matching degree greater than a preset threshold with the target account are identified as the interest tags corresponding to the target account.

[0044] Optionally, the fusion module is specifically used for:

[0045] The first text and the interest tag are concatenated to obtain the second text corresponding to the target account for the target network resource; or, the first text and the interest tag are input into a text fusion model to obtain the second text corresponding to the target account for the target network resource output by the text fusion model.

[0046] Optionally, the second acquisition module is specifically used for:

[0047] Based on the second text, guide text is generated, and the guide text and the preset theme image of the target network resource are input into a preset image generation model to obtain the image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0048] In a third aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0049] Memory, used to store computer programs;

[0050] When a processor executes a program stored in memory, it implements any of the steps described in the first aspect above.

[0051] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements any of the image generation methods described above.

[0052] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the image generation methods described above.

[0053] This invention provides an image generation method in which an electronic device acquires first text corresponding to a target network resource and interest tags corresponding to a target account; the first text describes the target network resource; the interest tags describe historical network resources that match the target account; based on the first text and the interest tags, a second text corresponding to the target account and the target network resource is generated; the second text describes the target network resource and historical network resources; based on the second text, a guiding text is generated and input into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0054] Based on the above processing, since the second text is used to describe the target online resource and the historical online resources matched with the target account, the thematic image generated based on the second text can represent both the target and historical online resources. That is, the thematic image can effectively represent the content of the target online resource and the content of the historical online resources that the user is interested in. When thematic images are subsequently recommended to the user, the user can browse thematic images that are of interest to them and can represent the content of the target online resource. This increases the probability that the user will choose the thematic image that interests them, thereby improving the effectiveness of electronic devices in displaying online resources to the user based on thematic images. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0056] Figure 1 This is a first flowchart of an image generation method provided in an embodiment of the present invention;

[0057] Figure 2 This is a second flowchart of the image generation method provided in an embodiment of the present invention;

[0058] Figure 3 This is a third flowchart of the image generation method provided in the embodiments of the present invention;

[0059] Figure 4 This is a fourth flowchart of the image generation method provided in the embodiments of the present invention;

[0060] Figure 5 This is a fifth flowchart of the image generation method provided in an embodiment of the present invention;

[0061] Figure 6 A structural diagram of an image generation apparatus provided in an embodiment of the present invention;

[0062] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0063] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0064] In related technologies, the main image of network resources is a fixed image. However, a fixed image may not be an image that users are interested in. Therefore, the effectiveness of electronic devices in displaying network resources to users based on fixed images is not high.

[0065] To address the aforementioned problems, this invention provides an image generation method applied to an electronic device. The electronic device can be a server or a client. The electronic device can acquire first text corresponding to a target online resource and interest tags corresponding to a target account. The first text describes the target online resource; the interest tags describe historical online resources matching the target account. Then, based on the first text and interest tags, a second text corresponding to the target online resource and the target account is generated. The second text describes both the target online resource and historical online resources. Furthermore, guiding text is generated based on the second text and input into a preset image generation model to obtain an image output by the model, which serves as the theme image of the target online resource matching the target account. The theme image can effectively represent the content of the target online resource and the content of historical online resources of interest to the user. Subsequent recommendation of theme images to the user allows them to browse theme images that are of interest and represent the target online resource, increasing the probability of the user selecting such images. This improves the effectiveness of the electronic device in displaying online resources to the user based on theme images.

[0066] See Figure 1 , Figure 1 This is a first flowchart of an image generation method provided in an embodiment of the present invention. The method may include the following steps:

[0067] S101: Obtain the first text corresponding to the target network resource, and obtain the interest tags corresponding to the target account.

[0068] The first text describes the target online resource; the interest tags describe the historical online resources that match the target account.

[0069] S102: Based on the first text and interest tags, generate a second text corresponding to the target account and the target network resources.

[0070] The second text describes the target network resources and historical network resources.

[0071] S103: Generate guiding text based on the second text, and input the guiding text into the preset image generation model to obtain the image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0072] Based on the image generation method provided in this invention, since the second text is used to describe the target network resource and the historical network resources matching the target account, the thematic image generated based on the second text can represent both the target network resource and the historical network resources. That is, the thematic image can effectively represent the content of the target network resource and the content of the historical network resources that the user is interested in. When thematic images are subsequently recommended to the user, the user can browse thematic images that are of interest to them and can represent the content of the target network resource. This increases the probability that the user will choose the thematic image that interests them, thereby improving the effectiveness of electronic devices in displaying network resources to the user based on thematic images.

[0073] Regarding step S101, the target online resource can be a movie, TV series, novel, etc. The electronic device can display a theme image of the target online resource on the display interface as an identifier for the target online resource. When it detects that a user has selected a theme image of an online resource, the electronic device can determine that the user needs to browse that online resource, and then the electronic device can display the online resource to the user based on the theme image.

[0074] Taking a movie as an example, the movie's theme image can be any video frame from the movie, or it can be the movie poster. The movie's theme image can display the movie's main actors and the movie's name, etc.

[0075] Taking a novel as an example, the novel's main image can be its cover. This image can display the protagonist's portrait and the novel's title, among other things.

[0076] The first text of the target online resource describes the target online resource. This first text could be a brief summary of the target online resource's content, its title, etc.

[0077] For each online resource, its content can be manually summarized to obtain a first text describing the resource. Alternatively, the online resource can be analyzed by an electronic device to obtain the first text describing it.

[0078] In this embodiment, film and television resources are used as an example to illustrate the concept. Film and television resources are stories, such as movies, TV series, and animated films. These resources have corresponding scripts.

[0079] In one implementation, the electronic device can directly use a text analysis model to perform text analysis on the script of the target network resource, obtaining the first text describing the target network resource. The text analysis model can be a GPT (Generative Pre-trained Transformers) model.

[0080] In another implementation, to improve the efficiency of obtaining the first text, the electronic device can perform slice-level analysis on the script of the target network resource. That is, the electronic device can divide the script of the target network resource into multiple script fragments, then determine some script fragments from the multiple script fragments, and perform text analysis on the determined script fragments to obtain the first text.

[0081] exist Figure 1 Based on this, see Figure 2 Step S101 may include the following steps:

[0082] S1011: Obtain the third text corresponding to the target network resource, and the type label of the target network resource.

[0083] The third text is the script of the target network resource; the third text includes at least one of the following: scenes in the target network resource, characters in the target network resource, and lines of dialogue of each character in the target network resource.

[0084] S1012: Based on the first preset keyword of the type label, extract the script fragment containing at least one keyword from the third text to obtain the fourth text.

[0085] S1013: Perform text analysis on the fourth text to obtain a brief description of the target network resource, which is used as the first text, and obtain the interest tags corresponding to the target account.

[0086] The script for a film or television resource is a third-party text corresponding to that resource. The script can include characters, scenes, and dialogue from the film or television resource. For example, the script for a target online resource could be: "[Scene] Inside a spaceship; [Character 1] Captain: We are heading to an unknown galaxy. [Character 2] Mechanic: The ship's engines are malfunctioning! [Scene] Planetary surface; [Character 1] Crew Member A: The creatures on this planet look very strange. [Character 2] Crew Member B: We need to explore this planet."

[0087] The type tags of film and television resources can indicate the category to which the content of the film and television resources belong. For example, the type tags can be "science fiction", "horror", "comedy" etc.

[0088] For each type tag, a first preset keyword can be set for that type tag. For each online resource, if a script fragment of that online resource includes at least one of the first preset keywords of that online resource's type tag, it can be determined that the script fragment has a high probability of describing the content of that online resource. For example, when the type tag is "science fiction", the first preset keywords can include: "galaxy", "space", "unknown", "alien", "universe", etc.

[0089] Furthermore, electronic devices can extract script fragments containing at least one keyword from third-party text based on the type tags of the target network resource.

[0090] In one implementation, the electronic device can directly detect a first preset keyword in the third text. When the first preset keyword is detected, the script segment to which the first preset keyword belongs is extracted to obtain the fourth text. For example, the lines of the character to which the first preset keyword belongs can be extracted as the fourth text.

[0091] In another implementation, Figure 2 Based on this, see Figure 3 Step S1012 may include the following steps:

[0092] S10121: Divide the third text according to the second preset keywords to obtain multiple script fragments.

[0093] S10122: For each script fragment, if the script fragment includes at least one first preset keyword corresponding to the type tag, the script fragment is determined to be the fourth text.

[0094] The second preset keyword is a word or phrase in the script. Each script segment obtained by dividing the script according to the second preset keyword is a separate text paragraph. For example, the second preset keyword can be "[scene]", which means dividing the script according to scenes to obtain multiple script segments. Each script segment includes a scene, the characters in that scene, and the characters' lines, etc.; or, the second preset keyword can be "[character 1]", which means dividing the script according to the characters' dialogue to obtain multiple script segments. Each script segment includes lines related to character 1, the scene in which they are located, etc., but is not limited to these.

[0095] Furthermore, for each script fragment, the electronic device can detect whether the fragment includes at least one first preset keyword corresponding to the type tag. If the fragment includes at least one first preset keyword corresponding to the type tag, it can effectively describe the content of the target online resource. Therefore, the electronic device can identify the fragment as the fourth text. Subsequent text analysis of the fourth text yields a more accurate summary of the target online resource's content.

[0096] Based on the above processing, the electronic device only needs to perform text analysis on the fourth text. Since the fourth text is a script fragment extracted from the third text, it contains less content, making text analysis of the fourth text more efficient. This improves the efficiency of acquiring the first text, and consequently, the efficiency of image generation. Furthermore, because the fourth text can better describe the content of the target network resource, the accuracy of the content summary of the target network resource obtained from text analysis of the fourth text is higher, which improves the accuracy of the first text and thus the accuracy of the generated subject image.

[0097] Furthermore, the electronic device can obtain a first text describing the target network resource based on a fourth text that can effectively describe the content of the target network resource.

[0098] In one implementation, the electronic device can directly concatenate the fourth texts and use the concatenated result as the first text. Since the fourth texts can effectively describe the content of the target network resource, the concatenated result of the fourth texts can also effectively describe the content of the target network resource, i.e., it can describe the target network resource.

[0099] In another implementation, step S1013 may include the following steps:

[0100] Step 1: Use a word segmenter to segment the fourth text to obtain the words to be analyzed.

[0101] Step 2: Input the words to be analyzed into the text analysis model to obtain a brief description of the target network resource output by the text analysis model, which serves as the first text.

[0102] Electronic devices can load pre-defined text analysis models and tokenizers. For example, when an electronic device calls the Tokenizer module from the Transformers model as a tokenizer, it can also call the Model For Causal LM module from the Transformers model as a text analysis model.

[0103] Then, the electronic device determines the corresponding encoding method and the corresponding number of characters based on the data format of the input data of the loaded text analysis model.

[0104] Then, the electronic device uses a word segmenter to segment the fourth text according to a determined number of characters, resulting in multiple words to be analyzed.

[0105] For each word to be analyzed, the electronic device encodes the word according to the determined encoding method, thus obtaining the input data for the text analysis model.

[0106] For example, based on the data format of the input data of text analysis model 1, the electronic device determines that the encoding method corresponding to text analysis model 1 is: Unicode encoding; the number of characters corresponding to text analysis model 1 is: 2.

[0107] Based on the data format of the input data of text analysis model 2, the electronic device determines that the encoding method corresponding to text analysis model 2 is: GBK (Chinese Internal Code Specification) encoding; the number of characters corresponding to text analysis model 2 is: 3.

[0108] If the preset text analysis model loaded by the electronic device is Text Analysis Model 1, then the electronic device determines that the corresponding encoding method is Unicode encoding, and the corresponding number of characters is 2. Therefore, the electronic device uses a word segmenter to segment the fourth text, obtaining two-character words to be analyzed; then, for each word to be analyzed, it performs Unicode encoding on the word to be analyzed, and uses the encoding result as the input data for the text analysis model.

[0109] Then, the electronic device inputs the encoded text analysis model's input data into the text analysis model to obtain the model's output data. Following a preset correspondence between the output data format and characters of the text analysis model, the electronic device decodes the output data to obtain a brief description of the target network resource, i.e., the first text.

[0110] Based on the above processing, the electronic device uses a text analysis model and a word segmenter to perform text analysis on the fourth text. This eliminates the need for technicians to manually determine the content summary of the target network resource based on the script fragments of the target network resource, thereby improving the efficiency and accuracy of obtaining the first text.

[0111] In some embodiments, to obtain a more concise first text, the maximum number of characters in the output data of the text analysis model can be set. For example, if the maximum number of characters can be set to 100, the electronic device will decode the output data of the text analysis model, and the resulting first text will have no more than 100 characters. Subsequently, obtaining the second text based on the first text with fewer characters is more efficient.

[0112] In one implementation, the electronic device can obtain the first text based on the script of the target network resource when it needs to display the theme image of the target network resource to the user.

[0113] In another implementation, the electronic device can predetermine the first text corresponding to each network resource. Subsequently, when the electronic device needs to display the theme image of the target network resource to the user, it can directly obtain the predetermined first text corresponding to the target network resource. Then, it can directly generate the theme image of the target network resource based on the obtained first text, which can improve the efficiency of image generation.

[0114] In some embodiments, the target network resource is a novel. The electronic device can analyze the novel's content to obtain a synopsis, which serves as the first text. The method by which the electronic device analyzes the novel's content is similar to the method used to analyze the script of film and television resources in the aforementioned embodiments, and can be found in the relevant descriptions of the foregoing embodiments.

[0115] Interest tags associated with the target account describe the historical online resources that match the target account. A match between historical online resources and the target account indicates that the user of the target account is more likely to be interested in those historical online resources.

[0116] In some embodiments, Figure 1 Based on this, see Figure 4 Step S101 may include the following steps:

[0117] S1014: Obtain the first text corresponding to the target network resource, and obtain the interaction information corresponding to the target account.

[0118] The interactive information is obtained based on the interactive instructions for historical network resources; the historical network resources are the network resources in the target account's historical access records.

[0119] S1015: For each feature tag of historical network resources, determine the matching degree between the target account and the feature tag based on the interaction information.

[0120] Among them, the feature tags of historical network resources are used to describe historical network resources.

[0121] S1016: Among the feature tags of historical network resources, the feature tags with a matching degree greater than a preset threshold with the target account are identified as the interest tags corresponding to the target account.

[0122] Since users of a target account may interact with historical web resources when browsing them, targeting resources of interest, the interaction information corresponding to the target account can indicate the historical web resources that the target account's user is interested in; that is, it can indicate historical web resources that match the target account. A historical web resource can have multiple feature tags, which are used to describe the historical web resource. For example, feature tags could be "dark blue-toned paintings," "rebellious style characters," or "futuristic urban scenes."

[0123] Furthermore, for each historical online resource, which may have multiple feature tags, different historical online resources may share the same feature tag. After acquiring the user's interaction information regarding browsed historical online resources, the electronic device can, for each feature tag, count the number of times the user interacted with historical online resources bearing that feature tag, and use the count as the total number of interactions for that target account with that feature tag. For example, the electronic device can calculate the sum of the number of interactions the user performed on historical online resources with that feature tag, or calculate the average number of interactions the user performed on historical online resources with that feature tag.

[0124] Furthermore, the electronic device can determine the matching degree between the target account and each feature tag based on the number of operations performed by the target account on each feature tag. For example, for each feature tag, the electronic device can calculate the ratio of the number of operations performed by the target account on that feature tag to the total number of operations performed by the target account on all feature tags, and use this ratio as the matching degree between the target account and that feature tag.

[0125] The match between a target account and a feature tag indicates the probability that users of that target account are interested in that feature tag. A higher match rate indicates a higher probability that users of that target account are interested in that feature tag; a lower match rate indicates a lower probability that users of that target account are interested in that feature tag.

[0126] Furthermore, in order to improve the accuracy of the interest tags corresponding to the target account, the electronic device can identify feature tags that match the target account more than a preset threshold as the interest tags corresponding to the target account.

[0127] Based on the above processing, electronic devices can obtain interest tags corresponding to the target account according to the user's interactive commands regarding historical online resources. These interest tags describe the content that the user is interested in. Subsequent thematic images generated based on the initial text and interest tags can effectively represent the content of the historical online resources that the user is interested in. Thematic images are then recommended to the user, allowing them to browse images that are of interest to them and represent the target online resources, thus improving the effectiveness of image display on the client side.

[0128] In one implementation of step S102, step S102 may include the following steps: concatenating the first text and interest tags to obtain the second text corresponding to the target account and the target network resource.

[0129] After obtaining the first text describing the target network resource and the interest tags describing historical network resources that match the target account, the electronic device can directly concatenate the first text and the interest tags to improve the efficiency of obtaining the second text, and use the concatenated result as the second text corresponding to the target account for the target network resource.

[0130] For example, if the first text is "Invaders from Planet A invade Earth, and Earth's defenders bravely fight back against the invaders"; and the interest tag is "rebellious character", then concatenating the first text and the interest tag will result in the second text being "Invaders from Planet A invade Earth, and Earth's defenders bravely fight back against the invaders; rebellious character".

[0131] In another implementation, step S102 may include the following steps: inputting the first text and interest tags into the text fusion model to obtain the second text corresponding to the target account and the target network resource output by the text fusion model.

[0132] After obtaining the first text and interest tags, electronic devices can use a text fusion model to fuse the first text and interest tags to improve the accuracy of the acquired second text and, consequently, the accuracy of the subsequent topic image generated based on the second text. This results in the second text corresponding to the target account and the target network resource. For example, the text fusion model could be a GPT model, a CNN (Convolutional Neural Network) model, or similar models.

[0133] For example, in the aforementioned case, if the first text and interest tags are input into the text fusion model, the second text output by the text fusion model can be "Invaders on planet A invade Earth, and rebellious Earth defenders bravely fight back against the invaders".

[0134] Based on the above processing, when the first text and interest tags are directly concatenated to obtain the second text, the efficiency of obtaining the second text can be improved, thereby improving the efficiency of generating thematic images; when the first text and interest tags are fused through a text fusion model to obtain the second text, the accuracy of the second text can be improved, thereby improving the accuracy of the generated thematic images.

[0135] In some embodiments, when obtaining the interest tags corresponding to the target account, the electronic device can also obtain the interest vector corresponding to the target account. Each element in the interest vector corresponds one-to-one with an interest tag. Each element represents the probability that a user is interested in the interest tag corresponding to that element. For example, the element corresponding to an interest tag can be the matching degree between the target account and that interest tag. For instance, the interest tag "dark blue-toned painting" corresponds to the first element in the interest vector; the interest tag "rebellious style character" corresponds to the second element in the interest vector; the interest tag "futuristic urban scene" corresponds to the third element in the interest vector; and the interest vector can be [0.9, 0.2, 0.1].

[0136] Furthermore, interest tags and the first text can be merged based on the numerical values ​​of the elements corresponding to the interest tags. For example, when users browse thematic images from various online resources, they generally first view the image region in the middle of the thematic image; and when using an image generation model to generate thematic images based on the second text, the earlier the text in the second text is positioned, the closer its corresponding image region in the generated thematic image will be to the middle of the thematic image. Therefore, interest tags with larger corresponding element values ​​can be merged with the earlier-positioned text in the first text.

[0137] In step S103, in order to obtain an image that better represents the content of the target network resource and the content of the historical network resources that the user is interested in, the electronic device generates guiding text based on the second text and inputs the guiding text into a preset image generation model. The image generation model outputs an image that conforms to the description of the second text, thus obtaining an image that can better represent the content of the target network resource and the content of the historical network resources that the user is interested in, which serves as the theme image of the target network resource that matches the target account.

[0138] For example, an electronic device can use prompt vectors to generate guiding text, and then output the guiding text to an image generation model to input information including content describing the target web resource and content of historical web resources that the user is interested in.

[0139] For example, the prompt vector could be "Based on your interests (content to be embedded), please generate thematic images related to these elements." Furthermore, the electronic device can embed the second text as the content to be embedded into the prompt vector, creating a new prompt vector carrying both the content of the target online resource and the content of historical online resources of interest to the user, thus obtaining the guiding text.

[0140] Then, the electronic device inputs the guiding text into the image generation model, which generates a corresponding image based on the content carried in the guiding text, thus obtaining a subject image of the target network resource that matches the target account.

[0141] Image generation models can include Stable Diffusion, Midjourney, and Dali-Wall E2.

[0142] In some embodiments, Figure 1 Based on this, see Figure 5 Step S103 may include the following steps:

[0143] S1031: Generate guiding text based on the second text, and input the guiding text and the preset target network resource theme image into the preset image generation model to obtain the image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0144] The preset theme image for the target online resource is a fixed image. The preset theme image can be a theme image provided by the producer of the target online resource. For example, if the target online resource is a movie, the preset theme image can be a movie poster provided by the producer; if the target online resource is a novel, the preset theme image can be a novel cover image provided by the producer.

[0145] To improve the accuracy of the generated thematic images, electronic devices can input both guiding text and a preset thematic image into a preset image generation model. The resulting image output by the model will be a thematic image of the target online resource that matches the target account. The generated thematic image is essentially an image obtained by adjusting the preset thematic image based on the guiding text. Therefore, the generated thematic image includes the content of the preset thematic image and can better describe the content of the target online resource and the historical online resources that the user is interested in.

[0146] For example, the target online resource could be a TV series, and the preset theme image could be the cover image of the TV series, displaying the male lead A, female lead B, and the title of the TV series. Then, through an image generation model, based on guiding text, the preset theme image can be adjusted to obtain a theme image that includes the male lead A, female lead B, and the title of the TV series, and better represents the content of the target online resource and the content of historical online resources that the user is interested in.

[0147] Based on the above processing, the image generation model can improve the accuracy of the content of the target network resource represented by the obtained theme image by following the input guiding text and the preset theme image of the target network resource.

[0148] In some embodiments, the electronic device is a server for an entertainment platform, such as a platform that provides online video streaming services, like a television platform or a movie platform. The target network resource is film and television resources.

[0149] Step 1: Obtain the script for the film and television resource.

[0150] In this step, the script of the film and television resource is the third text corresponding to the target network resource in the aforementioned embodiment. For example, the script could be: "[Scene] Inside the spaceship; [Character 1] Captain: We are heading to an unknown galaxy. [Character 2] Mechanic: The ship's engine is having problems! [Scene] Planetary surface; [Character 1] Crew A: The creatures on this planet look very strange. [Character 2] Crew B: We need to explore this planet."

[0151] Step 2: Divide the script according to the scene to obtain multiple script fragments.

[0152] In this step, the script is segmented according to the scene. That is, the second preset keyword in the aforementioned embodiment is "[scene]". The server divides the script according to the second preset keyword to obtain multiple script fragments.

[0153] Step 3: For each script segment, the server uses preset rules to determine whether the script segment is a science fiction plot slice, and extracts each science fiction plot slice.

[0154] In this step, the science fiction plot segment is the fourth text including at least one first preset keyword in the aforementioned embodiments. The preset rule is the same as in the aforementioned embodiments: determining whether the script fragment includes at least one first preset keyword corresponding to the type tag.

[0155] The server determines whether the script excerpt contains the first preset keyword of the genre tag "science fiction". When the genre tag is "science fiction", the first preset keyword can include: "galaxy", "space", "unknown", "alien", "universe", etc.

[0156] Step 4: Encode the sliced ​​text according to the preset text analysis model and word segmenter, process the encoded data using the text analysis model, and then decode the output data of the text analysis model to obtain the episode analysis results.

[0157] In this step, the preset text analysis model can be the Model For Causal LM module in the aforementioned embodiments; the tokenizer can be the Tokenizer module in the aforementioned embodiments. The sliced ​​text is the science fiction plot slice; the episode analysis result is the first text in the aforementioned embodiments.

[0158] The server determines the encoding method and the number of characters based on the data format of the input data to the loaded text analysis model. Then, according to the determined number of characters, a word segmenter is used to segment the science fiction plot segment, resulting in multiple words to be analyzed. For each word to be analyzed, it is encoded according to the determined encoding method, thus obtaining the input data for the text analysis model. This encoded input data is then fed into the text analysis model to obtain its output data. Finally, according to a preset correspondence between the data format and characters of the output data, the output data of the text analysis model is decoded to obtain a synopsis of the film and television resource, i.e., the analysis result of the series.

[0159] Step 5: Use machine learning algorithms to analyze user viewing behavior data and generate user interest tags.

[0160] In this step, the user viewing behavior data refers to the interaction information corresponding to the target account in the aforementioned embodiments. The interaction information is obtained based on interaction instructions for historical network resources; the historical network resources are the network resources in the target account's historical access records. The user interest tags are the interest tags corresponding to the target account in the aforementioned embodiments.

[0161] After acquiring user viewing behavior data, the server determines the matching degree between the target account and each feature tag of historical online resources, based on the user viewing behavior data. The feature tags of historical online resources are used to describe those resources. Furthermore, feature tags among the historical online resources whose matching degree with the target account is greater than a preset threshold are identified as user interest tags.

[0162] Step 6: Embed the generated user interest tags and the episode analysis results obtained from analyzing the sliced ​​text into the prompt vector to construct the guided generation prompt.

[0163] In this step, the prompt is generated using the prompt vector in the previous embodiment. For example, the prompt vector could be "Based on your interests (content to be embedded), please generate thematic images related to these elements."

[0164] The server can directly concatenate user interest tags and drama analysis results, and embed the concatenated result into the prompt vector as the content to be embedded, resulting in a newly generated prompt vector, which serves as the guiding text.

[0165] Alternatively, the server can input user interest tags and episode analysis results into the text fusion model to obtain the second text output by the text fusion model. Then, the second text output by the text fusion model is used as the content to be embedded and embedded into the prompt vector to obtain the newly generated prompt vector, which is the guiding text.

[0166] Step 7: Input the generated prompt into Stable Diffusion. The prompt will pass user interest information and content information of film and television resources to Stable Diffusion to obtain the image output by Stable Diffusion.

[0167] In this step, Stable Diffusion is the image generation model in the aforementioned embodiments; the image output by Stable Diffusion is the subject image of the target network resource that matches the target account in the aforementioned embodiments.

[0168] The server can input guiding text into a preset image generation model and obtain an image output by the image generation model, which can be used as the theme image of the target network resource that matches the target account.

[0169] Alternatively, the server can input the guiding text and the theme image of the preset film and television resources into the preset image generation model, and obtain the image output by the image generation model as the theme image of the target network resources that match the target account.

[0170] Based on the above processing, since the second text is used to describe the target online resource and the historical online resources matched with the target account, the thematic image generated based on the second text can represent both the target and historical online resources. That is, the thematic image can effectively represent the content of the target online resource and the content of the historical online resources that the user is interested in. When thematic images are subsequently recommended to the user, the user can browse thematic images that are of interest to them and can represent the content of the target online resource. This increases the probability that the user will choose the thematic image that interests them, thereby improving the effectiveness of electronic devices in displaying online resources to the user based on thematic images.

[0171] Furthermore, by deeply mining user viewing behavior data and combining it with drama series analysis, highly personalized user interest vectors and drama series tags were constructed. This provides strong support for generating personalized promotions, ensuring that the generated images closely match user preferences and drama series characteristics.

[0172] Based on the same inventive concept as the image generation method described above, embodiments of the present invention also provide an image generation apparatus. See [link to previous document]. Figure 6 , Figure 6 A structural diagram of an image generation apparatus provided in an embodiment of the present invention. The apparatus includes:

[0173] The first acquisition module 601 is used to acquire a first text corresponding to a target network resource and an interest tag corresponding to a target account; wherein, the first text is used to describe the target network resource; and the interest tag is used to describe historical network resources that match the target account.

[0174] The fusion module 602 is used to generate a second text corresponding to the target account for the target network resource based on the first text and the interest tag; wherein the second text is used to describe the target network resource and the historical network resource;

[0175] The second acquisition module 603 is used to generate guiding text based on the second text, and input the guiding text into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0176] Optionally, the target network resource is a film and television resource;

[0177] The first acquisition module 601 is specifically used for:

[0178] Obtain the third text corresponding to the target network resource, and the type tag of the target network resource; wherein, the third text is the script of the target network resource; the third text includes at least one of the following: the scene in the target network resource, the characters in the target network resource, and the lines of each character in the target network resource; based on the first preset keyword of the type tag, extract script fragments including at least one keyword from the third text to obtain a fourth text; perform text analysis on the fourth text to obtain a content summary of the target network resource, which serves as the first text.

[0179] Optionally, the first acquisition module 601 is specifically used for:

[0180] The third text is divided according to the second preset keywords to obtain multiple script fragments;

[0181] For each script fragment, if the script fragment includes at least one first preset keyword corresponding to the type tag, the script fragment is determined to be the fourth text.

[0182] Optionally, the first acquisition module 601 is specifically used for:

[0183] The fourth text is segmented using a word segmenter to obtain words to be analyzed; the words to be analyzed are then input into a text analysis model to obtain a summary of the target network resource output by the text analysis model, which serves as the first text.

[0184] Optionally, the first acquisition module 601 is specifically used for:

[0185] Obtain interaction information corresponding to the target account; wherein the interaction information is obtained based on interaction instructions for the historical network resources; the historical network resources are network resources in the historical access records of the target account; for each feature tag of the historical network resources, determine the matching degree between the target account and the feature tag based on the interaction information; wherein the feature tags of the historical network resources are used to describe the historical network resources; determine the feature tags among the feature tags of the historical network resources whose matching degree with the target account is greater than a preset threshold as the interest tags corresponding to the target account.

[0186] Optionally, the fusion module 602 is specifically used for:

[0187] The first text and the interest tag are concatenated to obtain the second text corresponding to the target account for the target network resource; or, the first text and the interest tag are input into a text fusion model to obtain the second text corresponding to the target account for the target network resource output by the text fusion model.

[0188] Optionally, the second acquisition module 603 is specifically used for:

[0189] Based on the second text, guide text is generated, and the guide text and the preset theme image of the target network resource are input into a preset image generation model to obtain the image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

[0190] Based on the image generation apparatus provided in this embodiment of the invention, since the second text is used to describe the target network resource and the historical network resources matching the target account, the thematic image generated based on the second text can represent both the target network resource and the historical network resources. That is, the thematic image can effectively represent the content of the target network resource and the content of the historical network resources that the user is interested in. When thematic images are subsequently recommended to the user, the user can browse thematic images that are of interest to them and can represent the content of the target network resource. Therefore, the probability of the user selecting a thematic image that interests them is higher, thereby improving the effectiveness of the electronic device in displaying network resources to the user based on thematic images.

[0191] This invention also provides an electronic device, such as... Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.

[0192] Memory 703 is used to store computer programs;

[0193] When the processor 701 executes the program stored in the memory 703, it implements the steps of any of the image generation methods in the above embodiments.

[0194] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0195] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0196] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0197] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0198] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the image generation methods described in the above embodiments.

[0199] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the image generation methods described in the above embodiments.

[0200] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0201] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0202] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0203] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An image generation method, characterized in that, The method includes: Obtain the first text corresponding to the target network resource and obtain the interest tags corresponding to the target account; wherein, the first text is used to describe the target network resource; the interest tags are used to describe historical network resources that match the target account; Based on the first text and the interest tags, a second text corresponding to the target account and the target network resource is generated; wherein, the second text is used to describe the target network resource and the historical network resource; Based on the second text, guide text is generated, and the guide text is input into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource that matches the target account; The target network resources are film and television resources; The step of obtaining the first text corresponding to the target network resource includes: obtaining the third text corresponding to the target network resource, and the type tag of the target network resource; wherein, the third text is the script of the target network resource; based on the first preset keyword of the type tag, extracting script fragments including at least one keyword from the third text to obtain the fourth text; performing text analysis on the fourth text to obtain a summary of the content of the target network resource, which is used as the first text.

2. The method according to claim 1, characterized in that, The third text includes at least one of the following: the scene in the target network resource, the characters in the target network resource, and the lines of each character in the target network resource.

3. The method according to claim 2, characterized in that, The fourth text, based on the first preset keyword of the type tag, extracts script fragments including at least one keyword from the third text, resulting in a fourth text, including: The third text is divided according to the second preset keywords to obtain multiple script fragments; For each script fragment, if the script fragment includes at least one first preset keyword corresponding to the type tag, the script fragment is determined to be the fourth text.

4. The method according to claim 3, characterized in that, The text analysis of the fourth text to obtain a brief description of the target network resource, which serves as the first text, includes: The fourth text is segmented using a word segmenter to obtain the words to be analyzed. The words to be analyzed are input into the text analysis model to obtain a brief description of the target network resource output by the text analysis model, which serves as the first text.

5. The method according to claim 1, characterized in that, The process of obtaining the interest tags corresponding to the target account includes: Obtain interaction information corresponding to the target account; wherein, the interaction information is obtained based on the interaction instructions for the historical network resources; the historical network resources are network resources in the historical access records of the target account; For each feature tag of the historical network resource, the matching degree between the target account and the feature tag is determined based on the interaction information; wherein, the feature tag of the historical network resource is used to describe the historical network resource; Among the feature tags of the historical network resources, the feature tags with a matching degree greater than a preset threshold with the target account are identified as the interest tags corresponding to the target account.

6. The method according to claim 1, characterized in that, The step of generating a second text corresponding to the target account and the target network resource based on the first text and the interest tags includes: The first text and the interest tag are concatenated to obtain the second text corresponding to the target account and the target network resource; or, The first text and the interest tag are input into the text fusion model to obtain the second text corresponding to the target account and the target network resource, which is output by the text fusion model.

7. The method according to claim 1, characterized in that, The step of generating guiding text based on the second text, and inputting the guiding text into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource matching the target account, includes: Based on the second text, guide text is generated, and the guide text and the preset theme image of the target network resource are input into a preset image generation model to obtain the image output by the image generation model, which serves as the theme image of the target network resource that matches the target account.

8. An image generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire first text corresponding to the target network resource and interest tags corresponding to the target account; wherein, the first text is used to describe the target network resource; and the interest tags are used to describe historical network resources that match the target account. The fusion module is used to generate a second text corresponding to the target account for the target network resource based on the first text and the interest tag; wherein the second text is used to describe the target network resource and the historical network resource; The second acquisition module is used to generate guiding text based on the second text, and input the guiding text into a preset image generation model to obtain an image output by the image generation model, which serves as the theme image of the target network resource that matches the target account; The target network resource is a film and television resource; the first acquisition module is specifically used to acquire the third text corresponding to the target network resource and the type tag of the target network resource; wherein, the third text is the script of the target network resource; based on the first preset keyword of the type tag, the script fragment including at least one keyword is extracted from the third text to obtain the fourth text; the fourth text is analyzed to obtain the content summary of the target network resource, which is used as the first text.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for generating cover image, equipment and storage medium

    CN113821677A

  • Picture generation method, live broadcast room image generation method and device

    CN116320524A