Game advertisement putting material making method based on AIGC technology

Through the game advertising material production method based on AIGC technology, personalized creative materials are generated based on game characteristics and delivery needs, solving the problem that materials cannot be intelligently customized in the existing technology, and efficient and diversified creative production is achieved.

CN119991208APending Publication Date: 2025-05-13WUHAN ZHIQITE ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510084185.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing game creatives generated through AIGC cannot intelligently customize personalized creatives based on game content, characteristics and delivery direction, resulting in poor material availability and cannot meet the needs of dynamically changing promotion channels and audiences.

Method used

A method for producing game advertising material based on AIGC technology is provided. By determining text prompt words, example pictures and example audio based on game characteristics and delivery needs, using preset pictures, audio and video generation models to generate corresponding materials, and optimize and integrate them to output the final advertising material.

Benefits of technology

It has achieved rapid batch production of creative materials that meet the characteristics of the game, and intelligently customized personalized creative materials based on different delivery channels and target audience characteristics, improving the production efficiency and creative diversity of game promotion materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991208A_ABST
    Figure CN119991208A_ABST
Patent Text Reader

Abstract

The invention discloses a game advertisement putting material making method based on an AIGC technology, and the method comprises the steps: determining a text prompt word, an example picture and an example audio according to the game characteristics and putting demands; generating a picture material according to the text cue word or the example picture by utilizing a preset picture generation model; utilizing a preset audio generation model to generate an audio material according to the text prompt word or the example audio; generating a video material according to the example picture, the picture material or the character prompt word by using a preset video generation model; and optimizing and integrating the picture material, the audio material and the video material, and outputting a final advertisement putting material according to a putting demand. Static pictures, dynamic posters and video advertisements conforming to game features can be rapidly manufactured in batches, personalized advertisement materials are intelligently customized according to different delivery channels and features of target audiences, and the production efficiency and creative diversity of game promotion materials are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technology, and in particular to a method for producing game advertisement delivery materials based on AIGC technology. Background Art

[0002] Against the backdrop of increasingly fierce competition in the global game market, game developers must adopt innovative and effective marketing strategies to stand out from many similar products, attract more players and increase market share. The quality of advertising materials directly affects the exposure and attractiveness of the game. Traditional game delivery material production usually relies on manual design and production. Designers produce materials one by one according to the characteristics of the game and advertising needs, and make manual adjustments and modifications. This process is time-consuming and labor-intensive, especially when multiple special effects and style changes are required according to different delivery channels and target audiences, which requires huge manpower and time costs.

[0003] With the iterative development of AIGC technology, generative algorithms have become a key innovation in the field of artificial intelligence. This type of algorithm can automatically generate diversified content such as pictures, videos, and audio, and has significant application value and broad development prospects in the fields of art, design, entertainment, virtual reality, etc. However, the existing game advertising materials generated by AIGC usually rely on preset templates or fixed material styles. It is impossible to intelligently customize personalized advertising materials according to the game content, characteristics, and delivery direction, and it is impossible to make different types of materials echo each other. When facing different market demands, it is slow to respond and has poor adaptability.

[0004] Therefore, it is necessary to propose a method for producing game advertising materials based on AIGC technology, which can quickly batch produce different types of advertising materials that meet the characteristics of the game. It can also intelligently customize personalized advertising materials according to the characteristics of different delivery channels and target audiences, so as to improve the production efficiency and creative diversity of game promotion materials. Summary of the invention

[0005] In view of this, the present invention provides a method for producing game advertising materials based on AIGC technology, so as to solve the technical problem that the existing game advertising material generation technology is unable to intelligently customize personalized advertising materials according to game content, characteristics and delivery direction, resulting in poor material availability and inability to meet dynamically changing promotion channels and target audiences.

[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a method for producing game advertisement delivery materials based on AIGC technology, comprising:

[0008] Determine text prompts, sample images, and sample audio based on game features and delivery requirements;

[0009] Generate image materials based on text prompts or sample images using a preset image generation model;

[0010] Generate audio material based on text prompts or sample audio using a preset audio generation model;

[0011] Generate video material according to the sample picture, picture material or text prompt word by using a preset video generation model;

[0012] The image materials, audio materials and video materials are optimized and integrated, and the final advertising delivery materials are output according to delivery requirements.

[0013] Further, the preset picture generation model includes a first text encoder, a picture encoder, a picture decoder and a picture generator; the output ends of the first text encoder and the picture encoder are connected to the input end of the picture generator, and the output end of the picture generator is connected to the input end of the picture decoder;

[0014] The first text encoder is used to segment and convert the text prompt words to obtain a text feature vector;

[0015] The picture encoder and picture decoder are constructed based on the VAE model. The picture encoder is used to compress the sample picture into the latent space to obtain the picture feature vector; the picture decoder is used to map the picture feature vector in the latent space back to the image space;

[0016] The image generator is built based on the DiT model, and is used to embed the input text feature vector into the latent space to generate the image feature vector, perform discrete operations on the image feature vector corresponding to the example image, and couple the text feature vector and the image feature vector to obtain the denoised latent space image features.

[0017] Furthermore, the first text encoder includes a word segmenter, and a CLIP model obtained by contrastive training or a T5 model trained by self-supervision;

[0018] The word segmenter is used to process the text prompt words into ids sequences;

[0019] The CLIP model or T5 model is used to convert the ids sequence into a text feature vector.

[0020] Further, the audio generation model includes a second text encoder, an audio generator and a vocoder connected in sequence;

[0021] The second text encoder is used to obtain semantic features of the text prompt word;

[0022] The audio generator is constructed based on Transformer Block, including an audio encoder and an audio decoder, wherein the audio encoder and the audio decoder are connected via a mapping layer, and the output end of the audio decoder is connected to the vocoder; the audio generator is used to generate a mel spectrogram sequence according to the audio file corresponding to the sample audio or text prompt word;

[0023] The vocoder is used to convert the mel spectrogram sequence into an audio waveform file.

[0024] Further, the video generation model includes a third text encoder, a video encoder, a video decoder and a video generator; the output ends of the third text encoder and the video encoder are connected to the input end of the video generator, and the output end of the video generator is connected to the input end of the video decoder;

[0025] The video encoder and video decoder take video clips as input and adopt a three-dimensional convolutional model. During model training, multiple frame clips are predicted simultaneously to maintain the consistency of the video before and after.

[0026] Furthermore, the image material, audio material and video material are optimized and integrated, including:

[0027] Remove background, restore image quality, and add and remove watermarks from image materials;

[0028] Convert audio formats, separate audio tracks and optimize sound quality;

[0029] Perform super-resolution enlargement, resize ratio adjustment, frame interpolation and format conversion on video materials.

[0030] Furthermore, the final advertising delivery materials are output according to the delivery requirements, including:

[0031] Deliver image materials, audio materials, and video materials separately based on delivery requirements;

[0032] The picture material is used as a watermark of the video material and then is released;

[0033] The audio material and the video material are matched and combined to generate a video material with both sound quality and picture quality and then released.

[0034] In a second aspect, the present invention further provides a device for producing game advertisement delivery materials based on AIGC technology, comprising:

[0035] An information extraction module is used to determine text prompt words, sample images, and sample audio according to game characteristics and delivery requirements;

[0036] A picture material generation module is used to generate picture materials according to text prompt words or sample pictures using a preset picture generation model;

[0037] An audio material generation module, used to generate audio materials according to text prompt words or sample audio using a preset audio generation model;

[0038] A video material generation module, used to generate video material according to the sample pictures, picture materials or text prompt words by using a preset video generation model;

[0039] The material processing module is used to optimize and integrate the image material, audio material and video material, and output the final advertising delivery material according to delivery requirements.

[0040] In a third aspect, the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for producing game advertising materials based on AIGC technology as described in any of the above technical solutions is implemented.

[0041] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, a method for producing game advertising materials based on AIGC technology as described in any of the above technical solutions is implemented.

[0042] Compared with the prior art, the present invention provides the following advantages:

[0043] (1) By utilizing preset information such as text prompts, sample images, and sample audio, the system can automatically generate images, audio, and video materials, thereby improving the efficiency of the production process and saving time and energy in manual creation and modification.

[0044] (2) Based on the characteristics of the game and the demand for delivery, the materials produced can accurately reflect the characteristics of the game and the needs of the target users. According to different game types, advertising needs and market trends, the rules for generating text prompts, pictures and audio are adjusted. Enterprises can customize the generation according to specific needs and quickly adapt to changing market demands.

[0045] (3) By integrating image, audio and video generation models, cross-modal content production can be achieved. By integrating image, audio and video materials, cross-modal content production can be achieved, ensuring that the final advertising materials are more vivid and expressive, and can be delivered on multiple platforms, thereby enhancing the attractiveness and expressiveness of the advertisements.

[0046] In summary, the present invention not only improves the efficiency and accuracy of material production, but also enhances the interactivity and user experience of advertisements, while reducing costs, allowing advertisers to adjust strategies more flexibly and respond quickly to market demands. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A flowchart of a method for producing game advertisement delivery materials based on the AIGC technology provided by the present invention;

[0048] Figure 2 A schematic diagram of the sample generation process provided by the present invention;

[0049] Figure 3 A schematic diagram of a scenario application provided by the present invention;

[0050] Figure 4 A schematic diagram of the structure of a device for producing game advertisement delivery materials based on AIGC technology provided by the present invention;

[0051] Figure 5 The present invention provides a schematic structural diagram of an electronic device. DETAILED DESCRIPTION

[0052] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.

[0053] See also Figure 1 This embodiment provides a method for producing game advertisement delivery materials based on AIGC technology, including:

[0054] Step S101: Determine text prompt words, sample pictures and sample audio according to game characteristics and delivery requirements;

[0055] Step S102: Generate picture materials according to text prompt words or sample pictures using a preset picture generation model;

[0056] Step S103: Generate audio material according to the text prompt words or sample audio using a preset audio generation model;

[0057] Step S104: Generate video material based on the sample image, image material or text prompt word using a preset video generation model

[0058] Step S105: Optimize and integrate the image material, audio material and video material, and output the final advertisement delivery material according to delivery requirements.

[0059] The method of this embodiment can significantly reduce the labor cost and time cost of producing game promotion materials. By automatically generating advertising materials through AIGC technology, not only can static images, dynamic posters and video ads that meet the characteristics of the game be quickly and batch-produced, but also personalized advertising materials can be intelligently customized according to the characteristics of different delivery channels and target audiences, greatly improving the production efficiency and creative diversity of game promotion materials.

[0060] As a specific embodiment, in step S101, the game delivery material production personnel first deeply analyze the content requirements, target users, delivery platform, delivery purpose, etc. of the game to be delivered. For example, according to the type of game, theme and storyline and gameplay characteristics, determine the visual style, atmosphere and specific BGM or sound effects of the game. After the analysis of the game to be delivered is completed, text prompts, sample pictures and sample audio are obtained. Here, it is necessary to ensure the consistency between the text prompts, sample pictures and sample audio so that they can maintain unity in style, emotion and core gameplay. For example, if the game is a doomsday survival game, the sample pictures and audio should also highlight this tense and desolate atmosphere. Secondly, adjust the content according to the characteristics of the target audience. For example, if the target is teenage users, the visual style, sound effects and text prompts in the advertisement can be more dynamic and fashionable; while for adult players, more mature and deep content may be required. Finally, since different platforms (social media, mobile advertising, video websites, etc.) have different requirements for advertising, the prompts and sample content should be optimized according to the characteristics of the platform. Through this step, the input text prompts, sample audio and sample images can be strongly related to the game to be launched, meeting the requirements of different game advertising delivery areas, and the generated materials can be set to various special effects and styles required.

[0061] As a preferred embodiment, the base layers of the preset picture generation model, audio generation model and video generation model are all generation models, which are initialized with a random noise from a prior distribution (such as a standard normal distribution), and then the noise is predicted by a deep learning model for iterative denoising to achieve the generation of new samples. Figure 2 As shown, Figure 2 The basic generation process of the generative model is shown. A random noise randomly sampled from a prior distribution, such as a standard normal distribution N(0,I), is used as the initial sample (representing a latent space vector of an initial image / audio / video clip). Denoising sampling is a cyclic iterative process, assuming that the number of iterations is N; in each step, the trained deep learning model is used to predict the noise of the current step based on text features, images or audio features, and the noise is used to update the new sample of the current step, and then iterative denoising is continued on the basis of the new sample. Finally, after performing N steps of iterative denoising, the final new sample image / audio / video is obtained after passing through the image decoder.

[0062] As a preferred embodiment, in step S102, the preset picture generation model includes a first text encoder, a picture encoder, a picture decoder and a picture generator; the output ends of the first text encoder and the picture encoder are connected to the input end of the picture generator, and the output end of the picture generator is connected to the input end of the picture decoder;

[0063] The first text encoder is used to segment and convert the text prompt words to obtain a text feature vector;

[0064] The picture encoder and picture decoder are constructed based on the VAE model. The picture encoder is used to compress the sample picture into the latent space to obtain the picture feature vector; the picture decoder is used to map the picture feature vector in the latent space back to the image space;

[0065] The image generator is built based on the DiT model, and is used to embed the input text feature vector into the latent space to generate the image feature vector, perform discrete operations on the image feature vector corresponding to the example image, and couple the text feature vector and the image feature vector to obtain the denoised latent space image features.

[0066] As a preferred embodiment, the first text encoder includes a word segmenter, and a CLIP model obtained by contrast training or a T5 model trained by self-supervision;

[0067] The word segmenter is used to process the text prompt words into ids sequences;

[0068] The CLIP model or T5 model is used to convert the ids sequence into a text feature vector.

[0069] As a specific embodiment, the input text prompt word is encoded by the first text encoder to generate a text feature vector of fixed dimension. This feature vector contains the semantic information of the input text and can capture the intention or content of the text description. At the same time, the image encoder is responsible for mapping the input example image to the latent space. The image encoder and decoder are composed of a VAE model of the Unet structure obtained by self-training. The image encoder compresses the input image into a low-dimensional latent vector. By dividing the image into multiple small blocks, it helps to capture information from different parts of the image. In this way, even if the image is very complex, the model can process and understand its local content.

[0070] The text feature vector and the image patch vector are input into the image generator. The image generator is a DiT model constructed by Transformer Block. During training, the image latent vector will be patchified into a series of discrete sequences, which will then be coupled with the input text or image feature vector and trained as the input of the DiT module through Flow matching. This allows the DiT module to acquire denoising capabilities. During generation, it can iterate step by step from pure noise to obtain the denoised latent vector image features, which are then mapped back to the pixel space through the image decoder to obtain the final generation result.

[0071] It should be noted that when only text prompts are input, the text feature vector contains semantic information in the text, such as words describing a scene, object or style. In this case, the model does not rely on the actual patch vector of the image. The image generator maps these text features into the latent space of the image by embedding, that is, embedding the text features into a latent space related to the image.

[0072] As a preferred embodiment, in step S103, the audio generation model includes a second text encoder, an audio generator and a vocoder connected in sequence;

[0073] The second text encoder is used to obtain semantic features of the text prompt word;

[0074] The audio generator is constructed based on Transformer Block, including an audio encoder and an audio decoder, wherein the audio encoder and the audio decoder are connected via a mapping layer, and the output end of the audio decoder is connected to the vocoder; the audio generator is used to generate a mel spectrogram sequence according to the audio file corresponding to the sample audio or text prompt word;

[0075] The vocoder is used to convert the mel spectrogram sequence into an audio waveform file.

[0076] As a specific embodiment, the structure of the second text encoder in the audio generation model is consistent with that of the first text encoder in the picture generation model. The main architecture of the audio generator is composed of Transformer Block, which is trained with vector quantization technology and is trained through supervised ASR tasks.

[0077] The main structure of the audio generator has 12 layers, which is an LLM built based on Transformer Block. The first six layers of Transformer Block are called encoders, followed by a Vector Quantization layer, and then six layers of Transformer Block to form a decoder. The trained encoder and Vector Quantization layer can convert the mel spectrogram sampled from the audio into a sequence.

[0078] The latent space vector of the mel spectrogram sequence (mel spectrogram is a kind of audio spectrum representation that can effectively capture the frequency and time characteristics of audio) is predicted with the text feature sequence or audio sequence as input; in addition, there will be a Unet structure decoder trained by Flowmatching method with the latent space vector of the mel spectrogram sequence predicted by LLM and the guide text feature vector as input, so that the decoder has the ability to denoise. When generating, it can iteratively generate the denoised mel spectrogram sequence step by step from pure noise, and then convert it into real audio by the vocoder.

[0079] The audio generation model can be used to generate pure music accompaniment, songs with accompaniment and specified lyrics, and dubbing audio with specified timbre and text content. After targeted processing of the generated audio materials, they can be directly released and can also be used in conjunction with video materials in the future.

[0080] As a preferred embodiment, in step S104, the video generation model includes a third text encoder, a video encoder, a video decoder and a video generator; the output ends of the third text encoder and the video encoder are connected to the input end of the video generator, and the output end of the video generator is connected to the input end of the video decoder;

[0081] The video encoder and video decoder take video clips as input and adopt a three-dimensional convolutional model. During model training, multiple frame clips are predicted simultaneously to maintain the consistency of the video before and after.

[0082] It can be seen that the architecture of the video generation model is consistent with that of the image generation model. The main difference is that the encoders and decoders of videos and images are different. The video codec takes video clips as input, and uses a three-dimensional convolution model instead of a two-dimensional convolution module internally. During training or generation, multiple frames are predicted simultaneously to maintain object consistency.

[0083] It should be noted that in addition to generating video materials, the video generation model can also output animated image materials with specified content and style.

[0084] As a preferred embodiment, in step S105, the picture material, audio material and video material are optimized and integrated, including:

[0085] Remove background, restore image quality, and add and remove watermarks from image materials;

[0086] Convert audio formats, separate audio tracks and optimize sound quality;

[0087] Perform super-resolution enlargement, resize ratio adjustment, frame interpolation and format conversion on video materials.

[0088] As a specific embodiment, the generated picture materials, audio materials and video materials may have some defects or need to be used in combination. At this time, the materials need to be optimized and integrated. It mainly includes: background removal, super-resolution enlargement (the deep learning model repairs the picture pixels while enlarging the blurry size to improve the picture quality and resolution), volume compression, watermark addition, watermark removal and other operations for picture materials; format conversion, track separation (audio such as BGM in conventional video files is composed of multiple different tracks, such as a music audio with BGM, can have various different tracks such as vocals, piano, bass, drums, etc., which can be separated and selected for use), sound quality optimization and other operations for audio materials; background removal, super-resolution enlargement, watermark addition, watermark removal, size ratio adjustment, frame filling, format conversion and other operations can be performed on video materials.

[0089] As a preferred embodiment, outputting the final advertisement delivery material according to delivery requirements includes:

[0090] Deliver image materials, audio materials, and video materials separately based on delivery requirements;

[0091] The picture material is used as a watermark of the video material and then is released;

[0092] The audio material and the video material are matched and combined to generate a video material with both sound quality and picture quality and then released.

[0093] Specifically, after processing various materials, they can be directly put into use separately or in combination. For example, the generated picture materials can be used as watermarks for pictures and videos, and the generated audio materials and video materials can be combined into videos with sound and picture quality. In addition, with text, dubbing audio and character video as inputs to the video material generation module, the characters in the video can be lip-synced and used in conjunction with the dubbing audio.

[0094] like Figure 3 As shown, Figure 3 The complete application of the game advertising material production method based on AIGC technology proposed in this application is demonstrated.

[0095] This embodiment also provides a game advertisement delivery material production system based on AIGC technology, such as Figure 4 As shown, the game advertisement delivery material production system 400 based on AIGC technology includes:

[0096] Information extraction module 401, used to determine text prompt words, sample pictures and sample audio according to game characteristics and delivery requirements;

[0097] Picture material generation module 402, used to generate picture material according to text prompt words or sample pictures using a preset picture generation model;

[0098] The audio material generation module 403 is used to generate audio material according to the text prompt words or sample audio using a preset audio generation model;

[0099] The video material generation module 404 is used to generate video material according to the example picture, picture material or text prompt word by using a preset video generation model;

[0100] The material processing module 405 is used to optimize and integrate the image material, audio material and video material, and output the final advertisement delivery material according to delivery requirements.

[0101] like Figure 5 As shown, the above-mentioned method for producing game advertisement delivery materials based on AIGC technology, the present invention also provides an electronic device 500, which can be a computing device such as a mobile terminal, a desktop computer, a notebook, a PDA, and a server. The electronic device includes a processor 501, a memory 502, and a display 503.

[0102] In some embodiments, the memory 502 may be an internal storage unit of a computer device, such as a hard disk or memory of a computer device. In other embodiments, the memory 502 may also be an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the memory 502 may also include both an internal storage unit of the computer device and an external storage device. The memory 502 is used to store application software and various types of data installed on the computer device, such as program codes for installing the computer device. The memory 502 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a game advertising material production method program 504 based on AIGC technology is stored on the memory 502, and the game advertising material production method program 504 based on AIGC technology can be executed by the processor 501, thereby realizing a game advertising material production method based on AIGC technology in each embodiment of the present invention.

[0103] In some embodiments, the processor 501 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 502, such as executing a method program for producing game advertising materials based on AIGC technology.

[0104] In some embodiments, the display 503 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 503 is used to display information on the computer device and to display a visual user interface. The components 501-503 of the computer device communicate with each other through a system bus.

[0105] This embodiment also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method for producing game advertising materials based on AIGC technology as described in any of the above technical solutions is implemented.

[0106] The computer-readable storage medium and computing device provided according to the above embodiments of the present invention can be implemented with reference to the content specifically described in the method for producing game advertising delivery materials based on AIGC technology as described above according to the present invention, and have similar beneficial effects as the method for producing game advertising delivery materials based on AIGC technology as described above, which will not be repeated here.

[0107] The game advertising delivery material production method based on AIGC technology provided by the present invention generates game delivery materials based on isomorphic AIGC technology, can quickly batch produce static pictures, dynamic posters and video advertisements that meet the characteristics of the game, and intelligently customize personalized advertising materials according to the characteristics of different delivery channels and target audiences, thereby greatly improving the production efficiency and creative diversity of game promotion materials.

[0108] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for producing game advertisement delivery materials based on AIGC technology, characterized in that: include: Determine text prompts, sample images, and sample audio based on game features and delivery requirements; Generate image materials based on text prompts or sample images using a preset image generation model; Generate audio material based on text prompts or sample audio using a preset audio generation model; Generate video material according to the sample picture, picture material or text prompt word by using a preset video generation model; The image materials, audio materials and video materials are optimized and integrated, and the final advertising delivery materials are output according to delivery requirements.

2. The method for producing game advertisement delivery materials based on AIGC technology according to claim 1, characterized in that: The preset picture generation model includes a first text encoder, a picture encoder, a picture decoder and a picture generator; the output ends of the first text encoder and the picture encoder are connected to the input end of the picture generator, and the output end of the picture generator is connected to the input end of the picture decoder; The first text encoder is used to segment and convert the text prompt words to obtain a text feature vector; The picture encoder and picture decoder are constructed based on the VAE model. The picture encoder is used to compress the example picture into the latent space to obtain the picture feature vector; the picture decoder is used to map the picture feature vector in the latent space back to the image space; The image generator is built based on the DiT model, and is used to embed the input text feature vector into the latent space to generate the image feature vector, perform discrete operations on the image feature vector corresponding to the example image, and couple the text feature vector and the image feature vector to obtain the denoised latent space image features.

3. The method for producing game advertisement delivery materials based on AIGC technology according to claim 2, characterized in that: The first text encoder includes a word segmenter, and a CLIP model obtained by contrastive training or a T5 model trained by self-supervision; The word segmenter is used to process the text prompt words into ids sequences; The CLIP model or T5 model is used to convert the ids sequence into a text feature vector.

4. The method for producing game advertisement delivery materials based on AIGC technology according to claim 1, characterized in that: The audio generation model includes a second text encoder, an audio generator and a vocoder connected in sequence; The second text encoder is used to obtain semantic features of the text prompt word; The audio generator is constructed based on Transformer Block, including an audio encoder and an audio decoder, wherein the audio encoder and the audio decoder are connected via a mapping layer, and the output end of the audio decoder is connected to the vocoder; the audio generator is used to generate a mel spectrogram sequence according to the audio file corresponding to the sample audio or text prompt word; The vocoder is used to convert the mel spectrogram sequence into an audio waveform file.

5. The method for producing game advertisement delivery materials based on AIGC technology according to claim 1, characterized in that: The video generation model includes a third text encoder, a video encoder, a video decoder and a video generator; the output ends of the third text encoder and the video encoder are connected to the input end of the video generator, and the output end of the video generator is connected to the input end of the video decoder; The video encoder and video decoder take video clips as input and adopt a three-dimensional convolutional model. During model training, multiple frame clips are predicted simultaneously to maintain the consistency of the video before and after.

6. The method for producing game advertisement delivery materials based on AIGC technology according to claim 1, characterized in that: Optimizing and integrating the image materials, audio materials and video materials, including: Remove background, restore image quality, and add and remove watermarks from image materials; Convert audio formats, separate audio tracks and optimize sound quality; Perform super-resolution enlargement, resize ratio adjustment, frame interpolation and format conversion on video materials.

7. The method for producing game advertisement delivery materials based on AIGC technology according to claim 1, characterized in that: Output the final advertising materials according to the delivery requirements, including: Deliver image materials, audio materials, and video materials separately based on delivery requirements; The picture material is used as a watermark of the video material and then is released; The audio material and the video material are matched and combined to generate a video material with both sound quality and picture quality and then released.

8. A device for producing game advertisement delivery materials based on AIGC technology, characterized in that: include: An information extraction module is used to determine text prompt words, sample images, and sample audio according to game characteristics and delivery requirements; A picture material generation module is used to generate picture materials according to text prompt words or sample pictures using a preset picture generation model; An audio material generation module, used to generate audio materials according to text prompt words or sample audio using a preset audio generation model; A video material generation module, used to generate video material according to the sample pictures, picture materials or text prompt words by using a preset video generation model; The material processing module is used to optimize and integrate the image material, audio material and video material, and output the final advertising delivery material according to delivery requirements.

9. An electronic device, characterized in that: It includes a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method for producing game advertising materials based on AIGC technology as described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for producing game advertisement delivery materials based on the AIGC technology as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Picture processing method and device, electronic equipment and storage medium

    CN121502021A