Media file generation method and device, equipment and medium

By generating multiple images and achieving cross-screen coverage of the main subject in the media file, the problem of insufficient content creation capabilities of multimedia editing software is solved. The generated media file is rich in information and the main subject is prominent, meeting the diverse needs of users.

CN121865062APending Publication Date: 2026-04-14BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing multimedia editing software struggles to meet users' diverse media file creation needs, has limited content creation capabilities, and is unable to generate media files with multi-screen splicing effects and cross-screen presentation of the main subject.

Method used

By acquiring a first image and a first text, multiple second images are generated, and a media file is generated based on these images and text. The media file presents a spliced ​​image of multiple images, with the area of ​​the same subject object covering multiple images, thus realizing multi-screen splicing and cross-screen presentation of the subject object.

Benefits of technology

It enhances the content creation capabilities of multimedia editing software, generates media files with rich information, highlights the main subject across screens, strengthens the correlation between different screens, and meets the diverse creative needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865062A_ABST
    Figure CN121865062A_ABST
Patent Text Reader

Abstract

The invention relates to a media file generation method and device, equipment and a medium, and the method comprises the steps: obtaining a first image and a first text, and generating a plurality of second images based on the first image and the first text; the first text is a picture description text of the media file to be generated; generating a media file based on at least one first main body object in the plurality of second images and the plurality of second images; the media file is an image or a video, and the media file presents a spliced picture obtained by splicing at least part of pictures corresponding to a plurality of second images; in the spliced picture presented by the media file, the object area of the same first main body object covers partial pictures corresponding to the at least two second images respectively. According to the method, the media file with a multi-picture splicing effect and a main body object cross-picture presentation effect can be obtained, the content creation capability of multimedia editing software can be well improved, and diversified media file creation requirements of users can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device and medium for generating media files. Background Technology

[0002] Media files such as images and videos, with their visual appeal and intuitive expressive power, have become a widely adopted core medium for information transmission in today's society. With the widespread use of multimedia editing software, more and more users are using it to create multimedia files. The inventors' research revealed that existing multimedia editing software offers limited content creation capabilities, mostly only optimizing user-provided images or adding special effects, failing to meet the diverse media file creation needs of users. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this application provides a method, apparatus, device and medium for generating media files.

[0004] In a first aspect, this application provides a method for generating a media file, the method comprising: acquiring a first image and a first text, generating a plurality of second images based on the first image and the first text; wherein the first text is a scene description text of a media file to be generated; generating a media file based on at least one first subject object in the plurality of second images and the plurality of second images; wherein the media file is an image or a video, the media file presents a spliced ​​image obtained by splicing at least a portion of the scene corresponding to each of the plurality of second images; and in the spliced ​​image presented by the media file, the object area of ​​the same first subject object covers the portion of the scene corresponding to each of at least two second images.

[0005] Secondly, this application also provides a media file generation apparatus, comprising: an image generation module, configured to acquire a first image and a first text, and generate a plurality of second images based on the first image and the first text; wherein the first text is a scene description text of the media file to be generated; and a media file generation module, configured to generate a media file based on at least one first subject object in the plurality of second images and the plurality of second images; wherein the media file is an image or a video, and the media file presents a spliced ​​image obtained by splicing at least a portion of the scenes corresponding to each of the plurality of second images; and in the spliced ​​image presented by the media file, the object area of ​​the same first subject object covers the portion of the scenes corresponding to at least two of the second images.

[0006] Thirdly, this application also provides an electronic device, the electronic device comprising: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the media file generation method provided in this application.

[0007] Fourthly, this application also provides a computer-readable storage medium storing a computer program for executing the media file generation method provided in this application.

[0008] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the media file generation method provided in this application.

[0009] The technical solution provided in this application can generate multiple second images based on an acquired first image and first text (screen description text of the media file to be generated). Based on at least one first subject object in the multiple second images and the multiple second images, a media file is generated. This media file is either an image or a video. The media file presents a stitched image obtained by splicing at least a portion of the screen corresponding to each of the multiple second images. Furthermore, in the stitched image presented by the media file, the object area of ​​the same first subject object covers the corresponding portions of at least two second images. Through this method, users only need to provide the first image and the screen description text of the media file to be generated to obtain a media file with multi-screen stitching effects and cross-screen presentation of the subject object. This can significantly improve the content creation capabilities of multimedia editing software and help meet users' diverse media file creation needs.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This application provides an illustration of an application scenario. Figure 2 A flowchart illustrating a method for generating media files provided in this application; Figure 3 A schematic diagram of a media file generation process provided in this application; Figure 4 A schematic diagram illustrating the generation effect of a media file provided in this application; Figure 5 A schematic diagram of a media file generation device provided in this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation

[0014] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0015] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the type, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0016] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.

[0017] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0018] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.

[0019] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.

[0020] To facilitate understanding of the application scenarios of the media file generation method provided in this application, exemplarily, Figure 1 This application provides an illustration of an application scenario for interaction between a user and a user terminal. Specifically, the user can input a reference image and a description text of the media file to be generated into a user terminal such as a mobile phone or computer. This description text can be used as model prompts. The user terminal can execute the media file generation method provided in this application based on the user input information to generate a media file with multi-screen splicing effects and cross-screen presentation of the main object, and display the media file on the interface. The user can intuitively view the media file through the user terminal interface, and can subsequently download and save the media file, etc. This application does not restrict the subsequent application and processing of the media file. Furthermore, Figure 1 The document does not explicitly specify the server-side functionality. In practical applications, the user client can generate media files independently, or it can interact with the server to collaboratively generate media files. Alternatively, the user client can provide reference images and a description of the scene to be generated to the server, which then generates the media file. It should be noted that... Figure 1 This is merely an application example provided in this application and should not be considered a limitation. In practical applications, users may also provide only reference images to the user terminal without providing the image description text of the media file to be generated, and the media file generation method can still be executed on this basis.

[0021] Figure 2 This is a flowchart illustrating a method for generating a media file provided in this application. This method can be executed by a media file generating device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 2 As shown, the method mainly includes the following steps S202 to S204: Step S202: Obtain the first image and the first text, and generate multiple second images based on the first image and the first text; wherein, the first text is the image description text of the media file to be generated.

[0022] The first image is a user-input image, and the first text can be user-input text or text intelligently generated based on the user-input image. In some specific examples, the user expects the media file to be generated to present the effect of multiple screens stitched together, so the first text can contain text descriptions corresponding to each of the multiple screens presented in the media file. In some implementations, an image import control and a text receiving control can be displayed on the interactive interface. The first image (i.e., the aforementioned reference image) is received through the image import control, and the first text is received through the text receiving control. The first image can be the user-input reference image, and the first text is the screen description text of the media file to be generated, which is input by the user. This application does not limit the content and format of the first text. For example, to ensure the subsequent generation effect, a prompt text can be displayed on the interactive interface to provide the user with an example description of the first text for the user's reference. In other implementations, in response to receiving the first image through the image import control but not receiving the first text through the text receiving control, the first text is generated based on the first image, such as by calling the generation model to actively generate the first text. In other implementations, only the image import control can be displayed on the interactive interface, and by default, only the user needs to import the first image, and the first text is generated directly based on the first image imported by the user in the subsequent implementations. In other implementations, an information receiving control that can simultaneously receive images and text can be provided, eliminating the need to separately set up image import controls and text receiving controls. The first image and the first text can be received directly through the information receiving control. If the information received by the information receiving control only contains the first image, the first text can then be generated based on the first image.

[0023] Step S204: Based on at least one first subject object in multiple second images and multiple second images, generate a media file; wherein the media file is an image or a video, and the media file presents a spliced ​​image obtained by stitching together at least a portion of the images corresponding to each of the multiple second images; and in the spliced ​​image presented by the media file, the object area of ​​the same first subject object covers the portion of the images corresponding to at least two second images.

[0024] This application does not limit the object category of the first subject object, such as it can be a person, animal, plant, or any object. Multiple second images may contain one or more subject objects. The first subject object is not limited here; for example, the first subject object includes one or more of the following: any subject object in multiple second images, the subject object that appears most frequently in multiple second images, the subject object located in the foreground in multiple second images, the subject object specified by the first text, the subject object presented in the first image, and the subject object contained in a selected third image from multiple second images. In some examples, a media file can be efficiently generated using a generative model based on at least one first subject object in multiple second images and the multiple second images. This media file can be a synthesized multi-grid image or a synthesized multi-grid video. It should be noted that this application does not limit the video length. If the video length is very short, it will present a dynamic image effect, making the media content more vivid and impactful.

[0025] Using the above method, users only need to provide the first image and the image description text of the media file to be generated to obtain a media file with multi-image splicing effect and subject object cross-image presentation effect. Compared with a single image, the final media file conveys richer information. The subject object cross-image effect can create a visual effect of the subject breaking through the border, making the relationship between different images more intuitive. It can better improve the content creation capabilities of multimedia editing software and help meet the diverse media file creation needs of users.

[0026] Considering that the first text may be abbreviated text input by the user, in order to further improve the generation effect of the second image, for example, the step of generating multiple second images based on the first image and the first text in step S202 above can be performed with reference to the following steps a and b: Step a: Based on the first text and the first image, generate the second text using the first generation model; wherein the second text is the content-enhanced text corresponding to the first text.

[0027] This application does not limit the specific implementation of the first generation model. The first generation model can expand the content of the first text based on the first image, such as by adding descriptive details, to obtain a second text with richer and more detailed content. For example, the first text contains descriptive text corresponding to each of the multiple frames presented in the media file to be generated, and the second text obtained by rewriting the first text also contains more detailed textual descriptions corresponding to each of the multiple frames presented in the media file to be generated. For example, the first image is a picture of a cat, and the first text describes the content of multiple scenes. Here is an example of the text description of scene one in the first text: "Scene 1 (Close-up): A pet's paw is on an open picture book, the page contains the moon and stars." After being rewritten using the first generation model, the example of the text description of scene one in the second text is: "Image 1: In a close-up shot, a white cat (facial features refer to the uploaded image) gently presses its pink and white paws on an open picture book, the pages of which are printed with moon and star patterns. Warm yellow spotlight shines from the lamp above, illuminating the cat's fluffy paw pads and the texture of the pages. The blurred background of the bedroom environment exudes a dark and warm feeling, and the grainy texture of the film makes the image more delicate and soft." It is understandable that the initial text input by the user is usually quite brief, and directly using it for image generation may result in poor output that fails to meet user needs. By using the image-guided text expansion method described above, the first image is used as context, enabling the first generation model to more accurately transform visual semantics into structured and usable text information. Based on this, the content of the first file is expanded to include visual details such as color, pose, lighting, background, and material, which helps to significantly improve the richness, realism, and appeal of the final generated image.

[0028] Step b involves generating multiple second images based on the first image and the second text using a second generation model. This application does not limit the specific implementation of the second generation model; it can be the same as or different from the first generation model. The second generation model can use the first image as a reference image. Based on the reference image, it generates corresponding second images according to the descriptive text for each scene contained in the second text. The multiple second images can have the same or different sizes; the size of the second images can be flexibly set according to requirements. This rewriting allows the second text to provide a deeper level of detail in the scene and the main objects within it, helping the second generation model generate detailed, rich, and user-relevant second images based on the first image and the second text.

[0029] For example, the first subject object is the subject object contained in the third image selected from multiple second images. The image selected from multiple second images is called the third image. This application does not limit the selection strategy of the third image. For example, the first subject object to be highlighted can be determined based on the first text, and the second image corresponding to the screen to which the first subject object belongs can be selected as the third image; or, for example, the screen located at a specified position can be determined based on the first text, and the second image corresponding to the screen located at the specified position can be selected as the third image. Based on this, the above step S204, that is, the step of generating a media file based on at least one first subject object in multiple second images and multiple second images, can be performed with reference to the following steps A to C: Step A involves enlarging the size of the third image among multiple second images (also known as image expansion processing) to obtain a fourth image. For example, the enlargement parameters of the third image can be obtained, and the third image among multiple second images can be enlarged based on these parameters to obtain the fourth image. For example, these enlargement parameters can be default parameters or determined based on the size of the main object in the third image. For instance, the enlargement parameters may require that the main object in the enlarged fourth image extends beyond the border of the original third image, thus achieving a better effect of the main object crossing the frame; this effect can also be called an out-of-frame effect.

[0030] Step B involves compositing the first main object in the fourth image, the fourth image, and multiple second images (excluding the third image) to obtain a composite image. This composite image can present a multi-screen stitching effect and a main object appearing across multiple screens. In some specific examples, step B can be performed as follows: Step B1 involves cropping the fourth image to obtain the fifth image; the fifth image has the same size as the third image. In practical applications, the cropping method for the fourth image can be flexibly set. For example, the cropping area of ​​the fourth image can be determined first, and then cropped based on this area to obtain the fifth image. This application does not limit the method for determining the cropping area of ​​the fourth image. For example, a reference point on the fourth image can be determined first, and the cropping area of ​​the fourth image can be determined based on the size of the third image and the reference point. The reference point can be the center point or vertex of the fourth image, or a specified vertex within the fourth image determined based on preset parameters; no limitation is imposed here. Through the above method, a fifth image with the same size as the third image can be obtained for easy stitching, and the content presented in the fifth image is magnified compared to the third image. In other words, the fifth image is an image obtained by magnifying a portion of the third image.

[0031] Step B2 involves performing image cutout processing on the first main object in the fourth image to obtain the object layer of the first main object. An image cutout algorithm can be used to cut out the first main object in the fourth image. The resulting object layer of the first main object can be a layer with an alpha channel, retaining only the object area of ​​the first main object while the rest of the area is transparent.

[0032] Step B3 involves compositing the fifth image, the object layer of the first main object, and multiple second images (excluding the third image) to obtain a composite image. In some specific examples, step B3 can be performed as follows: Steps B3.1 to B3.3. Step B3.1: Obtain the stitching positions corresponding to the fifth image and each of the multiple second images (excluding the third image). In some examples, multiple stitching positions can be pre-set, and then, based on a preset position determination strategy, the stitching positions corresponding to the fifth image and each of the multiple second images (excluding the third image) can be determined from these multiple stitching positions. This position determination strategy includes: random allocation, determining the stitching positions of the images corresponding to each image based on the description of each image in the first text, or, prioritizing the determination of the stitching position of the fifth image, and then determining the stitching positions of each of the multiple second images (excluding the third image). For ease of understanding, the following example illustrates the stitching of three images into a three-grid image. The stitching positions include the top, middle, and bottom positions. The middle position is set as the stitching position corresponding to the fifth image, and the other two images correspond to the top and bottom positions, respectively.

[0033] Step B3.2: Based on the stitching position, the fifth image is stitched together with multiple second images (excluding the third image) to obtain the sixth image. The sixth image is the one that presents the multi-image stitching effect.

[0034] Step B3.3: Based on the position of the first subject object in the fourth image, the object layer of the first subject object is overlaid on the sixth image to obtain a composite image. In some specific implementation examples, the fifth image is an image cropped from the magnified fourth image. The object layer of the first subject object is obtained based on the regional position of the first subject object in the fourth image, and it also presents the magnified first subject object. By overlaying the object layer of the first subject object on the sixth image according to the position of the first subject object in the fourth image, not only can the area of ​​the first subject object on the screen corresponding to the fifth image presented in the sixth image be completely covered, but also the effect of the subject object spanning across the screen can be presented.

[0035] Through steps B3.1 to B3.3 above, the fifth image, the object layer of the first main object, and the images of multiple second images other than the third image can be synthesized in an orderly and efficient manner, ensuring that the final synthesized image meets the requirements. Through steps B1 to B3 above, the method of cropping and cutting out the enlarged fourth image and finally synthesizing it based on the cropped fifth image, the cut-out object layer, and the images of multiple second images other than the third image is cleverly and conveniently realized to achieve multi-screen splicing effect and main object cross-screen effect.

[0036] Step C: Obtain the media file based on the synthesized image. In practical applications, the synthesized image can be used directly as the media file, or it can be post-processed, such as optimized, to further enhance the visual effect of the media file. Through the above steps A to C, media files with multi-screen splicing effects and cross-screen presentation of the main object can be obtained efficiently and reliably, helping to meet the diverse media file creation needs of users.

[0037] In some specific examples, step C above can be performed by referring to steps (1) to (2) below: Step (1): Based on the synthesized image, generate the corresponding caption for the synthesized image using a third generation model. This application does not limit the specific implementation of the third generation model; the third generation model can be the same as or different from the aforementioned first or second generation model. In practical applications, the caption for the synthesized image includes matching text for at least one of the multiple images presented by the synthesized image. Furthermore, the font format, font size, and other attribute information of the caption can be determined based on the style of the image related to the caption in the synthesized image. This application does not limit the conditions for generating the caption. For example, the default is to generate the caption for the synthesized image; or, in response to detecting that the first text carries a caption generation requirement, the above step (1) is executed; or, on the interactive interface, the user is provided with an option to generate the caption, and the above step (1) is executed only in response to receiving a caption generation instruction. Specific settings are flexible and not limited here.

[0038] Step (2): Overlay the text on the image onto the composite image to obtain the media file.

[0039] By overlaying text onto a composite image, and incorporating the text as part of the composition, it not only helps to efficiently convey the image's theme and transmit more information through the text, such as giving the image a sense of story and helping users understand it, but also helps to enhance the overall visual hierarchy.

[0040] In some other specific examples, step C above can be performed by referring to steps 1 to 2 below: Step 1: Based on the synthesized image, generate a synthesized video using the fourth generation model. This application does not limit the specific implementation of the fourth generation model; it can be the same as or different from the aforementioned first to third generation models. The fourth generation model can be a video generation model, generating a matching video based on the synthesized image. There is no limitation on the video length; furthermore, if the synthesized video is short, it can also present a dynamic image effect.

[0041] Step 2: Obtain the media file based on the synthesized video. In some examples, the synthesized video can be directly used as the media file. In other examples, the synthesized video can be further optimized, such as generating corresponding captions for the synthesized video through a third generation model, and then overlaying the corresponding captions frame by frame onto the synthesized video to obtain the media file, which helps to further enhance the appeal of the media file.

[0042] Based on the foregoing, this application provides a specific example of a method for generating media files. In this example, three images are used as the second image for illustration. See [link to relevant documentation]. Figure 3 The diagram illustrates a media file generation process. A first image and first text are used by a first generation model to generate second text. The first image and second text are then used by a second generation model to generate three second images: image one, image two, and image three. Figure 3 Image 2 is used as the third image mentioned above. It undergoes enlargement and image matting processes sequentially. The results are then combined with images 1 and 3 to create a three-panel image. This three-panel image can generate captions using a third generation model, and a composite video (or a GIF if short) can be generated using a fourth generation model. Frame-by-frame text overlay can then be applied to the composite video and captions to obtain the final media file. In the example above, three interconnected images can be generated from a single user image and then stitched together vertically to form a top-to-bottom cinematic narrative. The three-panel format is more compatible with vertical screen browsing habits and conveys richer information compared to traditional single images. It should be noted that... Figure 3 This is just one example. In practical applications, the three-grid image or three-grid video can be used directly as the final media file. Alternatively, the result of overlaying the three-grid image with the accompanying text can be used as the final media file. Image three can be used as the aforementioned third image, or both image two and image three can be used as the third image. The number of second images does not have to be three. For example, it can be presented as a four-grid or other grid effects. Furthermore, it is not limited to the above-mentioned method of merging multiple images vertically. It can also be merged horizontally, etc. There are no restrictions here.

[0043] For ease of understanding, a schematic diagram illustrating the effect of applying the media file generation method of this application is provided herein. See [link / reference]. Figure 4 The diagram illustrates the effect of generating a media file, showing that the user inputs a first image (containing a puppy) and a first text (the image description text of the media file to be generated). By performing the media file generation operation, the resulting media file is a three-panel image of a puppy spanning the screen. The media file generation operation mainly follows the steps of the media file generation method provided in this application. For example, Figure 4 The first text entered by the user includes the following: "The scene is set in a tranquil coniferous forest covered in thick snow. Overall style: realistic and warm texture, soft natural light, delicate plush texture, dynamic snowfall, and a tranquil and healing atmosphere. Scene 1 (close-up / face close-up): A close-up of the pet's face, with eyes closed and a relaxed expression. The plush texture of the white earmuffs and red scarf is clear and delicate, with tiny snowflakes falling on them. The background forest is blurred with a shallow depth of field, and the lighting is soft. Scene 2 (close-up / head tilt): The pet tilts its head back, looking at the falling snowflakes, with a peaceful expression. There is a small amount of snow on the red scarf, and the white earmuffs are fluffy and cute. The background is blurred, highlighting the dynamic of the falling snowflakes and the tranquil and healing atmosphere. Scene 3 (mid-long shot): The pet stands on the fluffy snow, wearing white plush earmuffs and a thick red knitted scarf. The background is a tall pine forest covered in snow, with snowflakes slowly falling from the sky. The pet's posture is natural, and the red scarf adds a touch of warmth to the cold snow scene."

[0044] Based on the first image, the first text is expanded using the first generation model to obtain the second text, which includes: "Image 1: Close-up shot, a small dog with light brown fur (facial features refer to the uploaded image) with its eyes closed, relaxed expression, wearing white earmuffs and a red scarf, the fur texture is clear and delicate, with small snowflakes falling on it. The background is a tranquil coniferous forest covered with thick snow, with shallow depth-of-field blurring, soft and natural lighting, dynamic snowfall, and a tranquil and healing atmosphere. Image 2: Close-up shot, centered composition, a small dog with light brown fur (facial features refer to the uploaded image) with its head tilted back, looking at the falling snowflakes, with a peaceful expression, wearing..." White earmuffs and a red scarf, the red scarf adorned with a few snowflakes, the white earmuffs fluffy and cute. The background is a tranquil coniferous forest covered in thick snow, blurred, with soft, natural lighting, and dynamic snowfall, creating a serene and healing atmosphere. Image 3: A wide shot, the subject is positioned slightly to the left, occupying the left side of the frame. A small dog with light brown fur (facial features referenced in the uploaded image) stands on fluffy snow, wearing white fluffy earmuffs and a thick red knitted scarf. Its posture is natural, the red scarf adding a touch of warmth to the stark snowy landscape. The background is a tranquil coniferous forest covered in thick snow, with tall cedar trees laden with snow, snowflakes gently falling from the sky, soft, natural lighting, and a serene and healing atmosphere. Based on the first image input by the user and the aforementioned second text, multiple second images are generated through the second generation model. Subsequently, by executing the aforementioned steps A to C, a three-panel effect image of a puppy appearing across the screen can be obtained. In practical applications, it is also possible to further obtain a three-panel video of a puppy appearing across the screen, etc., without any restrictions here.

[0045] Using the above method, users only need to provide the first image and the image description text of the media file to be generated to obtain a media file with multi-image splicing effect and subject object cross-image presentation effect. Compared with a single image, the final media file conveys richer information. The subject object cross-image effect can create a visual effect of the subject breaking through the border, making the relationship between different images more intuitive, which can better improve the content creation capabilities of multimedia editing software and help meet the diverse media file creation needs of users.

[0046] Corresponding to the aforementioned method for generating media files, this application further provides a media file generation apparatus. Figure 5 This application provides a schematic diagram of a media file generation device. This device can be implemented by software and / or hardware, and is generally integrated into an electronic device, such as... Figure 5 As shown, the media file generation device includes: The image generation module 502 is used to acquire a first image and a first text, and generate multiple second images based on the first image and the first text; wherein, the first text is the image description text of the media file to be generated; The media file generation module 504 is used to generate a media file based on at least one first subject object in a plurality of second images and the plurality of second images; wherein the media file is an image or a video, and the media file presents a spliced ​​image obtained by stitching together at least a portion of the images corresponding to each of the plurality of second images; and in the spliced ​​image presented by the media file, the object area of ​​the same first subject object covers the portion of the images corresponding to at least two of the second images.

[0047] With the above-mentioned device, users only need to provide the first image and the screen description text of the media file to be generated to obtain a media file with multi-screen splicing effect and the main object presented across screens. This can significantly improve the content creation capabilities of multimedia editing software and help meet users' diverse media file creation needs.

[0048] In some embodiments, the image generation module 502 is specifically used to: generate second text based on the first text and the first image using a first generation model; wherein the second text is content-enhanced text corresponding to the first text; and generate multiple second images based on the first image and the second text using a second generation model.

[0049] In some implementations, the first subject object is the subject object contained in the third image selected from the plurality of second images; the media file generation module 504 is specifically used to: enlarge the size of the third image in the plurality of second images to obtain a fourth image; perform a composite processing based on the first subject object in the fourth image, the fourth image, and the images in the plurality of second images other than the third image to obtain a composite image; and obtain a media file based on the composite image.

[0050] In some embodiments, the media file generation module 504 is specifically used to: crop the fourth image to obtain a fifth image; wherein the fifth image has the same size as the third image; perform image cutout processing based on the first subject object in the fourth image to obtain an object layer of the first subject object; and perform compositing processing based on the fifth image, the object layer of the first subject object, and the images other than the third image among the plurality of second images to obtain a composite image.

[0051] In some embodiments, the media file generation module 504 is specifically used to: obtain the splicing positions corresponding to the fifth image and each of the plurality of second images except the third image; based on the splicing positions, perform splicing processing on the fifth image and the plurality of second images except the third image to obtain a sixth image; based on the position of the first subject object in the fourth image, overlay the object layer of the first subject object on the sixth image to obtain a composite image.

[0052] In some implementations, the media file generation module 504 is specifically used to: generate a caption corresponding to the composite image based on the composite image using a third generation model; and overlay the caption onto the composite image to obtain a media file.

[0053] In some implementations, the media file generation module 504 is specifically used to: generate a synthesized video based on the synthesized image using a fourth generation model; and obtain a media file based on the synthesized video.

[0054] The media file generation apparatus provided in this application can execute the media file generation method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the method.

[0055] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0056] This application provides an electronic device, comprising: a storage device storing a computer program thereon; and a processing device for executing the computer program in the storage device to implement the steps of any of the methods in this application.

[0057] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device 600 suitable for implementing this application. The terminal device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Devices), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functionality and scope of this application.

[0058] like Figure 6 As shown, electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0059] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0060] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined in the methods of this application.

[0061] In addition to the methods and devices described above, embodiments of this application may also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the methods provided in this application. The computer program product may be written in any combination of one or more programming languages ​​to perform the operations of this application. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code may be executed entirely on a user's computing device, partially on a user's device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0062] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the methods provided in this application.

[0063] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0064] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the method of this application.

[0065] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0066] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating a media file, comprising: A first image and a first text are acquired, and multiple second images are generated based on the first image and the first text; wherein, the first text is a scene description text of the media file to be generated; A media file is generated based on at least one first subject object in the plurality of second images and the plurality of second images; wherein the media file is an image or a video, and the media file presents a spliced ​​image obtained by stitching together at least a portion of the images corresponding to each of the plurality of second images; and in the spliced ​​image presented by the media file, the object area of ​​the same first subject object is covered on the portion of the images corresponding to at least two of the second images.

2. The method according to claim 1, wherein, The generation of multiple second images based on the first image and the first text includes: Based on the first text and the first image, a second text is generated using a first generation model; wherein the second text is an expanded text corresponding to the first text. Based on the first image and the second text, multiple second images are generated using a second generation model.

3. The method according to claim 1, wherein, The first subject object is the subject object contained in the third image selected from the plurality of second images; the step of generating a media file based on at least one first subject object in the plurality of second images and the plurality of second images includes: The third image among the multiple second images is enlarged to obtain the fourth image; A composite image is obtained by combining the first main object in the fourth image, the fourth image, and the images in the plurality of second images other than the third image. Media files are obtained based on the synthesized image.

4. The method according to claim 3, wherein, The composite image is obtained by combining the first main object in the fourth image, the fourth image, and images from the plurality of second images other than the third image, including: The fourth image is cropped to obtain the fifth image; wherein the fifth image has the same size as the third image; Based on the first main object in the fourth image, a cutout process is performed to obtain the object layer of the first main object; A composite image is obtained by combining the fifth image, the object layer of the first subject object, and the images other than the third image from the plurality of second images.

5. The method according to claim 4, wherein, The composite image is obtained by combining the fifth image, the object layer of the first subject object, and the images other than the third image from the plurality of second images, including: Obtain the stitching position corresponding to the fifth image and each of the multiple second images except for the third image; Based on the stitching position, the fifth image is stitched together with the images other than the third image from the plurality of second images to obtain the sixth image; Based on the position of the first subject object in the fourth image, the object layer of the first subject object is overlaid on the sixth image to obtain a composite image.

6. The method according to claim 3, wherein, The process of obtaining a media file based on the synthesized image includes: Based on the synthesized image, a corresponding caption is generated using a third generation model; The accompanying text is overlaid on the composite image to obtain a media file.

7. The method according to claim 3, wherein, The process of obtaining a media file based on the synthesized image includes: Based on the synthesized image, a synthesized video is generated using a fourth generation model; Media files are obtained based on the synthesized video.

8. A media file generation apparatus, comprising: An image generation module is used to acquire a first image and a first text, and generate multiple second images based on the first image and the first text; wherein, the first text is a scene description text of the media file to be generated; A media file generation module is used to generate a media file based on at least one first subject object in the plurality of second images and the plurality of second images; wherein the media file is an image or a video, and the media file presents a spliced ​​image obtained by stitching together at least a portion of the images corresponding to each of the plurality of second images; and in the spliced ​​image presented by the media file, the object area of ​​the same first subject object covers the portion of the images corresponding to at least two of the second images.

9. An electronic device, the electronic device comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method for generating a media file according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for generating a media file according to any one of claims 1-7.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method for generating a media file according to any one of claims 1-7.