Method and system for generating multimodal content and sequential multimodal content

The method and system leverage multiple LLMs and multimodal models to iteratively refine narrative text and audio-video content, overcoming the limitations of existing video generation technologies by producing detailed and continuous multimodal content.

WO2026062683A1PCT designated stage Publication Date: 2026-03-26SEKHAR MADDULA NAGA VENKATA +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing video generation technologies are incapable of producing indefinite, detailed digital multimodal content based on user input, particularly lacking the ability to generate content with a detailed narrative text including characters.

Method used

A method and system utilizing multiple instances of Large Language Models (LLMs) and multimodal models to iteratively refine narrative text, script, and audio-video content, ensuring seamless transitions and inclusion of necessary elements, using prompt engineering to enhance and validate each stage.

Benefits of technology

Enables the generation of detailed, continuous, and high-quality multimodal content by iteratively refining narrative text, script, and audio-video components, addressing the limitations of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025051269_26032026_PF_FP_ABST
    Figure IN2025051269_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure describes method and system for generating multimodal content. The method comprising receiving input to generate multimodal content. The method further comprises of generating narrative text for the multimodal content using first instance of Large Language Models (LLM) based on the input. The method further comprises of generating script comprising one or more sections of the multimodal content, based on the narrative text using second instance of the LLM. The method further comprises of enhancing each of the one or more sub-sections using third instance of the LLM. The method further comprises of generating audio-video corresponding to each of one or more enhanced sub-sections using first instance of multimodal model. The method further comprises of converging the audio-video corresponding to each of one or more enhanced sub-sections for generating multimodal content, using second instance of multimodal model. The method further comprises of displaying multimodal content on electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE: METHOD AND SYSTEM FOR GENERATING MULTIMODAL CONTENT AND SEQUENTIAL MULTIMODAL CONTENT

[0002] TECHNICAL FIELD

[0003]

[0001] The present disclosure relates to the general filed of Generative Artificial Intelligence (GenAI), cloud computing and digital content generation. Particularly, but not exclusively, the present disclosure relates to a method and a system for generating multimodal content.

[0004] BACKGROUND ART

[0005]

[0002] Existing video generation technologies are capable of producing videos ranging from a few seconds to a few of minutes based on user input. While these technologies are able to generate short videos based on user inputs they do not have the capability to generate an indefinite, detailed digital multimodal content, including multimodal content based on a detailed narrative text including characters.

[0006]

[0003] Thus, it is desired to address the above-mentioned disadvantages or other shortcomings and at least provide a useful alternative.

[0007]

[0004] The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.

[0008] SUMMARY

[0009]

[0005] The present disclosure tries to address the aforesaid problem associated with existing technologies.

[0010]

[0006] In an embodiment, the present disclosure relates to a method for generating multimodal content. The method comprises of receiving an input to generate a multimodal content. The method further comprises of generating narrative text for the multimodal content using first instance of Large Language Models (LLM) based on the input. The method further comprises of generating script comprising one or more sections of the multimodal content, based on the narrative text using second instance of the LLM. The method further comprises of enhancing each of the one or more sub-sections using third instance of the LLM. The method further comprises of generating audio-video corresponding to each of the one or more enhanced sub- sections using first instance of multimodal model. The method further comprises of converging the audio-video corresponding to each of the one or more enhanced sub-sections for generating the multimodal content, using second instance of the multimodal model. The method further comprises of displaying the multimodal content on electronic device.

[0011]

[0007] In another embodiment, the present disclosure relates to a system for generating multimodal content. The system comprising a processor and a memory. The memory stores processor-executable instructions, which, on execution, cause the processor to receives an input to generate a multimodal content. Subsequently, the processor is configured to generate a narrative text for the multimodal content using first instance of Large Language Models (LLM) based on the input. Subsequently, the processor is configured to generate a script comprising one or more sections of the multimodal content, based on the narrative text using second instance of the LLM. Subsequently, the processor is configured to enhance each of the one or more sub-sections using third instance of the LLM. Subsequently, the processor is configured to generate audio-video corresponding to each of the one or more enhanced sub-sections using first instance of multimodal model. Subsequently, the processor is configured to converge the audio-video corresponding to each of the one or more enhanced sub-sections for generating the multimodal content, using second instance of the multimodal model. Subsequently, the processor is configured to display the multimodal content on electronic device.

[0012]

[0008] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.

[0013] BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS

[0014]

[0009] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of system and methods in accordance with embodiments of the present subject matter are now described below, by way of example only, and with reference to the accompanying figures.

[0010] Fig. 1 illustrates an environment diagram for generating multimodal content in accordance with some embodiments of the present disclosure.

[0015] [Oil] Fig. 2 shows a detailed block diagram of a system for generating multimodal content in accordance with some embodiments of the present disclosure.

[0016]

[0012] Fig. 3 illustrates a flowchart of a method for generating multimodal content in accordance with some embodiments of the present disclosure.

[0017]

[0013] Fig. 4 is a block diagram of an exemplary system for implementing embodiments consistent with the present disclosure.

[0018]

[0014] It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flowcharts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0019] DETAILED DESCRIPTION

[0020]

[0015] In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.

[0021]

[0016] While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the scope of the disclosure.

[0022]

[0017] The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises. . . a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or method.

[0023]

[0018] In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.

[0024]

[0019] As used herein, the term “Large Language Model (LLM)”, refers to a very large deep learning models that is pre-trained on vast amounts of data with an underlying transformer. It is a set of neural networks that consist of an encoder and a decoder with self-attention capabilities. The encoder and decoder extract meanings from a sequence of text and understand the relationships between words and phrases in it. Transformer LLMs are capable of unsupervised training. Transformers process entire sequences in parallel.

[0025]

[0020] As used herein, the term “multimodal model” refers to an architecture that can process a wide variety of inputs, including text, images, and audio, as prompts and convert those prompts into various outputs, such as text, images, and audio.

[0026]

[0021] As used herein, the term “agent” refers to an agent in a multi-agent architecture. A multiagent architecture is a branch of artificial intelligence, which consists of multiple decisionmaking agents which interact in a shared environment to achieve common or conflicting goals. Multi-agent architectures provide higher levels of adaptability in order to support applications that exhibit emergent behaviour. Agents interact with other components and have an open- ended task and is capable of communicating with other agents to achieve their assigned task. The multi-agent architecture provides discovery and communication services to its agents. An agent is an active task-oriented component that plays one or more roles in the multi-agent architecture environment. Each agent perpetually executes a control loop

[0027]

[0022] Fig. 1 illustrates an environment diagram comprising of a system 100 for generating multimodal content in accordance with some embodiments of the present disclosure. The system 100 is configured to receive an input 101 to generate a multimodal content. In an embodiment the input 101 is one of an action prompt to create the multimodal content based on existing multimodal content, a text based prompt comprising a brief of the multimodal content, or a prompt based on historical data of a user. In another embodiment, the input 101 may be received from a user.

[0028]

[0023] In an embodiment, the action prompt to create the multimodal content based on existing multimodal content may correspond to a click of a button to create a next episode based on a current episode of a multimodal content. In an embodiment, the text based prompt comprising a brief of the multimodal content may comprise of one of: one or more words or one or more sentences. An example of the text based prompt, without limitation to the following: “generate a story based on two characters”. In an embodiment, the text based prompt may comprise of one of a storyline comprising of one or more characters and a brief of the multimodal content expected. In an embodiment, the historical data comprises of user personalization data corresponding to a mobile application or a website or a website application. An example of the prompt based on historical data of the user, may be without limitation to: if the user personalization data of the user indicates that the user watches digital content based on action / thriller genre, the prompt may be: “generate a narrative text based on action / thriller genre”.

[0029]

[0024] In an embodiment, the system 100 may make API calls to one or more Large Language Models (LLM’s) 102. The LLM 102 may be a text based LLM which takes an input in the form of text and outputs a text. The system 100 may make API calls to one or more multimodal models 103 that take an input in the form of one of a text, an audio, or a video and outputs one of a text, an audio, or a video. The system 100 may make an API call to a first instance of the LLM 102a, to generate a narrative text for the multimodal content using the input 101. In an embodiment, the narrative text may correspond to a story line including one or more characters and a plotline based on the input 101. In an embodiment, the narrative text generated may include the one or more characters disclosed in the text based prompt input 101 and may be based on the brief of the multimodal content expected, disclosed in the text based prompt input 101. For example, without limitation to, based on the text based prompt: input 101, “generate a story based on two characters, at least one of them being a frog”, the following narrative may be generated by the system 100 using the first instance of the of the LLM 102a: “Two friends, John-the doe and Nightingale decided to go on a picnic to a forest. They trekked to the forest, and on the shores of a lake they set up their picnic. John-the doe and Nightingale ate the picnic that John-the doe had packed and had a gala time. They returned to their homes by sunset.”

[0025] In an embodiment, the system 100 may make an API call to a fourth instance of the LLM 102d, to validate the narrative text. The system 100 is configured to iteratively refine the narrative text using the fourth instance of the LLM 102d. The narrative text i s iteratively refined until the narrative text generated by the first instance of the LLM 102a meets a threshold requirement based on a plurality of first performance metrics. The plurality of first performance metrics includes, without limitation to, narrative continuity, omission of characters, efficiency, amount of narrative text generated etc. The output generated by the first instance of the LLM 102a, i.e., the narrative text, is iteratively passed to the fourth instance of the LLM 102d for validation. One or more improvements suggested by the fourth instance of the LLM 102d are communicated to the first instance of the LLM 102a through the system 100. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in the input 101 that may have been omitted by the first instance of the LLM 102a, removal of narrative discontinuity or further elaboration of the narrative text. The iterative process may be continued until the first instance of the LLM 102a and the fourth instance of the LLM 102d are satisfied with the narrative text. In an embodiment, the fourth instance of the LLM 102d may use prompt engineering to suggest improvements in the iterative refinement process. For example, the input 101 in particular requested the generati on of a story based on two characters and at least one of them being a frog. However, the narrative text generated : “Two friends, John-the doe and Nightingale decided to go on a picnic to a forest. They trekked to the forest, and on the shores of a lake they set up their picnic. John-the doe and Nightingale ate the picnic that John-the doe had packed and had a gala time. They returned to their homes by sunset” does not include the character “frog”. Therefore, an improvement such as: “include frog in the story” may be suggested by the fourth instance of the LLM 102d to the first instance of the LLM 102a through the system 100. The first instance of the LLM 102a may revise the narrative text to: “Two friends, Frog and Nightingale decided to go on a picnic to a forest. They trekked to the forest, and on the shores of a lake they set up their picnic. Frog and Nightingale ate the picnic that Frog had packed and had a gala time. They returned to their homes by sunset”, based on the fourth instance of the LLM’s 102d suggestion. This revised narrative text may be validated narrative text.

[0030]

[0026] In an embodiment, the system 100 script generates a script based on the validated narrative text generated by the first instance of the LLM 102a, using an API call to a second instance of the LLM 102b. The script may comprise of one or more sections of the multimodal content and each section may comprise of one or more sub-sections. The script based on the validated narrative text may comprise of a detailed script. In an embodiment, the system 100 may use prompt engineering to generate the script. For example, without limitation to, based on the validated narrative text, “Two friends, Frog and Nightingale decided to go on a picnic to a forest. They trekked to the forest, and on the shores of a lake they set up their picnic. Frog and Nightingale ate the picnic that Frog had packed and had a gala time. They returned to their homes by sunset.”, the following script with one or more sections and their corresponding one or more sub sections may be generated by the second instance of the LLM 102b:

[0031] Section 1 : The Decision

[0032] Section 1 Sub-Section 1 : The Meeting

[0033] INT. LIVING ROOM - MORNING

[0034] Nightingale is sitting on a couch, scrolling through her phone. Frog enters, holding a picnic basket.

[0035] Frog: It’s a perfect day for a picnic. How about we head to the forest?

[0036] Nightingale: Sounds like a plan!

[0037] Section 2: The Journey

[0038] Section 2_Sub- Section 1 : The Trek

[0039] EXT. FOREST TRAIL - DAY

[0040] Nightingale and Frog walk along a winding trail, surrounded by towering trees. They chat and laugh, enjoying the fresh air.

[0041] Section 2_Sub-Section 2: Arrival at the Lake

[0042] EXT. LAKE SHORE - DAY

[0043] They arrive at a serene lake. The water glistens under the sunlight. Frog sets down the picnic basket, and they spread out a blanket on the grass.

[0044] Section 3 : The Picnic

[0045] Section 3 Sub -Section 1 : A Meal by the Lake

[0046] EXT. LAKE SHORE - DAY

[0047] Frog opens the basket, revealing sandwiches, fruits, and a thermos. Nightingale pours coffee while Frog hands her a sandwich. They eat, laugh, and enjoy the peaceful surroundings. Nightingale: This was a great idea.

[0048] Frog: Glad you think so. Section 4: The Return

[0049] Section 4_Sub-Section 1 : Heading Home

[0050] EXT. FOREST TRAIL - LATE AFTERNOON

[0051] The sun is beginning to set as Nightingale and Frog walk back along the trail, their laughter echoing through the trees.

[0052] Section 4_Sub-Section 2: Parting Ways

[0053] EXT. VILLAGE ROAD - SUNSET

[0054] They arrive back at their village. The sky is a soft orange hue.

[0055] Nightingale: Same time next week?

[0056] Frog: Absolutely.

[0057] They wave goodbye and walk to their respective homes.

[0058] THE END

[0059]

[0027] In an embodiment, the system 100 validates the script generated using the second instance of the LLM 102b. Validating the script comprises of, using an API call to contact the fifth instance of the LLM 102e to validate the script. The system 100 iteratively refines the script using the fifth instance of the LLM 102e. The script is iteratively refined until the script generated by the second instance of the LLM 102b meets a threshold requirement based on a plurality of second performance metrics. The plurality of second performance metrics includes, without limitation to, sub-section continuity, omission of characters, efficiency, amount of subsection generated. One or more improvements suggested by the fifth instance of the LLM 102e are communicated to the second instance of the LLM 102b through the system 100. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in the narrative text that may have been omitted by the second instance of the LLM 102b (for example, without limitation to, in the above cited example, if either of the two characters are omitted in a sub-section, this may be suggested as an improvement), removal of discontinuity between one or more section and each of their corresponding subsections (for example, without limitation to, in the above cited example, if the frog and the nightingale decide to go on a picnic in one sub-section, and in the subsequent sub-section, if the frog and nightingale are described to be heading back home without elaborating on the aspects of the picnic, a suggestion pertaining to adding a scene for the absent connecting subsection may be sent to the second instance of the LLM 102b) or further elaboration of the one or more sections of the script (for example, without limitation to, if the amount of content generated under each sub-section is insufficient, then “addition of more content” may be suggested as an improvement). The iterative process may be continued until both the second instance of the LLM 102b and the fifth instance of the LLM 102e are satisfied with the script. In an embodiment, the fifth instance of the LLM 102e may use prompt engineering to suggest improvements in the iterative refinement process.

[0060]

[0028] In an embodiment, the system 100 enhances each of the validated one or more subsections generated by the second instance of the LLM 102b using an API call to a third instance of the LLM 102c. In an embodiment, enhancing the one or more sub-sections may comprise of adding contextual information to each of the one or more sub-sections. In an embodiment, the contextual information is at least one of meta-data of a sub-section, location data (for example, without limitation to, in the above cited example, location data may correspond to dense forest, lake shore etc for each of the one or more sub-sections), background data, character data, object class data (object class may correspond to without limitation to, animal, human, thing, vehicle etc), clothing of characters (for example, in the above cited example, the clothing of the frog may correspond to a blue dress and the clothing of the nightingale may correspond to a purple frock), appearance of characters (for example, in the above cited example, the appearance of the frog may correspond to happy and the appearance of the nightingale may correspond to sad) etc.

[0061]

[0029] In an embodiment, the system 100 validates the enhanced one or more sub-sections generated by the third instance of the LLM 102c. Validating the one or more sub-sections comprises of, using an API call to contact the sixth instance of the LLM 102f, to iteratively refine the one or more enhanced sub-sections until the enhanced one or more sub-sections meets a threshold requirement based on a plurality of third performance metrics. The plurality of third performance metrics includes, without limitation to, narrative continuity, omission of characters, sufficiency of metadata, drastic change in metadata, etc. One or more improvements suggested by the sixth instance of the LLM 102f are communicated to the third instance of the LLM 102c through the system 100. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in one or more previous enhanced sub-sections that may have been omitted by the third instance of the LLM 102c (for example, without limitation to, in the above cited example, if either of the two characters are omitted in an enhanced sub-section, this may be suggested as an improvement), removal of discontinuity between one or more sections and each of their corresponding enhanced sub- sections (for example, without limitation to, in the above cited example, if the frog and the nightingale decide to go on a picnic in one enhanced sub-section, and in the subsequent enhanced sub-section, if the frog and nightingale are described to be heading back home without elaborating on the aspects of the picnic, a suggestion pertaining to adding a scene for the absent connecting sub-section may be sent to the third instance of the LLM 102c) or further elaboration of metadata corresponding to each of the one or more enhanced sub-sections (for example, without limitation to, if the amount of metadata generated for each of the enhanced sub-sections is insufficient, then “addition of more metadata” may be suggested as an improvement) and removal of drastic changes in the metadata (for e.g., if there is a drastic change in the appearance a character between two consecutive sub-sections, corresponding to a section, removal of such a discrepancy may be necessitated). The iterative process may be continued until the third instance of the LLM 102c and the sixth instance of the LLM 102f are satisfied with the one or more enhanced sub-sections. In an embodiment, the sixth instance of the LLM 102f may use prompt engineering to suggest improvements in the iterative refinement process.

[0062]

[0030] In an embodiment, the system 100 generates an audio-video corresponding to each of the one or more enhanced sub-sections by making an API call to the first instance of the multimodal model 103 a.

[0063]

[0031] In an embodiment, the system 100 validates the audio-video corresponding to each of the one or more enhanced sub-sections, generated by the first instance of the multimodal model 103a. Validating the audio-video corresponding to each of the one or more enhanced subsections comprises of, using an API call to contact a third instance of the multimodal model 103 c, to iteratively refine the audio-video corresponding to each of the one or more enhanced sub-sections until the audio-video corresponding to each of the one or more enhanced subsections meets a threshold requirement based on a plurality of fourth performance metrics. The plurality of fourth performance metrics includes, without limitation to, narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections. One or more improvements suggested by the third instance of the multimodal model 103 c are communicated to the first instance of the multimodal model 103 a through the system 100. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in one or more previous audio-video corresponding to each of one or more enhanced sub-sections that may have been omitted by the first instance of the multimodal model 103 a (for example, without limitation to, in the above cited example, if either of the two characters are omitted in the audio-video corresponding to each of one or more enhanced sub-sections, this may be suggested as an improvement) and removal of discontinuity between audio-video corresponding to each of one or more enhanced sub-sections (for example, without limitation to, in the above cited example, if the frog and the nightingale decide to go on a picnic in an audio-video corresponding to an enhanced subsection, and in the subsequent audio-video corresponding to the subsequent enhanced subsection, if the frog and nightingale are described to be heading back home without elaborating on the aspects of the picnic, a suggestion pertaining to adding a scene for the absent connecting enhanced sub-section may be sent to the first instance of the multimodal model 103a). The iterative process may be continued until both the first instance of the multimodal model 103a and the third instance of the multimodal model 103c are satisfied with the audio-video corresponding to the one or more enhanced sub-sections. In an embodiment, the third instance of the multimodal model 103 c may use prompt engineering to suggest improvements in the iterative refinement process.

[0064]

[0032] In an embodiment, the system 100 converges the audio-video corresponding to each of the one or more enhanced sub-sections by making an API call to a second instance of the multimodal model 103b, for generating the multimodal content. The converging of the audiovideo corresponding to each of the one or more enhanced sub-sections ensures seamless transition between an audio-video corresponding to one of the enhanced sub-sections and an audio-video corresponding to its subsequent enhanced sub-section.

[0065]

[0033] In an embodiment, the system 100 validates the multimodal content. Validating the multimodal content comprises of, the system 100 using an API call to contact a fourth instance of the multimodal model 103d, to iteratively refine the multimodal content until the multimodal content generated by the system 100 using the second instance of the multimodal model 103b meets a threshold requirement based on a plurality of fifth performance metrics. The plurality of fifth performance metrics includes, without limitation to, narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections. One or more improvements suggested by the by the fourth instance of the multimodal model 103 d are communicated to the second instance of the multimodal model 103b through the system 100. The one or more improvements / suggestions may include, without limitation to, removal of discontinuity between one or more audio-video corresponding to each of one or more enhanced sub-sections in the multimodal content (for example, without limitation to, in the above cited example, if the frog and the nightingale decide to go on a picnic in an audio-video corresponding to an enhanced sub-section, and in the subsequent audio-video corresponding to the subsequent enhanced sub-section, if the frog and nightingale are described to be heading back home without elaborating on the aspects of the picnic, a suggestion pertaining to adding a scene for the absent connecting enhanced sub-section may be sent to the second instance of the multimodal model 103b) and removal of seams between transition of one or more audio-videos corresponding to each of one or more enhanced sub-sections. The iterative process may be continued until both the second instance of the multimodal model 103b and the fourth instance of the multimodal model 103d are satisfied with the multimodal content. In an embodiment, the fourth instance of the multimodal model 103 d may use prompt engineering to suggest improvements in the iterative refinement process.

[0066]

[0034] In an embodiment the display unit 105 is configured to display the validated multimodal content on an electronic device.

[0067]

[0035] In an embodiment, the multimodal content may be displayed to the user via an interface configured to enable development, upload and share the multimodal content generated based on the input 101. In another embodiment, the multimodal content may be displayed via one or more mobile applications or streaming platforms. In an embodiment, the multimodal content may be stored in a cloud server.

[0068]

[0036] Fig. 2 shows a detailed block diagram of the system 100 for generating multimodal content in accordance with some embodiments of the present disclosure.

[0069]

[0037] The system 100 for generating multimodal content includes an Input-Output (I / O) interface 201, a processor 203, a memory 205 and the display unit 105. In the present embodiment, data 207 is stored within the memory 205.

[0070]

[0038] The I / O interface 201 is configured to receive the input 101. In another embodiment, the I / O interface 201 is configured to receive the input 101 from a user. The I / O interface 201 employs communication protocols or methods such as, without limitation, audio, analog, digital, monoaural, Radio Corporation of America (RCA) connector, stereo, IEEE®- 1394 high speed serial bus, serial bus, Universal Serial Bus (USB), infrared, Personal System / 2 (PS / 2) port, Bayonet Neill-Concelman (BNC) connector, coaxial, component, composite, Digital Visual Interface (DVI), High-Definition Multimedia Interface (HDMI®), Radio Frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE® 802.1 Ib / g / n / x, Bluetooth, cellular e.g., Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System for Mobile communications (GSM®), Long-Term Evolution (LTE®), Worldwide interoperability for Microwave access (WiMax®), or the like.

[0071]

[0039] The memory 205 is communicatively coupled to the processor 203 of the system 100 for generating multimodal content. The memory 205, also, stores processor-executable instructions which cause the processor 203 to execute the instructions for generating multimodal content. The memory 205 includes, without limitation, memory drives, removable disc drives, etc. The memory drives may further include a drum, magnetic disc drive, magnetooptical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.

[0072]

[0040] The processor 203 includes at least one data processor for generating multimodal content. The processor 203 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc.

[0073]

[0041] The display unit 105 may correspond to one of a Liquid Crystal Display (LCD) screen, a plasma screen, a Light Emitting Diode (LED) screen or an Organic LED (OLED) screen.

[0074]

[0042] The data 207 includes, for example, the input 101 received by the I / O interface. In an embodiment the input 101 is one of an action prompt to create the multimodal content based on existing multimodal content, a text based prompt comprising a brief of the multimodal content, or a prompt based on historical data of a user. In an embodiment, the input 101 may be received from the user. In an embodiment, the action prompt to create the multimodal content based on existing multimodal content may correspond to a click of a button to create the multimodal content. For example, a button may be clicked to generate a next episode based on a current episode of a TV series. In another example, the button may be clicked to generate a new movie based on a currently watched movie or a movie displayed on a screen. In an embodiment, the text based prompt may comprise a brief of the multimodal content, for example one or more words or one or more sentences such as continue the story of the AVENGERS™. In an embodiment, the text based prompt may comprise of one of a storyline comprising of one or more characters and a brief of the multimodal content expected. In an embodiment, the historical data comprises of user personalization data corresponding to a mobile application or a web application.

[0075]

[0043] The miscellaneous data 211 stores data, including meta data, and temporary files, generated by the units 213 of the system 100 for performing the various functions of the system 100.

[0076]

[0044] In the embodiment of the present disclosure, the data 207 in the memory 205 are processed by the one or more units 213 (also, referred as units) of the system 100. In an embodiment, the one or more units 213 may be implemented as software and may be implemented as virtual functions. In another embodiment, the one or more units 213 may be implemented as dedicated hardware units (e.g., circuits). As used herein, the term unit refers to, for example, an Application Specific Integrated Circuit (ASIC), an electronic circuit, a Programmable System-on-Chip (PSoC), a combinational logic circuit, and / or other suitable components that provide the described functionality. In one embodiment of the present disclosure, the one or more units 213 are communicatively coupled to the processor 203 for generating multimodal content. The one or more units 213 when configured with the functionality defined in the present disclosure results in a novel hardware.

[0077]

[0045] In one implementation, the one or more units 213 include, but are not limited to, the narrative generation agent 215, the script generation agent 217, the sub-section enhancing agent 219, the content generation agent 221, the converging agent 223, the display unit 105, the narrative text validation agent 225, the script validation agent 227, the enhanced sub-section validation agent 229, the audio-video validation agent 231, and the content validation agent 233. The one or more units 213 also, includes miscellaneous units 235 to perform various miscellaneous functionalities of the system 100.

[0078]

[0046] In an embodiment, the narrative generation agent 215, the script generation agent 217, the sub-section enhancing agent 219, the narrative text validation agent 225, the script validation agent 227 and the enhanced sub-section validation agent 229, each use a different instance of the one or more LLM’s 102 to generate their corresponding outputs.

[0079]

[0047] In an embodiment, the content generation agent 221, the converging agent 223, the audiovideo validation agent 231, and the content validation agent 233 each use a different instance of the multimodal model 103 to generate their corresponding outputs.

[0048] In an embodiment, the narrative generation agent 215 is configured to receive the input 101 to generate a multimodal content through the I / O interface 201 from the user. In an embodiment the input 101 is one of an action prompt to create the multimodal content based on existing multimodal content, a text based prompt comprising a brief of the multimodal content, or a prompt based on historical data of a user. In another embodiment, the input 101 may be received from a user. In an embodiment, the input 101 may be stored in the memory 205 as input 209.

[0080]

[0049] In an embodiment, the action prompt to create the multimodal content based on existing multimodal content may correspond to a click of a button to create a next episode based on a current episode of a multimodal content. In an embodiment, the text based prompt comprising a brief of the multimodal content may comprise of one of: one or more words or one or more sentences. An example of the text based prompt, without limitation to the following: “generate a story based on two characters”. In an embodiment, the text based prompt may comprise of one of a storyline comprising of one or more characters and a brief of the multimodal content expected. In an embodiment, the historical data comprises of user personalization data corresponding to a mobile application or a website or a website application. An example of the prompt based on historical data of the user, may be without limitation to: if the user personalization data of the user indicates that the user watches digital content based on action / thriller genre, the prompt may be: “generate a narrative text based on action / thriller genre”.

[0081]

[0050] In an embodiment, the narrative generation agent 215 is configured to generate the narrative text for the multimodal content using the first instance of the LLM 102a based on the input 101. In an embodiment, the narrative text may correspond to a story line including one or more characters and a plotline based on the input 101. In an embodiment, the narrative text generated may include the one or more characters disclosed in the text based prompt input 101 and may be based on the brief of the multimodal content expected, disclosed in the text based prompt input 101.

[0082]

[0051] In an embodiment, the narrative text validation agent 225 is configured to receive the narrative text from the narrative generation agent 215 and is configured to iteratively refine the narrative text using the fourth instance of the LLM 102d. The narrative text i s iteratively refined until the narrative text generated by the narrative generation agent 215 108 meets a threshold requirement based on a plurality of first performance metrics. The plurality of first performance metrics includes, without limitation to, inclusion of aspects from the text based prompt comprising the brief of the multimodal content, narrative continuity, omission of characters, efficiency of generating the narrative text, amount of narrative text generated etc. One or more improvements suggested by the narrative text validation agent 225 are communicated to the narrative generation agent 215 for regenerating the narrative text based on the improvements. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in the input 101 that may have been omitted by the narrative generation agent 215, removal of narrative discontinuity or further elaboration of the narrative text. The iterative process may be continued until both the narrative generation agent 215 and the narrative text validation agent 225 are satisfied with the narrative text. In an embodiment, the narrative text validation agent 225 may use prompt engineering to suggest improvements in the iterative refinement process.

[0083]

[0052] In an embodiment, the script generation agent 217 is configured to generate the script based on the validated narrative text generated by the narration generation agent 215, using the second instance of the LLM 102b. The script may comprise of one or more sections of the multimodal content and each section may comprise of one or more sub-sections. The script based on the validated narrative text may comprise of a detailed script. In some embodiments, the script may include details about the characters' movements, actions, expressions, and dialogue, as well as scene descriptions and changes. The script may also provide visual and audio details. For example the script may be related to scene settings, action, dialogues. In an embodiment, the script generation agent 217 may use prompt engineering to generate the script.

[0084]

[0053] In an embodiment, the script validation agent 227 is configured to validate the output of the script generation agent 217, i.e., the script. The script validation agent 227 is configured to receive from the script generation agent 217 the script and iteratively refine the script using the fifth instance of the LLM 102e. The script is iteratively refined until the script generated by the script generation agent 217 meets a threshold requirement based on a plurality of second performance metrics. The plurality of second performance metrics includes, without limitation to, sub-section continuity, omission of characters, efficiency, amount of sub-section generated. One or more improvements suggested by the script validation agent 227 are communicated to the script generation agent 217. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in the narrative text that may have been omitted by the script generation agent 217, removal of discontinuity between one or more section and each of their corresponding sub-sections or further elaboration of the one or more sections of the script. The iterative process may be continued until both the script generation agent 217 and the script validation agent 227 are satisfied with the script. In an embodiment, the script validation agent 227 may use prompt engineering to suggest improvements in the iterative refinement process.

[0085]

[0054] In an embodiment, the sub-section enhancing agent 219 is configured to enhance each of the validated one or more sub-sections generated by the script generation agent 217, using the third instance of the LLM 102c. In an embodiment, enhancing the one or more sub-sections may comprise of adding contextual information to each of the one or more sub-sections. In an embodiment, the contextual information is at least one of meta-data of a sub-section, location data, background data, character data, object class data (object class may correspond to without limitation to, animal, human, thing, vehicle etc), clothing of characters, appearance of characters etc. For example, a forest scene may be enhanced by including background sound that occur in a forest. In another example, a city scene may be enhanced by including street lights, post boxes, and the like.

[0086]

[0055] In an embodiment, the enhanced sub-section validation agent 229 is configured to validate the output of the sub-section enhancing agent 219 i.e., the enhanced one or more subsections. In an embodiment, the enhanced sub-section validation agent 229 is configured to receive from the sub-section enhancing agent 219, the one or more enhanced sub-sections and iteratively refining the one or more enhanced sub-sections using the sixth instance of the LLM 102f. The one or more enhanced sub-sections are iteratively refined until the enhanced one or more sub-sections generated by the sub-section enhancing agent 219 meets a threshold requirement based on a plurality of third performance metrics. The plurality of third performance metrics includes, without limitation to, narrative continuity, omission of characters, sufficiency of metadata, drastic change in metadata, etc. One or more improvements suggested by the enhanced sub-section validation agent 229are communicated to the subsection enhancing agent 219. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in one or more previous enhanced sub-sections that may have been omitted by the sub-section enhancing agent 219, removal of discontinuity between one or more sections and each of their corresponding enhanced sub-sections or further elaboration of metadata corresponding to each of the one or more enhanced sub-sections and removal of drastic changes in the metadata (for e.g., if there is a drastic change in the appearance a character between two consecutive sub-sections, corresponding to a section, removal of such a discrepancy may be necessitated). The iterative process may be continued until both the sub-section enhancing agent 219 and the enhanced sub-section validation agent 229are satisfied with the one or more enhanced sub-sections. In an embodiment, the enhanced sub-section validation agent 229may use prompt engineering to suggest improvements in the iterative refinement process.

[0087]

[0056] In an embodiment, the content generation agent 221 is configured to generate an audiovideo corresponding to each of the one or more enhanced sub-sections using the first instance of the multimodal model 103 a. In an embodiment, the

[0088]

[0057] In an embodiment, the audio-video validation agent 231 is configured to validate the output of the content generation agent 221 i.e., the audio-video corresponding to each of the one or more enhanced sub-sections. In an embodiment, the audio-video validation agent 231 is configured to receive from the content generation agent 221, the audio-video corresponding to each of the one or more enhanced sub-sections and iteratively refine the audio-video corresponding to each of the one or more enhanced sub-sections using the third instance of the multimodal model 103 c. The audio-video corresponding to each of the one or more enhanced sub-sections are iteratively refined until the audio-video corresponding to each of the one or more enhanced sub-sections generated by the content generation agent 221 meets a threshold requirement based on a plurality of fourth performance metrics. The plurality of fourth performance metrics includes, without limitation to, narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections. One or more improvements suggested by the audio-video validation agent 231 are communicated to the content generation agent 221. The one or more improvements / suggestions may include, without limitation to, inclusion of one or more characters present in one or more previous audio-video corresponding to each of one or more enhanced sub-sections that may have been omitted by the content generation agent 221, and removal of discontinuity between audio-video corresponding to each of one or more enhanced sub-sections. The iterative process may be continued until both the content generation agent 221 and the audio-video validation agent 231 are satisfied with the one or more enhanced subsections. In an embodiment, the audio-video validation agent 231 may use prompt engineering to suggest improvements in the iterative refinement process.

[0058] In an embodiment, the converging agent 223 is configured to converge the audio-video corresponding to each of the one or more enhanced sub-sections using the second instance of the multimodal model 103b, for generating the multimodal content. The converging of the audio-video corresponding to each of the one or more enhanced sub-sections ensures seamless transition between an audio-video corresponding to one of the enhanced sub-sections and an audio-video corresponding to its subsequent enhanced sub-section.

[0089]

[0059] In an embodiment, the content validation agent 233 is configured to validate the output of the converging agent 223 i.e., the multimodal content. In an embodiment, the content validation agent 233, configured to validate the multimodal content is configured to receive from the converging agent 223, the multimodal content and iteratively refine the multimodal content using the fourth instance of the multimodal model 103d. The multimodal content is iteratively refined until the multimodal content generated by the converging agent 223 meets a threshold requirement based on a plurality of fifth performance metrics. The plurality of fifth performance metrics includes, without limitation to, narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections. One or more improvements suggested by the content validation agent 233 are communicated to the converging agent 223. The one or more improvements / suggestions may include, without limitation to, removal of discontinuity between one or more audio-video corresponding to each of one or more enhanced sub-sections in the multimodal content and removal of seams between transition of one or more audio-videos corresponding to each of one or more enhanced sub-sections. The iterative process may be continued until both the converging agent 223 and the content validation agent 233 are satisfied with the multimodal content. In an embodiment, the content validation agent 233 may use prompt engineering to suggest improvements in the iterative refinement process.

[0090]

[0060] In an embodiment the display unit 105 is configured to display the multimodal content on an electronic device.

[0091]

[0061] In an embodiment, the multimodal content may be displayed to the user via an interface configured to enable development, upload and share the multimodal content generated based on the input 101. In another embodiment, the multimodal content may be displayed via one or more mobile applications or streaming platforms. In an embodiment, the multimodal content may be stored in a cloud server.

[0062] Fig. 3 illustrates a flowchart showing a method for generating multimodal content in accordance with some embodiments of the present disclosure.

[0092]

[0063] As illustrated in Fig. 3, the method 300 includes one or more operation steps for generating multimodal content in accordance with some embodiments of the present disclosure. The method 300 may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform particular functions or implement particular abstract data types.

[0093]

[0064] The order in which the method 300 is described is not intended to be construed as a limitation, and any number of the described method operation steps can be combined in any order to implement the method. Additionally, individual operation steps may be deleted from the methods without departing from the scope of the subject matter described herein. Furthermore, the method can be implemented in any suitable hardware, software, firmware, or combination thereof.

[0094]

[0065] At operation step 301, the processor 203 of the system 100 receives an input 101 to generate a multimodal content. In an embodiment the input 101 is one of an action prompt to create the multimodal content based on existing multimodal content, a text based prompt comprising a brief of the multimodal content, or a prompt based on historical data of a user. In another embodiment, the input 101 may be received from the user.

[0095]

[0066] In an embodiment, the action prompt to create the multimodal content based on existing multimodal content may correspond to a click of a button to create a next episode based on a current episode of a multimodal content. In an embodiment, the text based prompt comprising a brief of the multimodal content may comprise of one of: one or more words or one or more sentences. In an embodiment, the text based prompt may comprise of one of a storyline comprising of one or more characters and a brief of the multimodal content expected. In an embodiment, the historical data comprises of user personalization data corresponding to a mobile application or a website or a website application.

[0096]

[0067] At operation step 303, the processor 203 of the system 100 generates a narrative text for the multimodal content using the first instance of the LLM 102a based on the input 101. In an embodiment, the narrative text may correspond to a story line including one or more characters and a plotline based on the input 101. In an embodiment, the narrative text generated may include the one or more characters disclosed in the text based prompt input 101 and may be based on the brief of the multimodal content expected, disclosed in the text based prompt input 101.

[0097]

[0068] At operation step 305, the processor 203 of the system 100 generates a script based on the validated narrative text generated by the processor 203 of the system 100 at step 303, using the second instance of the LLM 102b. The script may comprise of one or more sections of the multimodal content and each section may comprise of one or more sub-sections. The script based on the validated narrative text may comprise of a detailed script. In an embodiment, the processor 203 may use prompt engineering to generate the script.

[0098]

[0069] At operation step 307, the processor 203 of the system 100, enhances each of the validated one or more sub-sections generated by processor 203 of the system 100 at step 305, using the third instance of the LLM. In an embodiment, enhancing the one or more sub-sections may comprise of adding contextual information to each of the one or more sub-sections. In an embodiment, the contextual information is at least one of meta-data of a sub-section, location data, background data, character data, object class data (object class may correspond to without limitation to, animal, human, thing, vehicle etc), clothing of characters, appearance of characters etc.

[0099]

[0070] At operation step 309, the processor 203 of the system 100 generates an audio-video corresponding to each of the one or more enhanced sub-sections using the first instance of the multimodal model 103 a.

[0100]

[0071] At operation step 311, the processor 203 of the system 100 converges the audio-video corresponding to each of the one or more enhanced sub-sections using the second instance of the multimodal model 103b, for generating the multimodal content. The converging of the audio-video corresponding to each of the one or more enhanced sub-sections ensures seamless transition between an audio-video corresponding to one of the enhanced sub-sections and an audio-video corresponding to its subsequent enhanced sub-section.

[0101]

[0072] At operation step 313, the processor 203 of the system 100, is configured to display the multimodal content on an electronic device.

[0073] In another embodiment Fig. 3 illustrates a flowchart showing a method for generating multimodal content in accordance with some embodiments of the present disclosure.

[0102]

[0074] As illustrated in Fig. 3, the method 300 includes one or more operation steps for generating multimodal content in accordance with some embodiments of the present disclosure.

[0103]

[0075] At operation step 301, the narrative generation agent 215 of the system 100 receives an input 101 to generate a multimodal content. In an embodiment the input 101 is one of an action prompt to create the multimodal content based on existing multimodal content, a text based prompt comprising a brief of the multimodal content, or a prompt based on historical data of a user. In another embodiment, the input 101 may be received from the user.

[0104]

[0076] In an embodiment, the action prompt to create the multimodal content based on existing multimodal content may correspond to a click of a button to create a next episode based on a current episode of a multimodal content. In an embodiment, the text based prompt comprising a brief of the multimodal content may compri se of one of: one or more words or one or more sentences. In an embodiment, the text based prompt may comprise of one of a storyline comprising of one or more characters and a brief of the multimodal content expected. In an embodiment, the historical data comprises of user personalization data corresponding to a mobile application or a website or a website application.

[0105]

[0077] At operation step 303, the narrative generation agent 215 of the system 100 generates a narrative text for the multimodal content using the first instance of the LLM 102a based on the input 101. In an embodiment, the narrative text may correspond to a story line including one or more characters and a plotline based on the input 101. In an embodiment, the narrative text generated may include the one or more characters disclosed in the text based prompt input 101 and may be based on the brief of the multimodal content expected, disclosed in the text based prompt input 101.

[0106]

[0078] At operation step 305, the script generation agent 217 of the system 100 generates a script based on the validated narrative text generated by the narration generation agent 215, using the second instance of the LLM 102b. The script may comprise of one or more sections of the multimodal content and each section may comprise of one or more sub-sections. The script based on the validated narrative text may comprise of a detailed script. In an embodiment, script generation agent 217 may use prompt engineering to generate the script.

[0107]

[0079] At operation step 307, the sub-section enhancing agent 219 of the system 100, enhances each of the validated one or more sub-sections generated by the script generation agent 217, using the third instance of the LLM 102c. In an embodiment, enhancing the one or more subsections may comprise of adding contextual information to each of the one or more subsections. In an embodiment, the contextual information is at least one of meta-data of a subsection, location data, background data, character data, object class data (object class may correspond to without limitation to, animal, human, thing, vehicle etc), clothing of characters, appearance of characters etc.

[0108]

[0080] At operation step 309, the content generation agent 221 of the system 100 generates an audio-video corresponding to each of the one or more enhanced sub-sections generated by the sub-section enhancing agent 219, using the first instance of the multimodal model 103 a.

[0109]

[0081] At operation step 311, the converging agent 223 of the system 100, converges the audiovideo corresponding to each of the one or more enhanced sub-sections generated by the content generation agent 221, using the second instance of the multimodal model 103b, for generating the multimodal content. The converging of the audio-video corresponding to each of the one or more enhanced sub-sections ensures seamless transition between an audio-video corresponding to one of the enhanced sub-sections and an audio-video corresponding to its subsequent enhanced sub-section.

[0110]

[0082] At operation step 313, the display unit 105 of the system 100, is configured to display the multimodal content generated by the converging agent 223, on an electronic device.

[0111]

[0083] In an embodiment, the multimodal content may be displayed to the user via an interface configured to enable development, upload and share the multimodal content generated based on the input 101. In another embodiment, the multimodal content may be displayed via one or more mobile applications or streaming platforms. In an embodiment, the multimodal content may be stored in a cloud server.

[0112]

[0084] In some embodiments, Fig. 4 illustrates a block diagram of an exemplary computer system 400 for implementing embodiments consistent with the present disclosure. In some embodiments, the computer system 400 may be the system 100 that comprises a processor (also referred as a processor 402 in this Fig. 4) that is used for generating multimodal content. In another embodiment, the computer system 400 may correspond to a server hosting one of: the interface configured to enable development, uploading and sharing of the multimodal content generated based on the input 101, a mobile application for streaming the multimodal content or streaming platform for streaming the multimodal content. The processor 402 may include at least one data processor for executing program components for executing user or systemgenerated business processes. The processor 402 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc.

[0113]

[0085] The processor 402 may be disposed in communication with input devices 410 and output devices 411 via I / O interface 401. The I / O interface 401 may employ communication protocols / methods such as, without limitation, audio, analog, digital, stereo, IEEE-1394, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), High-definition multimedia interface (HDMI), Radio Frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE 802. n / b / g / n / x, Bluetooth, cellular (e.g., Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System For Mobile Communications (GSM), Long-Term Evolution (LTE), WiMax, or the like), etc.

[0114]

[0086] Using the I / O interface 401, computer system 400 may communicate with input devices 410 (to receive the input 101) and output devices 411 (to display the multimodal conent).

[0115]

[0087] In some embodiments, the processor 402 may be disposed in communication with a communication network 409 via a network interface 403. The network interface 403 may communicate with the communication network 409. The network interface 403 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), Transmission Control Protocol / Internet Protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc. Using the network interface 403 and the communication network 409, the computer system 400 may communicate with an external server, a computer system or electronic device.

[0116]

[0088] The communication network 409 can be implemented as one of the different types of networks, such as intranet or Local Area Network (LAN) and such within the organization. The communication network 409 may either be a dedicated network or a shared network, which represents an association of the different types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Intemet Protocol (TCP / IP), Wireless Application Protocol (WAP), etc., to communicate with each other.

[0117]

[0089] Further, the communication network 409 may include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, etc. In some embodiments, the processor 402 may be disposed in communication with a memory 405 (e.g., Random Access Memory (RAM), ROM, etc. not shown in Fig. 4) via a storage interface 404. The storage interface 404 may connect to memory 405 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as Serial Advanced Technology Attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), fibre channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.

[0118]

[0090] The memory 405 may store a collection of program or database components, including, without limitation, a user interface 406, an operating system 407, a web browser 408, an application 412 etc. In some embodiments, the computer system 400 may store user / application data, such as the data, variables, records, etc. as described in this invention. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase.

[0119]

[0091] Operating system 407 may facilitate resource management and operation of computer system 400. Examples of operating systems include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (E.G, BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc ), LINUX® DISTRIBUTIONS (E.G, RED HAT®, UBUNTU®, KUBUNTU®, etc ), IBM®OS / 2®, MICROSOFT® WINDOWS® (XP®, VISTA® / 7 / 8, 10 etc ), APPLE® IOS®, GOOGLE™ ANDROID™, BLACKBERRY® OS, or the like. User interface 406 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, user interfaces may provide computer interaction interface elements on a display system operatively connected to computer system 400, such as cursors, icons, check boxes, menus, scrollers, windows, widgets, etc. Graphical User Interfaces (GUIs) may be employed, including, without limitation, Apple® Macintosh® operating systems’ Aqua®, IBM® OS / 2®, Microsoft® Windows® (e.g., Aero, Metro, etc.), web interface libraries (e.g., ActiveX®, Java®, Javascript®, AJAX, HTML, Adobe® Flash®, etc.), or the like.

[0120]

[0092] The computer system 400 may implement web browser 408 stored program components. Web browser 408 may be a hypertext viewing application, such as MICROSOFT® INTERNET EXPLORER®, GOOGLE™ CHROME™, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers 408 may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), etc. The computer system 400 may implement a mail server stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as ASP, ACTIVEX®, ANSI® C++ / C#, MICROSOFT®, NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system 400 may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc.

[0121]

[0093] The application 412 may be a software program that may be designed to perform a specific function or task directly for a user or for another software program. The application 412 may be one of a native applications, a hybrid application, a Progressive Web Application (PWA), an inventory management application, a mobile application, or a web application.

[0122]

[0094] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present invention. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non -transitory. Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, non-volatile memory, hard drives, Compact Disc (CD), Read-Only Memory (ROMs), Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media.

[0123]

[0095] Some of the advantages of the present disclosure are listed below.

[0124]

[0096] In existing video generation technologies are not are capable generating an indefinite, detailed digital multimodal content, including content based on a detailed narrative text including characters. The method and system of the present disclosure overcome these shortcomings by providing a method and system capable of generating indefinite, detailed digital multimodal content, including content based on detailed narrative text including characters.

[0125]

[0097] With respect to the use of substantially any plural and singular terms herein, those having skill in the art can translate from the plural to the singular and from the singular to the plural as is appropriate to the context or application. The various singular or plural permutations may be expressly set forth herein for sake of clarity.

[0126]

[0098] One or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which a software (program) readable by an information processing apparatus may be stored. The information processing apparatus includes a processor and a memory, and the processor executes a process of the software. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., be non-transitory. Examples include RAM, ROM, volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.

[0127]

[0099] The described operations may be implemented as a method, a system, or an article of manufacture using at least one of standard programming and engineering techniques to produce software, firmware, hardware, or any combination thereof. The described operations may be implemented as code maintained in a “non-transitory computer readable medium”, where a processor may read and execute the code from the computer readable medium. The processor is at least one of a microprocessor and a processor capable of processing and executing the queries. A non-transitory computer readable medium may include media such as magnetic storage medium (e.g., hard disk drives, floppy disks, tape, etc.), optical storage (CD ROMs, DVDs, optical disks, etc.), volatile and non-volatile memory devices (e.g., Programmable ROM (PROMs), Electrically Erasable PROM (EEPROMs), ROMs, RAMs, Dynamic RAM (DRAMs), Static RAM (SRAMs), Flash Memory, firmware, programmable logic, etc.), etc. Further, non-transitory computer-readable media include all computer-readable media except for a transitory. The code implementing the described operations may further be implemented in hardware logic (e.g., an integrated circuit chip, Programmable Gate Array (PGA), ASIC, etc.).

[0128]

[0100] The terms “an embodiment”, “embodiment”, “embodiments”, “the embodiment”, “the embodiments”, “one or more embodiments”, “some embodiments”, and “one embodiment” mean “one or more (but not all) embodiments of the invention(s)” unless expressly specified otherwise.

[0129]

[0101] The terms “including”, “comprising”, “having” and variations thereof mean “including but not limited to”, unless expressly specified otherwise.

[0130]

[0102] The enumerated listing of items does not imply that any or all of the items are mutually exclusive, unless expressly specified otherwise.

[0131]

[0103] The terms “a”, “an” and “the” mean “one or more”, unless expressly specified otherwise.

[0132]

[0104] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention.

[0133]

[0105] When a single device or article is described herein, it will be readily apparent that more than one device or article (whether or not they cooperate) may be used in place of a single device or article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be readily apparent that a single device or article may be used in place of the more than one device, or article, or a different number of devices or articles may be used instead of the shown number of devices or programs. At least one of the functionalities and the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality or features. Thus, other embodiments of the invention need not include the device itself.

[0134]

[0106] The illustrated operations of Fig. 3 show certain events occurring in a certain order. In alternative embodiments, certain operations may be performed in a different order, modified, or removed. Moreover, steps may be added to the above-described logic and still conform to the described embodiments. Further, operations described herein may occur sequentially or certain operations may be processed in parallel. Yet further, operations may be performed by a single processing unit or by distributed processing units.

[0135]

[0107] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.

[0136]

[0108] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

CLAIMS:

1. A method for generating multimodal content, comprising: receiving, by a narrative generation agent 215, an input to generate a multimodal content; generating, by the narrative generation agent 215, a narrative text for the multimodal content using a first instance of a Large Language Models (LLM) based on the input; generating, by a script generation agent 217, a script comprising one or more sections of the multimodal content, wherein each section comprises one or more subsections, based on the narrative text generated by the narration generation agent, using a second instance of the LLM; enhancing, by a sub-section enhancing agent 219, each of the one or more subsections using a third instance of the LLM, wherein enhancing comprises of adding contextual information to each of the one or more sub-sections; generating, by a content generation agent 221, an audio-video corresponding to each of the one or more enhanced sub-sections using a first instance of a multimodal model; converging, by a converging agent 223, the audio-video corresponding to each of the one or more enhanced sub-sections for generating the multimodal content, using a second instance of the multimodal model; and displaying, by a display unit 105, the multimodal content on an electronic device.

2. The method as claimed in claim 1, comprises validating the narrative text, wherein validating the narrative text comprises of: receiving, by a narrative text validation agent 225 from the narrative generation agent 215, the narrative text; and iteratively refining, by the narrative text validation agent 225, the narrative text until the narrative text generated by the narrative generation agent 215 meets a threshold requirement based on a plurality of first performance metrics using a fourth instance of the LLM, wherein the plurality of first performance metrics includes: narrative continuity, omission of characters, efficiency, amount of narrative text generated.

3. The method as claimed in claim 1, comprises validating the script, wherein validating the script comprises of: receiving, by a script validation agent 227 from the script generation agent217, the script; and iteratively refining, by the script validation agent 227, the script until the script generated by the script generation agent 217 meets a threshold requirement based on a plurality of second performance metrics using a fifth instance of the LLM, wherein the plurality of second performance metrics includes: sub-section continuity, omission of characters, efficiency, amount of sub-section generated.

4. The method as claimed in claim 1, comprises validating the one or more enhanced subsections, wherein validating the one or more enhanced sub-sections comprises of: receiving, by an enhanced sub-section validation agent 229 from the subsection enhancing agent 219, the one or more enhanced sub-sections; and iteratively refining, by the enhanced sub-section validation agent 229, the one or more enhanced sub-sections until the one or more enhanced sub-sections generated by the sub-section enhancing agent 219 meets a threshold requirement based on a plurality of third performance metrics using a sixth instance of the LLM, wherein the plurality of third performance metrics includes: narrative continuity, omission of characters, sufficiency of metadata, drastic change in metadata.

5. The method as claimed in claim 1, comprises validating the audio-video corresponding to each of the one or more enhanced sub-sections, wherein validating the audio-video corresponding to each of the one or more enhanced sub-sections comprises of: receiving, by an audio-video validation agent 231 from the content generation agent 221, the audio-video corresponding to each of the one or more enhanced subsections; and iteratively refining, by the audio-video validation agent 231, the audio-video corresponding to each of the one or more enhanced sub-sections until the audio-video corresponding to each of the one or more enhanced sub-sections generated by the content generation agent 221 meets a threshold requirement based on a plurality of fourth performance metrics using a third instance of the multimodal model, wherein the plurality of fourth performance metrics includes: narrative continuity, omission ofcharacters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections.

6. The method as claimed in claim 1, comprises validating the multimodal content, wherein validating the multimodal content comprises of: receiving, by a content validation agent 233 from the converging agent 223, the multimodal content; and iteratively refining, by the content validation agent 233, the multimodal content, until the multimodal content generated by the converging agent 223 meets a threshold requirement based on a plurality of fifth performance metrics using a fourth instance of the multimodal model, wherein the plurality of fifth performance metrics includes: narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections.

7. The method as claimed in claim 1, wherein the input is one of: an action prompt to create the multimodal content based on existing multimodal content, a text based prompt comprising a brief of the multimodal content, or a prompt based on historical data of a user, wherein historical data comprises of user personalization data.

8. The method as claimed in claim 1, wherein the contextual information is at least one of: meta-data of a sub-section, location data, background data, character data, object class data, clothing of characters, appearance of characters.

9. A system for generating multimodal content, comprising of a narrative generation agent 215, a script generation agent 217, a sub-section enhancing agent 219, a content generation agent 221, a converging agent 223, a display unit 105, a narrative text validation agent 225, a script validation agent 227, an enhanced sub-section validation agent 229, an audio-video validation agent 231, a content validation agent 233, a processor; and a memory, wherein the memory stores processor-executable instructions, which, on execution, cause the processor to: receive an input to generate multimodal content; generate a narrative text for the multimodal content using a first instance of a Large Language Models (LLM) based on the input;generate a script comprising one or more sections of the multimodal content, wherein each section comprises one or more sub-sections, based on the narrative text using a second instance of the LLM; enhance each of the one or more sub-sections using a third instance of the LLM, wherein enhancing comprises of adding contextual information to each of the one or more sub-sections; generate an audio-video corresponding to each of the one or more enhanced sub-sections using a first instance of a multimodal model; converge the audio-video corresponding to each of the one or more enhanced sub-sections for generating the multimodal content, using a second instance of the multimodal model; and display the multimodal content on an electronic device.

10. The system as claimed in claim 9, comprises validating by the processor the narrative text, wherein the processor configured to validate the narrative text is configured to: receive the narrative text; and iteratively refine the narrative text until the narrative text generated by the processor meets a threshold requirement based on a plurality of first performance metrics using a fourth instance of the LLM, wherein the plurality of first performance metrics includes: narrative continuity, omission of characters, efficiency, amount of narrative text generated.

11. The system as claimed in claim 9, comprises validating by the processor the script, wherein the processor configured to validate the script is configured to: receive the script; and iteratively refine the script until the script generated by the processor meets a threshold requirement based on a plurality of second performance metrics using a fifth instance of the LLM, wherein the plurality of second performance metrics includes: sub-section continuity, omissi on of characters, efficiency, amount of sub-section generated.

12. The system as claimed in claim 9, comprises validating by the processor the one or more enhanced sub-sections, wherein the processor configured to validate the one or more enhanced sub-sections is configured to:receive the one or more enhanced sub-sections; and iteratively refine the one or more enhanced sub-sections until the one or more enhanced sub-sections generated by the processor meets a threshold requirement based on a plurality of third performance metrics using a sixth instance of the LLM, wherein the plurality of third performance metrics includes: narrative continuity, omission of characters, sufficiency of metadata, drastic change in metadata.

13. The system as claimed in claim 9, comprises validating by the processor the audiovideo corresponding to each of the one or more enhanced sub-sections, wherein the processor configured to validate the audio-video corresponding to each of the one or more enhanced sub-sections is configured to: receive the audio-video corresponding to each of the one or more enhanced sub-sections; and iteratively refine the audio-video corresponding to each of the one or more enhanced sub-sections until the audio-video corresponding to each of the one or more enhanced sub-sections generated by the processor meets a threshold requirement based on a plurality of fourth performance metrics using a third instance of the multimodal model, wherein the plurality of fourth performance metrics includes: narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections.

14. The system as claimed in claim 9, comprises validating by the processor the multimodal content, wherein the processor configured to validate the multimodal content is configured to: receive the multimodal content; and iteratively refine the content until the multimodal content generated by the processor, meets a threshold requirement based on a plurality of fifth performance metrics using a fourth instance of the multimodal model, wherein the plurality of fifth performance metrics includes: narrative continuity, omission of characters, discrepancies between the audio-video corresponding to each of one or more enhanced sub-sections.

15. The system as claimed in claim 9, wherein the input is one of: an action prompt to create the multimodal content based on existing multimodal content, a text based promptcomprising a brief of the multimodal content, or a prompt based on historical data of a user, wherein historical data comprises of user personalization data.

16. The system as claimed in claim 9, wherein the contextual information is at least one of: meta-data of a sub-section, location data, background data, character data, object class, clothing of characters, appearance of characters.

Citation Information

Patent Citations

  • Video content automatic generation method and system

    CN116484048A

  • Generative artificial intelligence-based novel tweet video generation method and system

    CN117078782A

  • Method and system for automatically generating story video in meta universe

    CN117177003A