Idiom story video generation method and device, equipment and medium

Idiom story scripts are generated through natural language processing technology, storyboard scripts and visual elements are automatically generated, and time alignment and optimization are performed in combination with voice narration and background sound effects. This solves the problems of module fragmentation and insufficient cultural adaptability in idiom story video generation, realizes personalized video generation and optimization, and improves the overall coordination and expressiveness of the video.

CN120658923APending Publication Date: 2025-09-16SHENZHEN KINGSUN SCIENCE & TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510809715.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing technology for generating idiom story videos has problems such as module fragmentation, insufficient cultural adaptability, limited dynamic expression, and lack of personalization, resulting in a lack of overall coordination in the video content in terms of story logic, picture style, and voice intonation. It is unable to vividly present the dramatic effect of the idiom story and lacks user interaction functions.

Method used

Generate idiom story scripts through natural language processing technology, automatically generate storyboards and visual elements, combine voice narration and background sound effects, perform timing alignment and optimization, and adjust video style and format according to user needs to provide personalized story presentation.

Benefits of technology

It achieves the overall coordination and coherence of idiom story videos, improves the expressiveness of character actions and scene switching, ensures the overall coordination of content, pictures, sounds and interactivity, and supports personalized video generation and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658923A_ABST
    Figure CN120658923A_ABST
Patent Text Reader

Abstract

The invention provides an idiom story video generation method and device, equipment and a medium, and the method comprises the following steps: generating an idiom story script through a natural language processing technology based on an input idiom; automatically generating a split script and a corresponding visual element according to the story script; based on the story script, generating a voice paraphrasing and background sound effect; generating a dynamic video and performing time sequence alignment and optimization by combining the visual element, the voice frame and the background sound effect; and adjusting the story style, the visual style and the output format of the dynamic video according to user requirements, and generating an idiom story video. According to the method, the problems that in the prior art, all modules are split, culture adaptability is insufficient, dynamic expressive force is limited, and individuation is lacked are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital content creation, and more specifically to a method, device, equipment and medium for generating idiom story videos. Background Art

[0002] At present, there are still many technical bottlenecks in the field of artificial intelligence in the generation of idiom story videos. First, existing technologies mostly adopt modular independent processing methods, and the links such as text generation, image generation and speech synthesis are separated from each other, resulting in the lack of overall coordination of the final output video content in terms of story logic, picture style and voice intonation. Secondly, the general artificial intelligence model lacks special optimization for the characteristics of Chinese idioms, and the generated storylines often deviate from the original meaning of the idioms, making it difficult to accurately convey the connotation of traditional culture. Furthermore, video generation technology mostly adopts the method of converting static images into videos, which makes the character movements stiff and the scene transitions abrupt, and cannot vividly show the dramatic effect of idiom stories. In addition, existing systems generally lack user interaction functions and cannot adjust the content difficulty and presentation form according to different application scenarios (such as children's education or cultural science popularization). These problems seriously restrict the application effect of artificial intelligence in the field of traditional cultural communication. Summary of the Invention

[0003] The purpose of the present invention is to overcome the defects of the prior art and provide a method, device, equipment and medium for generating idiom story videos, which aims to solve the problems of module fragmentation, insufficient cultural adaptability, limited dynamic expression and lack of personalization in the prior art.

[0004] To achieve the above object, the present invention adopts the following technical solutions:

[0005] In a first aspect, an embodiment of the present invention provides a method for generating an idiom story video, comprising the following steps:

[0006] Based on the input idioms, idiom story scripts are generated through natural language processing technology;

[0007] Automatically generate storyboards and corresponding visual elements based on the story script;

[0008] Based on the story script, generate voice narration and background sound effects;

[0009] Combining the visual elements, the voice narration, and the background sound effects to generate a dynamic video and perform timing alignment and optimization;

[0010] The story style, visual style and output format of the dynamic video are adjusted according to user needs to generate an idiom story video.

[0011] In a second aspect, an embodiment of the present invention provides an idiom story video generation device, comprising:

[0012] A story script generation module is used to generate idiom story scripts based on input idioms through natural language processing technology;

[0013] A storyboard and visual generation module, configured to automatically generate a storyboard script and corresponding visual elements based on the story script;

[0014] A speech synthesis module, for generating voice narration and background sound effects based on the story script;

[0015] A video synthesis module, for combining the visual elements, voice narration, and background sound effects to generate dynamic video and perform timing alignment and optimization;

[0016] The user customization module is used to adjust the dynamic video according to user needs and generate the final idiom story video.

[0017] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned method for generating idiom story videos when executing the computer program.

[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for generating an idiom story video as described above is implemented.

[0019] Compared with the existing technology, the beneficial effects of the present invention are as follows: idiom story scripts are generated through natural language processing technology, which solves the problem of content fragmentation and ensures the integrity and coherence of the story content; storyboards and visual elements are automatically generated based on the story script to achieve the coordination and unity of the picture style and story content; matching voice narration and background sound effects are generated according to the story script to ensure the consistency of the voice intonation and the story tone; visual elements, voice and sound effects are integrated into dynamic videos through timing alignment optimization to enhance the expressiveness of character actions and scene switching; finally, the video style and output format are adjusted according to user needs to provide a personalized story presentation method. This method realizes an end-to-end automated process from idiom input to video output, ensuring the overall coordination of the generated idiom story video in terms of content, picture, sound and interactivity.

[0020] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A schematic diagram of a flow chart of a method for generating an idiom story video provided by an embodiment of the present invention;

[0022] Figure 2 A schematic block diagram of a device for generating an idiom story video provided by an embodiment of the present invention;

[0023] Figure 3 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0028] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0029] See also Figure 1 As shown, the embodiment of the present invention discloses a method for generating an idiom story video, comprising the following steps:

[0030] S100, based on the input idiom, generating an idiom story script through natural language processing technology;

[0031] Specifically, natural language processing technology is used to perform semantic analysis and cultural connotation mining on input idioms, constructing a story framework that aligns with the idiom's context, addressing the inefficiency and semantic deviations of traditional manual scripting. Specifically, the system first receives the idiom input from the user and simultaneously acquires optional parameters such as the target audience and story style to clarify the direction of story generation. Next, using a pre-trained language model combined with a built-in idiom knowledge base, preliminary story content, including characters, settings, and plot, is generated. To ensure that the generated content deviates from the original meaning of the idiom, a rules engine is used to logically verify the preliminary content. For example, this checks whether the plot conforms to the basic logic of the idiom's context and whether the character's behavior is consistent with the idiom's connotation. The final output is a structured story script, providing an accurate textual foundation for subsequent storyboard generation, visual design, and other steps. This design combines AI-powered automated generation with rule-based verification to improve script creation efficiency while ensuring cultural accuracy.

[0032] It can be understood that a pre-trained language model refers to a neural network model that has been pre-trained on large-scale text data. It has the ability to understand language semantics and generate coherent text, and can be adapted to the idiom story generation task through fine-tuning; the idiom knowledge base is a database that stores idiom-related knowledge in a structured manner, including the source allusions, core meanings, common role settings, etc. of the idioms; the rule engine is a module that verifies the generated content based on preset logical rules (such as the meaning of the idiom must be reflected, the character behavior must be consistent with the background of the times, etc.), which can be implemented through regular expressions, decision trees and other technologies.

[0033] Taking the idiom "waiting for a rabbit by a tree" as an example, after the user enters the idiom and selects the "primary school students" target audience and the "fun version" style, the system invokes the language model and, combined with the knowledge base information about the idiom "a farmer waiting for a rabbit by a tree, satirizing the mentality of taking chances," generates preliminary story content: "A farmer was working in the fields when a rabbit crashed into a tree and died. The farmer then waited under the tree every day for the rabbit, eventually leaving the fields barren." The rule engine verifies that the content includes the core characters of "farmer" and "rabbit," the key scenes of "farmland" and "tree," and whether it embodies the moral of "unearned gain is undesirable." If verification is successful, a script containing the characters, scenes, and moral is output.

[0034] S200, automatically generating a storyboard script and corresponding visual elements according to the story script;

[0035] Specifically, the text-based story script is converted into a visual storyboard script and visual materials to construct the narrative framework and visual foundation of the video, solving the problem that storyboard design in traditional video production relies on manual experience and the creation of visual elements is inefficient. Specifically, the system needs to first perform semantic analysis on the storyboard, extract key plots, and split it into multiple storyboards according to narrative logic. Each storyboard contains elements such as shot number, scene description, and character action. Subsequently, a screen description prompt is generated for each storyboard, and visual elements such as character image and scene background are rendered through an image generation model (such as Stable Diffusion), and dynamic actions are added to the characters in combination with animation templates. This design greatly improves the efficiency of storyboard design and visual generation through artificial intelligence automation technology, and the generated storyboard sequence is more consistent with the story logic, ensuring the smoothness and visual expressiveness of the video narrative.

[0036] It can be understood that the storyboard script is the blueprint for video production, which contains information such as the shot size (long shot, close shot), shooting angle (straight shot, overhead shot), character actions, lines, etc., which is used to guide the generation of visual elements and video editing; visual element generation technology combines natural language processing and computer vision, and controls the image generation model through text prompts to output characters, scenes, special effects and other materials that meet the storyboard requirements; animation templates are predefined character action libraries that can quickly give static images dynamic effects.

[0037] S300, generating voice narration and background sound effects based on the story script;

[0038] Specifically, through speech synthesis and sound effect matching technology, the story script is given audio expressiveness, solving the problems of high audio production costs and poor emotional adaptability in traditional video production. The system must first extract the narration text and dialogue content from the story script, identify the storyboard scene corresponding to each text (such as the farmland scene corresponding to "The farmer waits for the rabbit by the tree"), and then select the parameters of the speech synthesis model (such as tone and speed) according to the emotional tone of the scene (such as irony, cheerfulness) to generate emotional voice narration. At the same time, based on the scene type (such as nature, life scenes), match the environmental sound effects (such as birdsong, wind) from the sound effect library or generate them through artificial intelligence. Finally, adjust the audio intensity according to the storyboard time node to achieve audio and picture synchronization. This design shortens the time taken for traditional manual dubbing through automated audio generation, and improves the emotional consistency between audio and picture, enhancing the immersiveness of the video.

[0039] It is understandable that the speech synthesis model is an artificial intelligence tool based on deep learning, which can convert text into natural speech and support the adjustment of intonation, speaking speed, and emotional parameters; the sound effect library is a database that stores various types of environmental sounds and action sounds, and is classified and managed by scene type (such as nature, city) and emotional type (such as cheerful, horror); artificial intelligence sound effect generation technology can generate customized sound effects through text prompts.

[0040] S400, combining the visual elements, the voice narration, and the background sound effects to generate a dynamic video and perform timing alignment and optimization;

[0041] Specifically, the multimodal materials are integrated into a coherent video, and the viewing experience is enhanced by time alignment and special effects addition, solving the problems of cumbersome video synthesis processes and unstable quality in existing technologies. The system first arranges the dynamic visual elements into an initial sequence according to the order of the storyboards, adjusts the playback speed of the voice narration according to the length of the storyboards, and uses a dynamic time warping algorithm to achieve alignment of the audio and video time axes; then inserts transition effects (such as fade in and fade out, wipe) at the storyboard switching point, generates subtitles based on the voice content and superimposes them, and finally mixes the background sound effects to output the video. For example, when synthesizing the video of "Foolish Old Man Moves the Mountains", it is necessary to ensure that the visual action of "Foolish Old Man Swinging the hoe" is synchronized with the voice of "Digging the Mountain". When switching to the "Wise Old Man Laughing" storyboard, a "wipe" transition is added to enhance the narrative rhythm.

[0042] It is understandable that the dynamic time warping algorithm can achieve synchronization by stretching / compressing the voice frames when the voice duration is inconsistent with the storyboard duration; the transition special effects library includes a variety of transition effects such as fade in and fade out, slide, dissolve, etc., which can be selected according to the emotion of the scene, such as "black transition" for serious scenes and "circular wipe" for fun scenes; subtitle generation is based on speech recognition technology, which converts the narration text into a subtitle file synchronized with the timeline.

[0043] S500: Adjust the story style, visual style and output format of the dynamic video according to user needs to generate an idiom story video.

[0044] Specifically, this step is used to meet the personalized needs of users and solve the problem of the single style of traditional video generation. The system provides an interactive interface to support users to select the complexity of the story (such as children's version, academic version), visual style (cartoon style, ink style), and output format (horizontal screen, vertical screen). According to the selection, the adapted story script or visual elements are regenerated and the video is resynthesized. For example, if the user selects "academic version + ink style + horizontal screen", the system will add allusion source analysis to the story script, call the ink style image generation model, and output a 16:9 format video.

[0045] It is understandable that story style conversion is achieved through parameter adjustment of the large language model. For example, the children's version uses simple vocabulary + repeated sentences, and the academic version uses professional terminology + in-depth analysis; visual style conversion is controlled by the style parameters of the image generation model; output format conversion involves adjustments to resolution, frame rate, and encoding methods, such as the vertical screen resolution is 1080×1920 and the horizontal screen resolution is 1920×1080.

[0046] Furthermore, the steps of adjusting the story style, visual style and output format of the dynamic video according to user needs and generating the idiom story video also include:

[0047] S600: Performing quality verification on the generated idiom story video content and outputting the verification result;

[0048] Specifically, automated quality testing is used to ensure that the video meets technical standards and solve the problem of high missed detection rate in manual review. The system verifies from the following dimensions: audio and video synchronization, visual quality, audio quality, and content compliance. After verification, a report is generated, listing the problem items and locations. It is understandable that the quality verification algorithm includes: audio and video synchronization detection: calculating the time difference between voice and video frames; visual quality assessment: PSNR (peak signal-to-noise ratio), SSIM (structural similarity) indicators; audio quality assessment: volume standardization, distortion detection; content compliance: keyword filtering (such as sensitive words), idiom meaning matching (compared with the knowledge base).

[0049] S610. Automatically optimize the idiom story video content according to the verification result and output the final version of the video.

[0050] Specifically, it automatically fixes quality issues, improves video yield, and reduces manual intervention. The system performs corresponding optimization based on the verification report:

[0051] Audio and video out of sync: Re-run the DTW algorithm for alignment; Low resolution: Call an image super-resolution model (such as Real-ESRGAN) to improve clarity; Audio issues: Adjust the volume ratio or regenerate the voice; Content issues: Trigger the script regeneration process. After optimization, recheck until all indicators meet the standards. It is understandable that automatic optimization tools include: Audio and video synchronization repair (based on deep learning timing alignment model); Image super-resolution (Real-ESRGAN, Topaz Video Enhance, etc.); Audio repair (Adobe Audition's automatic noise reduction and volume balancing functions); Content repair (script error correction function of large language models).

[0052] In one embodiment, the step of generating an idiom story script based on the input idiom by using natural language processing technology includes:

[0053] Receive an idiom and optional parameters input by a user, wherein the optional parameters include a target audience and a story style;

[0054] Specifically, this step obtains the core elements of user needs and provides a clear guide for the personalized generation of story scripts, thereby solving the problem of repeated revisions caused by vague needs in traditional content generation. The system needs to provide an idiom input box and parameter selection items in the interactive interface, where the target audience parameters are used to determine the complexity of the story language and the depth of the plot, and the story style parameters are used to limit the narrative tone and the degree of cultural expansion. For example, when the user selects the "children" audience, the language of the story subsequently generated by the system needs to be simplified and the plot needs to be intuitive; when the "academic version" style is selected, it is necessary to add content such as textual research of allusions and in-depth analysis of the meaning. This design enables the system to generate adapted story scripts according to different application scenarios, improve the practicality and pertinence of the content, and avoid the lack of content adaptability caused by "one-size-fits-all" generation.

[0055] It is understandable that the target audience refers to the expected recipients of the story. Audiences of different age groups or knowledge levels have different abilities to understand and interests in the story. For example, children prefer simple plots and vivid characters, while adults can accept more complex moral analysis. Story style refers to the narrative characteristics and expression forms of the story, including language style (popular, elegant), content depth (basic explanation, academic analysis), emotional tone (fun, serious), etc.

[0056] Generate preliminary story content through pre-trained language model combined with idiom knowledge base;

[0057] Specifically, this step utilizes artificial intelligence technology to automate the creation of story content, addressing the time-consuming and labor-intensive nature of manual editing and the lack of consistent content. The pre-trained language model possesses powerful language understanding and generation capabilities. By accessing an idiom knowledge base, it can generate a story framework that meets the requirements based on the input idioms and parameters. This design, through the collaboration between the language model and the knowledge base, ensures that AI-generated stories not only adhere to the cultural connotations of idioms but also adjust the content depth based on parameters, achieving a "batch generation + personalized customization" creative model.

[0058] It is understandable that the pre-trained language model is based on the Transformer architecture, learns language rules through massive text data, and has the ability to understand context and generate semantics; the idiom knowledge base is constructed using knowledge graph technology, which contains triple data of idioms (such as "waiting for the rabbit by the tree - allusion - the farmer waiting by the tree"), attribute features (historical background, meaning labels), etc., which can guide the language model to generate content that meets specific requirements.

[0059] A rule engine is used to verify whether the content of the preliminary story conforms to the original meaning of the idiom and its logical rationality;

[0060] Specifically, AI-generated stories are quality-controlled using pre-set logical rules to avoid deviations from the cultural connotations of idioms and illogical plots. Furthermore, the rule engine must include two types of rules: semantic verification, such as checking whether the story captures the core meaning of the idiom and whether character behavior aligns with the setting of the allusion; and logical verification, such as checking whether the plot development adheres to chronological order and whether causal relationships are reasonable. During verification, the engine automatically analyzes the text of the preliminary story, extracting key characters, scenes, and plot points, and compares them with the core idiom elements in the knowledge base. If the story "Foolish Old Man Moving the Mountains" lacks the ending "The Emperor is moved by his sincerity," the story is deemed inconsistent with the original intent. Similarly, if the story "Carving a Boat to Seek a Sword" finds a logical inconsistency, such as the Chu people searching for their sword after the boat docks, the story is deemed logically unreasonable. This automated verification mechanism ensures that AI-generated stories meet application standards for cultural accuracy and logical rigor, reducing manual review costs.

[0061] It can be understood that the rule engine is a rule-based system consisting of a rule base (storing idiom verification rules), an inference mechanism (analyzing whether the text complies with the rules), and a fact base (storing story content data). Verification can be achieved through regular expression matching of keywords, semantic network analysis of plot logic, etc.

[0062] In one embodiment, the step of automatically generating a storyboard script and corresponding visual elements based on the story script includes:

[0063] Parse the story script, extract key plots and split them into multiple storyboards;

[0064] Specifically, this step transforms the textual story into a visual lens language, providing structured guidance for subsequent visual element generation. This addresses the inefficient reliance on manual experience in traditional video production for storyboard design. The system performs semantic analysis on the storyboard, identifying the timeline and logical nodes of event development, extracting key plot points (such as "farmer tilling the fields," "rabbit hitting a tree," and "waiting for a rabbit by a tree"), and then breaks them down into separate storyboards based on the narrative rhythm. Each storyboard must include shot number, scene description, character actions, and plot points. For example, "Shot 1: Daytime view of a farmland, farmer hoeing, demonstrating his diligence; Shot 2: Close-up of a tree, rabbit crashing into the tree at high speed, highlighting the unexpected event." This design transforms the abstract textual story into a concrete visual sequence, clarifying the specific requirements of visual generation. It also lays the foundation for the subsequent timing of video synthesis, ensuring the narrative fluidity and logic of the final video.

[0065] It can be understood that storyboards are the basic units in video production, which contain information such as the shot size (long shot, close shot), shooting angle (straight shot, overhead shot), character actions, scene atmosphere, etc., which are used to guide the generation of visual elements and video editing; key plots are the core events in the story that drive the development of the narrative, such as the generation of contradictions, climax, ending and other nodes.

[0066] Generate corresponding screen description prompts for each storyboard;

[0067] Specifically, the abstract description of the storyboard is converted into text instructions that the image generation model can understand, ensuring that the generated visual elements closely match the storyboard requirements and addressing the semantic ambiguity and detail deviations that often occur when AI generates images. The prompts must encompass the core elements of the storyboard: character traits (e.g., "a farmer in coarse linen clothing"), scene setting (e.g., "a golden wheat field with distant green hills"), action poses (e.g., "a farmer wielding a hoe to till the field"), and emotional atmosphere (e.g., "a sunny harvest scene"). For example, the prompt generated for the "Farmer Waiting for Rabbits" storyboard requires clarity: "A farmer, dressed in Song Dynasty clothing, squats under a tree, his eyes gazing blankly into the distance. A hoe rests beside him. The farmland is overgrown with weeds, and the image is grayish, reflecting the abandoned state of farming." This design uses structured prompts to guide the image generation model to accurately render visual content, ensuring that the generated character imagery, scene layout, and atmosphere meet the storyboard's narrative needs, reducing image rework caused by unclear prompts.

[0068] Based on the screen description prompt words, using an image generation model to generate a character image and a scene image;

[0069] Specifically, AI-powered image generation technology converts text prompts into visual assets, addressing the time-consuming and labor-intensive nature of traditional hand-drawn characters and scenes, as well as the difficulty of mass production. The system feeds the storyboard's description prompts into an image generation model (such as StableDiffusion or MidJourney). The model analyzes the semantic information (such as character identity and scene type) and visual features (such as color and composition) contained in the prompts to generate corresponding character images (such as a farmer's appearance and clothing) and scene images (such as farmland and trees). For example, given the prompt "a cartoon-style farmer with a round head and a simple smile, dressed in blue cloth, holding a wooden hoe, standing in a green field," the image generation model can generate characters and scenes that fit the childlike style. To improve the accuracy of the generated images, control models such as ControlNet can be incorporated to constrain the generation process (such as specifying character poses and scene perspective). This design automates the generation of visual assets, significantly improving efficiency compared to traditional hand-drawing, and supports the generation of diverse visual content based on different stylistic parameters (such as Chinese style and ink painting).

[0070] It is understandable that the image generation model is an artificial intelligence system based on deep learning. By learning massive image-text pairs, it has the ability to generate images based on text descriptions; ControlNet is a conditional control model that can add additional constraints (such as line drawings and depth maps) in the image generation process to make the generation results more in line with expectations.

[0071] Combine predefined animation templates to add dynamic actions to characters and form dynamic visual elements.

[0072] Specifically, this step gives the static character image dynamic expressiveness, solving the problem of stiff and lack of smoothness in the conversion of static images into videos in the prior art. The system needs to establish a predefined animation template library, which includes bone binding and key frame settings for common actions (such as "walking", "waving", and "squatting"). For example, the farmer's "farming" action template in the Spine animation tool includes key frames such as swinging a hoe, bending over, and standing. For each character in the storyboard, a matching animation template is selected, and the generated static character image is mapped to the skeleton structure of the template. After adjusting the details (such as the swing amplitude of clothing and changes in facial expressions), a dynamic action sequence is generated. For example, for the "farmer swinging a hoe and farming" storyboard, the "farming" template is called to make the static farmer image produce a cyclic action of swinging a hoe and bending over, and at the same time add a particle effect of the hoe sinking into the soil. This design greatly improves the production efficiency of dynamic visual elements through templated action generation, while ensuring natural and smooth movements, making the character performance more vivid and enhancing the narrative tension of the video.

[0073] It can be understood that predefined animation templates are pre-designed standardized action sequences, including character skeleton binding, keyframe animation, physical effects, etc., which can be repeatedly applied to similar action scenes; skeleton binding is to associate 2D images with virtual bone structures, and realize dynamic deformation of images by controlling bone movement, which is commonly found in 2D animation tools such as Spine and Rive.

[0074] In one embodiment, the step of generating voice narration and background sound effects based on the story script includes:

[0075] Extracting the narration text and dialogue content in the story script and identifying the storyboard scenes corresponding to each paragraph of text;

[0076] Specifically, this step establishes a mapping relationship between text and visual scenes, provides an accurate basis for subsequent voice emotion adjustment and sound effect matching, and solves the problem of audio and picture disconnection. The system parses the story script through natural language processing technology, extracts narration text (such as descriptive sentences) and dialogue content (such as character lines), and analyzes scene keywords (such as "farmland" and "big tree") and action keywords (such as "farming" and "waiting") in the text to match them to specific scenes in the storyboard script. For example, the narration text "The rabbit suddenly jumped out of the grass" contains keywords such as "grass" and "jump out", which can be identified as the "rabbit hits the tree" storyboard scene. This design ensures that the audio content corresponds to the visual scene one by one, lays the foundation for the synchronization of sound and picture, and avoids the logical error of "seaside scene with mountain sound effects".

[0077] Based on the emotional tone of the storyboard scene, select the appropriate speech synthesis model parameters to generate voice narration;

[0078] Specifically, this step aims to make the emotional expression of the voice narration highly consistent with the picture scene, solving the problem that traditional speech synthesis has a single emotion and cannot fit the plot. The system predefines the mapping relationship between emotional tone and voice parameters, such as "sarcasm" corresponds to falling tone + slow speed, "cheerfulness" corresponds to rising tone + fast speed, and "tragic" corresponds to deep timbre + pause. At the same time, in order to improve the emotional authenticity, an emotional classification model can be introduced to first score the emotional expression of the storyboard scene (such as the degree of contempt from -1 to 1), and then dynamically adjust the parameter weights. This design makes voice narration the core carrier of emotional expression in the video, and the accuracy of emotional transmission is greatly improved.

[0079] It can be understood that the parameters of the speech synthesis model include intonation (pitch changes), speaking speed (words / second), timbre (bright / thick), pauses (inter-sentence intervals), etc.; the emotion classification model outputs emotion labels (such as "joy" and "anger") and intensity values ​​by analyzing the text description of the storyboard scene and the character's actions.

[0080] Combine the type of storyboard scene and match the corresponding background sound effects from the sound effect library;

[0081] Specifically, the realism of the environment is enhanced through scene-adapted sound effects, solving the problem of insufficient immersion caused by the disconnection between sound effects and scenes. The system divides the storyboard scenes into natural scenes (such as forests and rivers), life scenes (such as farmland and streets), fantasy scenes (such as wonderland and monster caves), and other types. The sound effect library is stored by type: natural scenes match bird songs and water sounds, life scenes match labor sounds and human voices, and fantasy scenes match magic sound effects and monster roars. For example, the "farming" storyboard of "Waiting for the Rabbit" belongs to a life scene, which matches the sound of a hoe hitting the soil and the cry of a cow; the "rabbit hitting a tree" storyboard contains natural elements and superimposes the rustling sound of grass. To improve the matching accuracy, audio retrieval technology can be used to retrieve corresponding files in the sound effect library based on keywords in the scene description (such as "early morning" and "wilderness").

[0082] According to the time node of the storyboard script, the intensity of the voice narration and the background sound effect is adjusted to synchronize the voice narration, background sound effect and storyboard scene.

[0083] Specifically, accurate synchronization of audio and video is achieved through audio timing adjustment, solving the problem of audio and picture being out of sync in traditional manual editing. The system obtains the start time and duration of each shot in the storyboard script, stretches or compresses the voice narration in time (such as 5 seconds of storyboard corresponding to 5 seconds of voice), and adds audio fade-in and fade-out effects when the shot switches (such as 0.5 seconds of fading and 0.5 seconds of fading). For background sound effects, the sound effect peak (such as the sudden increase in the impact sound) is triggered according to the scene action node (such as the moment when the rabbit hits the tree).

[0084] It is understandable that audio timing adjustment technology includes the dynamic time warping (DTW) algorithm, which is used to match the duration of voice and the duration of the storyboard; audio fade-in and fade-out is achieved by adjusting the volume envelope, which is commonly used in audio processing tools such as Audition; the sound effect triggering mechanism is based on the action time point in the storyboard (for example, if the "hitting the tree" action occurs at the 2nd second of the storyboard, the impact sound will be played at the 2nd second).

[0085] In one embodiment, the step of combining the visual elements, the voice narration, and the background sound effects to generate a dynamic video and perform timing alignment and optimization includes:

[0086] Arranging the visual elements in a shot order to form an unsynchronized initial video sequence;

[0087] Specifically, this step establishes the basic narrative order of the video, providing structured material for subsequent temporal alignment and resolving the issue of chaotic arrangement of multiple materials. The system arranges the generated dynamic visual elements (such as character animations and scene images) in sequence according to the storyboard script's numbering sequence (e.g., Shot 1, Shot 2), forming an initial video sequence without audio synchronization. Each visual element is labeled with the storyboard ID, duration, and content description (e.g., "Shot 3: Farmer squatting under a tree, duration 4 seconds") to facilitate subsequent retrieval and adjustment.

[0088] It is understandable that dynamic visual elements include character animations with transparent channels, background scene images, and special effects materials (such as particle effects), which are stored in PNG sequence or video file format; video sequence arrangement is achieved through timeline editing tools, such as the timeline panel of Adobe Premiere, which supports dragging and dropping to adjust the order and set the material duration.

[0089] Using the duration of each frame in the initial video sequence, adjusting the playback speed of the corresponding voice narration;

[0090] Specifically, this step aims to solve the problem of mismatch between voice duration and frame duration, ensuring that the audio and picture are consistent in time. The system obtains the actual duration of each frame in the initial sequence, calculates the target duration of the corresponding paragraph of voice narration, and adjusts the playback speed through audio time stretching technology (such as WSOLA algorithm). To avoid voice distortion after acceleration, a speed adjustment threshold (such as ±30%) can be set, which triggers text regeneration (such as shortening the narration content) when exceeded.

[0091] It is understandable that the WSOLA (Waveform Similarity Overlap-Add) algorithm adjusts the duration without changing the pitch by analyzing the similarity of speech waveforms; audio speed adjustment tools such as Audition's "time stretching" function support real-time preview of the adjustment effect; the speed adjustment threshold is preset according to the characteristics of the speech synthesis model, usually not more than 30% for fast speed and not less than 70% for slow speed.

[0092] Aligning the adjusted voice narration with the initial video sequence on a time axis to obtain an intermediate video with synchronized audio and video;

[0093] Specifically, this step achieves precise synchronization between audio and video, resolving the inefficiency and large errors associated with traditional manual alignment. The system imports the adjusted voice narration into video editing software and, based on the start time of the storyboard, drags the voice track to align the narration with the visual action.

[0094] It is understandable that timeline alignment is achieved through the multi-track editing function of the video editing software, with the voice track and video track displayed side by side; waveform reference lines are used to identify accents and pauses in the voice, and video frame previews are used to locate key action scenes; synchronization error assessment can be achieved by comparing the audio and video during playback, and manually or automatically marking the positions that exceed the threshold.

[0095] Inserting a transition effect at the point where the storyboards of the intermediate video are switched;

[0096] Specifically, this step smooths out the transitions between scenes, enhancing the visual fluidity of the video and resolving the issue of abrupt cuts. The system selects transition effects based on the emotional relevance of the scenes: "Fade" for similar scenes (e.g., farmland at daytime to farmland at dusk), "Wipe" for contrasting scenes (e.g., quiet farmland to bustling street market), and "Magic Flash" for fantasy scenes.

[0097] It is understandable that the transition special effects library contains various preset effects, which are stored by scene type; the special effects parameters can be adjusted, such as the speed of fade-in and fade-out, the direction of wipe, and the color of flash; the transition trigger mechanism automatically inserts special effects clips on the timeline based on the end / start time of the storyboard.

[0098] Generate subtitles based on the synchronized voice content and overlay them on the video after special effects are inserted;

[0099] Specifically, subtitles are used to improve the comprehensibility of content and solve the problem of comprehension barriers in scenarios with fast speech speed or dialect. The system performs real-time speech recognition on the synchronized voice narration, generates subtitle text with time code, sets the subtitle style, and superimposes it on non-critical areas of the video screen. Subtitle generation needs to consider the language version (such as Chinese and English), font size (to adapt to different screens), and display / hide synchronously with the voice. For example, the children's version of the video uses large color fonts, and the academic version uses small white fonts + pinyin annotations.

[0100] It is understandable that speech recognition technology uses deep learning models (such as Whisper) to convert audio into text and mark timestamps; subtitle style settings include font, font size, color, and position; subtitle file formats are SRT, ASS, etc., and support importing into video editing software.

[0101] Mix the video with superimposed subtitles with background sound effects to output a dynamic video.

[0102] Specifically, this step solves the problem of unbalanced audio mixing by integrating all audio elements to form a complete multimedia work. The system mixes the video with superimposed subtitles with the background sound effects and voice narration according to the volume ratio, and adjusts the sound effect transition at the switching point of the storyboard to avoid sudden changes in volume. The output format is selected according to user needs (such as MP4, MOV), and parameters such as resolution (such as 1080p), frame rate (24 / 30fps), and bit rate are set. For example, the short video platform outputs 9:16 vertical screen MP4, and the educational courseware outputs 16:9 horizontal screen MOV.

[0103] It is understandable that the audio mixing is adjusted through the volume envelope to ensure that the voice narration is always clear and audible, and the background sound effects serve as an atmosphere supplement; the video output parameters can preset commonly used profiles and support one-click selection; a final check is performed before output to check the audio and video synchronization, subtitle display, and file integrity.

[0104] In one embodiment, the step of adjusting the story style, visual style, and output format of the dynamic video according to user needs to generate an idiom story video includes:

[0105] Provides a user interface to support selection of story complexity, visual style, and video format;

[0106] Specifically, this step establishes an interactive channel between the user and the system, so that personalized needs can be accurately obtained and the problem of vague demand communication is solved. The interactive interface adopts a graphical design, including: story complexity selection: children's version (simple language + less plot), standard version (standard narrative), academic version (in-depth analysis); visual style selection: cartoon style, ink style, realistic style, paper-cut style and other preset styles, with preview images; video format selection: horizontal screen (16:9), vertical screen (9:16), square screen (1:1), and display the applicable platform (such as horizontal screen is suitable for TV, vertical screen is suitable for TikTok); the interface supports real-time preview of style effects. For example, when "ink style" is selected, a sample image is automatically generated for user confirmation.

[0107] It is understandable that the interactive interface is developed based on Web technology (HTML+CSS+JavaScript) or integrated into desktop applications; the style preview image is quickly rendered through a lightweight image generation model without the need for complete video generation; the user selection data is passed to the back-end processing in JSON format, including style parameters, format parameters, etc.

[0108] Regenerate story scripts or visual elements of corresponding styles based on user selections;

[0109] Specifically, the system dynamically adjusts content output based on user needs to address the limitations of a "one-size-fits-all" generation model. After receiving the user's selection, if the story style is adjusted (e.g., children's version → academic version), the original story script is input into the large language model, with instructions such as "add allusion source" and "deepen the moral analysis" added to generate a new script. If the visual style is adjusted (e.g., cartoon style → ink painting style), the original storyboard prompts are supplemented with style keywords such as "ink brushstrokes" and "white space composition" and the image generation model is re-invoked to generate visual elements. For example, when the story "Waiting for the Rabbit" changes from a cartoon style to an ink painting style, the prompts need to add descriptions such as "rice paper texture," "thick and thin ink colors," and "brush strokes."

[0110] It is understandable that the story script is regenerated through fine-tuning of the instructions of the large language model, such as "Convert the following story into an academic version, and add a philosophical analysis of the source and meaning of "Han Feizi""; the visual elements are regenerated by modifying the style parameters of the prompt words, such as adding an ink style instruction after the original prompt words.

[0111] Resynthesize the dynamic video based on the adjusted content and output a customized idiom story video.

[0112] Specifically, the adjusted script and visual elements are integrated into the final product to ensure personalized needs are met and solve the problem of low efficiency caused by multiple revisions. The system re-synthesizes the video, generates voice narration using the adjusted story script, replaces the original material with the new visual elements, and outputs it in the format selected by the user. For example, taking the ink painting style + academic version of the "Foolish Old Man Moves Mountains" video as an example, the following steps are included:

[0113] Step 1: Use a calm tone for the voice narration and include readings of allusions and quotations;

[0114] Step 2: The visual elements are in the style of ink and wash landscape, and the image of Yugong is based on traditional Chinese painting figures;

[0115] Step 3: Output to 16:9 horizontal screen format, suitable for academic lectures.

[0116] Step 4: Automatically retain the previous timing alignment and special effects settings during the synthesis process, and only replace the main content to improve the efficiency of re-synthesis.

[0117] It is understandable that the existing timing alignment data (such as the synchronization points between voice and storyboard) is reused during re-synthesis without the need for readjustment; the visual element replacement adopts the "intelligent matching" algorithm to ensure that the duration and action nodes of the new material are consistent with the original storyboard; the output format parameters are automatically configured according to the user's selection, such as the resolution and bit rate of the horizontal screen.

[0118] The embodiment of the present invention further provides an idiom story video generation device, which is used to execute any embodiment of the above-mentioned idiom story video generation method. Figure 2 , Figure 2 : is a schematic block diagram of an idiom story video generation device provided by an embodiment of the present invention. The idiom story video generation device 700 includes:

[0119] A story script generation module 710 is used to generate an idiom story script based on the input idiom by using natural language processing technology;

[0120] A storyboard and visual generation module 720, configured to automatically generate a storyboard and corresponding visual elements based on the storyboard;

[0121] A speech synthesis module 730 is used to generate voice narration and background sound effects based on the story script;

[0122] A video synthesis module 740 is used to combine the visual elements, voice narration and background sound effects to generate a dynamic video and perform timing alignment and optimization;

[0123] The user customization module 750 is used to adjust the dynamic video according to user needs to generate a final idiom story video.

[0124] In one embodiment, the story script generation module 710 is specifically configured to:

[0125] Receive an idiom and optional parameters input by a user, wherein the optional parameters include a target audience and a story style;

[0126] Generate preliminary story content through pre-trained language model combined with idiom knowledge base;

[0127] A rule engine is used to verify whether the content of the preliminary story conforms to the original meaning of the idiom and its logical rationality;

[0128] Based on the verification results, the story script is output, including characters, scenes and moral information.

[0129] In one embodiment, the storyboard and visual generation module 720 is specifically configured to:

[0130] Parse the story script, extract key plots and split them into multiple storyboards;

[0131] Generate corresponding screen description prompts for each storyboard;

[0132] Based on the screen description prompt words, using an image generation model to generate a character image and a scene image;

[0133] Combine predefined animation templates to add dynamic actions to characters and form dynamic visual elements.

[0134] In one embodiment, the speech synthesis module 730 is specifically configured to:

[0135] Extracting the narration text and dialogue content in the story script and identifying the storyboard scenes corresponding to each paragraph of text;

[0136] Based on the emotional tone of the storyboard scene, select the appropriate speech synthesis model parameters to generate voice narration;

[0137] Combine the type of storyboard scene and match the corresponding background sound effects from the sound effect library;

[0138] According to the time node of the storyboard script, the intensity of the voice narration and the background sound effect is adjusted to synchronize the voice narration, background sound effect and storyboard scene.

[0139] In one embodiment, the video synthesis module 740 is specifically configured to:

[0140] Arranging the visual elements in a shot order to form an unsynchronized initial video sequence;

[0141] Using the duration of each frame in the initial video sequence, adjusting the playback speed of the corresponding voice narration;

[0142] Aligning the adjusted voice narration with the initial video sequence on a time axis to obtain an intermediate video with synchronized audio and video;

[0143] Inserting a transition effect at the point where the storyboards of the intermediate video are switched;

[0144] Generate subtitles based on the synchronized voice content and overlay them on the video after special effects are inserted;

[0145] Mix the video with superimposed subtitles with background sound effects to output a dynamic video.

[0146] In one embodiment, the user customization module 750 is specifically configured to:

[0147] Provides a user interface to support selection of story complexity, visual style, and video format;

[0148] Regenerate story scripts or visual elements of corresponding styles based on user selections;

[0149] Resynthesize the dynamic video based on the adjusted content and output a customized idiom story video.

[0150] In this embodiment, the idiom story video generation device 700 further includes a video verification module 760, which is specifically used to:

[0151] Perform quality verification on the generated idiom story video content and output the verification results;

[0152] Automatically optimize the idiom story video content based on the verification results and output the final version of the video.

[0153] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned idiom story video generation device and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.

[0154] The idiom story video generating device 700 can be implemented as a computer program. Figure 3 Runs on the computer equipment shown.

[0155] See also Figure 3 , Figure 3 800 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 800 can be a terminal or a server. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.

[0156] The computer device 800 includes a processor 820 , a memory, and a network interface 850 connected via a system bus 810 , wherein the memory may include a non-volatile storage medium 830 and an internal memory 840 .

[0157] The non-volatile storage medium 830 may store an operating system 831 and a computer program 832. The computer program 832 includes program instructions, which, when executed, may cause the processor 820 to execute a method for generating an idiom story video.

[0158] The processor 820 is used to provide computing and control capabilities to support the operation of the entire computer device 800.

[0159] The internal memory 840 provides an environment for the operation of the computer program 832 in the non-volatile storage medium 830. When the computer program 832 is executed by the processor 820, the processor 820 can execute a method for generating an idiom story video.

[0160] The network interface 850 is used to communicate with other devices through the network. Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 800 to which the solution of the present application is applied. The specific computer device 800 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0161] It should be understood that in the embodiment of the present application, the processor 820 may be a central processing unit (CPU), and the processor 820 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0162] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0163] Therefore, the present application also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps:

[0164] S100, based on the input idiom, generating an idiom story script through natural language processing technology;

[0165] S200, automatically generating a storyboard script and corresponding visual elements according to the story script;

[0166] S300, generating voice narration and background sound effects based on the story script;

[0167] S400, combining the visual elements, the voice narration, and the background sound effects to generate a dynamic video and perform timing alignment and optimization;

[0168] S500: Adjust the story style, visual style and output format of the dynamic video according to user needs to generate an idiom story video.

[0169] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0170] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0171] The non-Company software tools or components appearing in the embodiments of this application are merely examples and do not represent actual use.

[0172] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and other division methods may be used in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not implemented.

[0173] The steps in the method of the embodiment of the present application can be adjusted in order, combined, and deleted according to actual needs. The units in the device of the embodiment of the present application can be combined, divided, and deleted according to actual needs. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.

[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application.

[0175] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for generating idiom story videos, characterized in that: The following steps are involved: Based on the input idioms, idiom story scripts are generated through natural language processing technology; Automatically generate storyboards and corresponding visual elements based on the story script; Based on the story script, generate voice narration and background sound effects; Combining the visual elements, the voice narration, and the background sound effects to generate a dynamic video and perform timing alignment and optimization; The story style, visual style and output format of the dynamic video are adjusted according to user needs to generate an idiom story video.

2. The idiom story video generation method according to claim 1, characterized in that: The step of generating an idiom story script based on the input idiom by using natural language processing technology includes: Receive an idiom and optional parameters input by a user, wherein the optional parameters include a target audience and a story style; Generate preliminary story content through pre-trained language model combined with idiom knowledge base; A rule engine is used to verify whether the content of the preliminary story conforms to the original meaning of the idiom and its logical rationality; Based on the verification results, the story script is output, including characters, scenes and moral information.

3. The idiom story video generation method according to claim 1, characterized in that: The step of automatically generating a storyboard script and corresponding visual elements according to the story script includes: Parse the story script, extract key plots and split them into multiple storyboards; Generate corresponding screen description prompts for each storyboard; Based on the screen description prompt words, using an image generation model to generate a character image and a scene image; Combine predefined animation templates to add dynamic actions to characters and form dynamic visual elements.

4. The idiom story video generation method according to claim 1, characterized in that: The step of generating voice narration and background sound effects based on the story script includes: Extracting the narration text and dialogue content in the story script and identifying the storyboard scenes corresponding to each paragraph of text; Based on the emotional tone of the storyboard scene, select the appropriate speech synthesis model parameters to generate voice narration; Combine the type of storyboard scene and match the corresponding background sound effects from the sound effect library; According to the time node of the storyboard script, the intensity of the voice narration and the background sound effect is adjusted to synchronize the voice narration, background sound effect and storyboard scene.

5. The idiom story video generation method according to claim 1, characterized in that: The steps of combining the visual elements, the voice narration, and the background sound effects to generate a dynamic video and perform timing alignment and optimization include: Arranging the visual elements in a shot order to form an unsynchronized initial video sequence; Using the duration of each frame in the initial video sequence, adjusting the playback speed of the corresponding voice narration; Aligning the adjusted voice narration with the initial video sequence on a time axis to obtain an intermediate video with synchronized audio and video; Inserting a transition effect at the point where the storyboards of the intermediate video are switched; Generate subtitles based on the synchronized voice content and overlay them on the video after special effects are inserted; Mix the video with superimposed subtitles with background sound effects to output a dynamic video.

6. The idiom story video generation method according to claim 1, characterized in that: The step of adjusting the story style, visual style and output format of the dynamic video according to user needs to generate an idiom story video includes: Provides a user interface to support selection of story complexity, visual style, and video format; Regenerate story scripts or visual elements of corresponding styles based on user selections; Resynthesize the dynamic video based on the adjusted content and output a customized idiom story video.

7. The method for generating idiom story videos according to claim 1, wherein: The steps after adjusting the story style, visual style and output format of the dynamic video according to user needs and generating the idiom story video also include: Perform quality verification on the generated idiom story video content and output the verification results; Automatically optimize the idiom story video content based on the verification results and output the final version of the video.

8. A device for generating idiom story videos, characterized in that: include: A story script generation module is used to generate idiom story scripts based on input idioms through natural language processing technology; A storyboard and visual generation module, configured to automatically generate a storyboard script and corresponding visual elements based on the story script; A speech synthesis module, for generating voice narration and background sound effects based on the story script; A video synthesis module, for combining the visual elements, voice narration, and background sound effects to generate dynamic video and perform timing alignment and optimization; The user customization module is used to adjust the dynamic video according to user needs and generate the final idiom story video.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the idiom story video generation method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the idiom story video generation method as described in any one of claims 1 to 7 are implemented.