Text-to-video method and device based on deep semantic analysis, equipment and medium

By using deep semantic analysis and deep learning models, financial videos are automatically generated, solving the problem of insufficient intelligence and automation of existing tools in the financial field, and achieving efficient, professional and coherent video content generation.

CN119399330BActive Publication Date: 2025-11-25CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411439474.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-11-25
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

Existing video production tools lack intelligent and automated processing in the financial field, making it difficult to accurately understand financial jargon and complex terminology. This results in generated video content that deviates from the intended meaning, with insufficient visual and auditory presentation, disjointed content, and a lack of professionalism, thus increasing the complexity of production.

Method used

By extracting keywords and named entities from text through deep semantic analysis, analyzing text structure and core themes, generating semantic analysis results, establishing mapping relationships between visual elements, generating animation effects using deep learning models, constructing a video timeline, and integrating visual and auditory elements to generate a complete video.

Benefits of technology

It enables automatic generation of video content, improves production efficiency, ensures professionalism and content consistency, reduces human intervention, and enhances the dynamic presentation and visual continuity of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399330B_ABST
    Figure CN119399330B_ABST
Patent Text Reader

Abstract

The application relates to the fields of artificial intelligence and financial technology, and discloses a text-to-video method based on deep semantic analysis, which extracts keywords and named entities from the text to be converted, analyzes the core theme and text structure of the text, performs semantic analysis to generate a semantic analysis result, retrieves matched images, video clips, sound effects and background music according to the semantic analysis result, and generates a preliminary visual layout framework through spatial layout; a video timeline is constructed, layout schemes, background music and sound effects are integrated to generate complete video content, and finally, a video file is output. The application accurately extracts keywords, named entities, core themes and text structures in the text through semantic analysis, and intelligently matches images, video clips, sound effects and background music that are consistent with the semantics, thereby avoiding the problem of requiring a large amount of manual intervention. Through a deep learning model, adaptive animation effects are generated, so that the dynamic presentation of the video content and the coherence of the visual effects are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and fintech, and particularly to a text-to-video method, apparatus, device, and storage medium based on deep semantic analysis. Background Technology

[0002] In the financial sector, video creation tools are increasingly becoming important tools for promoting financial services and educating users. Currently, financial institutions use videos to showcase complex financial products, investment strategies, and risk control methods, in order to convey information to customers more intuitively. However, existing digital media tools face multiple challenges in producing financial videos, severely limiting their effectiveness and efficiency.

[0003] First, existing video production tools lack intelligent and automated processing. While some tools support automation, users still need to manually adjust video content in most cases, especially when the video involves highly technical content such as financial data and investment analysis. This not only increases editing time but also affects the video's coherence and the professionalism of the content, leading to inefficiencies for financial institutions when promoting and educating customers.

[0004] Secondly, existing tools have limitations in semantic understanding and content matching. The financial industry involves a large number of complex terms, data, analyses, and technical jargon; however, existing text-to-video technologies struggle to accurately understand the specialized terminology, complex industry jargon, and their inherent relationships within the financial field. Inaccurate semantic understanding can cause the generated video content to deviate from the financial institution's intent, failing to effectively convey key information and reducing the video's value.

[0005] Furthermore, existing technologies fall short in scene matching and sound effect synthesis for visual and auditory presentation. Financial videos often demand rigor and professionalism in information delivery, but current video creation tools frequently fail to intelligently match financial scenarios and generate suitable data visualizations. For example, when showcasing investment returns or risk predictions, existing tools may not automatically generate appropriate graphics and visual effects, requiring significant manual intervention from users. Moreover, the selection of sound effects and background music lacks intelligent support, making it difficult to automatically generate suitable background sound effects based on the video's content and tone, further increasing the complexity of video production.

[0006] Finally, although some platforms support automatically generating videos from text or PowerPoint presentations, the resulting videos often suffer from incoherent content and insufficient professionalism, especially when dealing with financial information. The logical consistency of the content, the accuracy of the data, and the presentation of charts often require financial professionals to repeatedly edit and revise the generated results, consuming a significant amount of time and effort. Summary of the Invention

[0007] The main objective of this invention is to provide a text-to-video method, apparatus, device, and storage medium based on deep semantic analysis, aiming to solve the technical problem that existing technologies cannot intelligently match suitable visual elements, background music, and sound effects based on text semantic analysis to automatically generate coherent and professional video content.

[0008] To achieve the above objectives, this invention provides a text-to-video method based on deep semantic analysis, comprising:

[0009] Obtain the text to be converted, perform word segmentation and part-of-speech tagging on the text to be converted, and extract keywords and named entities from the text to be converted;

[0010] Analyze the core theme and text structure of the text to be converted, including headings, subheadings, and / or paragraphs;

[0011] Semantic analysis is performed based on the keywords, named entities, text structure, and core themes to generate semantic analysis results.

[0012] Based on the semantic analysis results, a mapping relationship between semantics and visual elements is established, and images, video clips, sound effects, and background music that match the semantic analysis results are retrieved from the material library.

[0013] Based on the semantic analysis results, the text content corresponding to the keywords and named entities is spatially arranged with images and video clips to generate a preliminary visual layout framework.

[0014] The layout content features of the preliminary visual layout framework are analyzed using a deep learning model.

[0015] Animation effects are generated based on the layout content features, and the animation effects are applied to the text content and visual elements in the initial visual layout framework.

[0016] Based on the initial visual layout framework with applied animation effects, background music, and sound effects, construct the video timeline;

[0017] Based on the video timeline, the initial visual layout framework, background music, and sound effects are integrated to generate the integrated complete video content.

[0018] Furthermore, to achieve the above objectives, the present invention provides a text-to-video device based on deep semantic analysis, comprising:

[0019] The text processing module is used to acquire the text to be converted, perform word segmentation and part-of-speech tagging on the text to be converted, and extract keywords and named entities from the text to be converted.

[0020] The text analysis module is used to analyze the core theme and text structure of the text to be converted, the text structure including titles, subtitles and / or paragraphs;

[0021] The semantic analysis module is used to perform semantic analysis based on the keywords, named entities, text structure, and core themes, and generate semantic analysis results.

[0022] The material retrieval module is used to establish a mapping relationship between semantics and visual elements based on the semantic analysis results, and to retrieve images, video clips, sound effects and background music that match the semantic analysis results from the material library;

[0023] The layout generation module is used to spatially arrange the text content corresponding to the keywords and named entities with images and video clips based on the semantic analysis results, and generate a preliminary visual layout framework.

[0024] The layout feature analysis module is used to analyze the layout content features of the preliminary visual layout framework using a deep learning model.

[0025] An animation effect generation module is used to generate animation effects based on the layout content features and apply the animation effects to the text content and visual elements in the initial visual layout framework.

[0026] The timeline building module is used to construct a video timeline based on a preliminary visual layout framework with applied animation effects, background music, and sound effects.

[0027] The video generation module is used to integrate the initial visual layout framework, background music, and sound effects according to the video timeline to generate integrated complete video content.

[0028] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a text-to-video program based on deep semantic analysis stored in the memory and executable on the processor, wherein when the text-to-video program based on deep semantic analysis is executed by the processor, it implements the steps of the text-to-video method based on deep semantic analysis as described above.

[0029] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a text-to-video program based on deep semantic analysis, wherein the text-to-video program based on deep semantic analysis, when executed by a processor, implements the steps of the text-to-video method based on deep semantic analysis as described above.

[0030] Beneficial Effects: This invention relates to the fields of artificial intelligence and fintech, and discloses a text-to-video method based on deep semantic analysis. It extracts keywords and named entities from the text to be converted, analyzes the core theme and text structure, performs semantic analysis to generate semantic analysis results, retrieves matching images, video clips, sound effects, and background music based on the semantic analysis results, generates a preliminary visual layout framework through spatial layout, uses a deep learning model to generate animation effects and applies them to the layout scheme, constructs a video timeline, integrates the layout scheme, background music, and sound effects to generate complete video content, and finally outputs a video file. This invention accurately extracts keywords, named entities, core themes, and text structure from the text through semantic analysis, and intelligently matches images, video clips, sound effects, and background music that match the semantics, avoiding the problem of requiring a lot of manual intervention in existing technologies. Simultaneously, it generates adaptive animation effects through a deep learning model and applies them to the preliminary visual layout framework, ensuring the dynamic presentation of video content and the coherence of visual effects. After constructing the video timeline and integrating all elements, it automatically generates complete video content, significantly improving video production efficiency and ensuring professionalism and content consistency. Attached Figure Description

[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0032] Figure 1 This is a schematic diagram of an application environment for a text-to-video method based on deep semantic analysis in one embodiment of the present invention;

[0033] Figure 2 This is a flowchart illustrating an embodiment of the text-to-video method based on deep semantic analysis of the present invention.

[0034] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the text-to-video device based on deep semantic analysis of the present invention.

[0035] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0036] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0037] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0038] The text-to-video method based on deep semantic analysis provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can extract keywords and named entities from the text to be converted through the user terminal, analyze the core theme and text structure, perform semantic analysis to generate semantic analysis results, retrieve matching images, video clips, sound effects, and background music based on the semantic analysis results, generate a preliminary visual layout framework through spatial layout, generate animation effects using a deep learning model and apply them to the layout scheme, construct a video timeline, integrate the layout scheme, background music, and sound effects to generate complete video content, and finally output as a video file. This invention accurately extracts keywords, named entities, core themes, and text structure from text through semantic analysis, and intelligently matches images, video clips, sound effects, and background music that match the semantics, avoiding the problem of requiring a lot of manual intervention in existing technologies. Simultaneously, it generates adaptive animation effects through a deep learning model and applies them to the preliminary visual layout framework, ensuring the dynamic presentation of video content and the coherence of visual effects. After constructing the video timeline and integrating all elements, it automatically generates complete video content, significantly improving video production efficiency and ensuring professionalism and content consistency. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0039] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the text-to-video method based on deep semantic analysis provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0040] like Figure 2 As shown, the text-to-video method based on deep semantic analysis proposed in this invention includes the following steps:

[0041] S10, Obtain the text to be converted, perform word segmentation and part-of-speech tagging on the text to be converted, and extract keywords and named entities from the text to be converted;

[0042] In this embodiment, the text to be converted can be in various formats, such as ordinary TXT text files, PPT files, or structured or unstructured files like financial reports. For subsequent word segmentation and semantic analysis, the system needs to extract and preprocess the input text data. The file input module can accept files manually uploaded, via API, or read from a specified path. The system supports multiple file types and will parse the content of different file formats during preprocessing, such as parsing PPT text content or extracting text from PDF files.

[0043] Word segmentation breaks down the text to be converted into its smallest linguistic units (vocabularies) so that the system can identify specific words in the text. This is a crucial operation, especially in Chinese. Part-of-speech tagging (POS) adds a corresponding part-of-speech tag to each segmented word, such as noun, verb, or adjective, to aid in subsequent semantic understanding. Existing word segmentation algorithms can be used, such as the jieba segmenter for Chinese or other open-source natural language processing libraries. For financial terminology, a customized segmentation dictionary can be developed to ensure accurate segmentation. Based on natural language processing techniques (such as Stanford NLP or SpaCy), POS tagging is performed on the segmented text. The tagged words are then accompanied by corresponding part-of-speech tags for subsequent semantic analysis.

[0044] Keyword extraction identifies the most representative and meaningful words from text to aid in understanding its core content. Named entity recognition (NER) detects proper nouns in text, such as names of people, places, and organizations, which is particularly important in the financial field. Based on word segmentation and part-of-speech tagging results, techniques such as TF-IDF (Term Frequency-Inverse Document Frequency) or TextRank are used to extract keywords with high weights from the text. Keywords can be sorted according to their frequency and importance. Named entity recognition models (such as BERT, SpaCy, and StanfordNER) are used to scan the text and identify proper nouns such as names of people, places, and organizations. A list of entities with labels (such as names of people and places) is generated based on the recognition results. Keywords and named entities are stored in the system and used in subsequent steps for semantic analysis and visual mapping.

[0045] Example Explanation: Taking an annual investment report in the financial sector as an example, the system receives the input investment report document, automatically segments the financial terms (such as "price-to-earnings ratio" and "return on capital"), and extracts keywords. Simultaneously, the named entity recognition module accurately extracts key information from the report, such as company names, financial institution names, and geographical locations, providing semantic information for subsequent video generation. For example, the keyword "asset management" will be matched with relevant icons in the image library, while the institution name "JP Morgan" will automatically match with the corresponding company logo and be used for display in the video, thus enabling the generated video to more accurately convey professional content in the financial field.

[0046] Through the above steps, intelligent word segmentation, keyword extraction, and named entity recognition of text data are achieved, greatly improving the accuracy of semantic analysis.

[0047] S20, Analyze the core theme and text structure of the text to be converted, the text structure including title, subtitle and / or paragraph;

[0048] In this embodiment, the core theme of the text to be converted is analyzed, and the structural hierarchy of the text (such as titles, subtitles, paragraphs, etc.) is identified to provide logic and coherence for subsequent video generation. Core theme analysis mainly uses semantic modeling technology to identify the main idea of ​​the text, while text structure analysis aims to understand the hierarchical layout of the text, thereby providing a framework for matching visual elements and animation effects.

[0049] Based on high-frequency keywords, named entities, contextual relationships, and syntactic structures in the text, the core theme of the text is extracted using a pre-defined topic model (such as LDA topic model or BERT pre-trained model). This core theme is usually the main idea or purpose expressed in the text.

[0050] It identifies structural information in text, such as headings, subheadings, and paragraphs, and automatically distinguishes different levels of text blocks based on features such as line spacing, font size, and punctuation. Headings typically have greater weight, while paragraphs are used to carry more detailed content.

[0051] After analyzing the structure of headings, subheadings, paragraphs, etc., a hierarchical text structure diagram is constructed based on their content and logical relationships. This hierarchical structure will be used for subsequent visual element layout and animation effect generation.

[0052] Topic extraction is performed using a pre-defined topic model (such as LDA). The model generates the main topics of the text based on keyword distribution and assigns corresponding topic weights to each paragraph or chapter. In the financial field, especially in annual reports or market analysis reports, common topics include "risk control," "asset allocation," and "market outlook." The system is trained and matched to these topics using a training corpus.

[0053] Headings and subheadings are identified based on text formatting information (such as font size and bold style), and their content is extracted. Headings are typically given higher weight in a document. Paragraphs are identified using sentence segmentation techniques from natural language processing. Paragraphs usually consist of multiple sentences, each addressing a small theme and forming the overall framework of the document.

[0054] Based on the identified titles, subtitles, and paragraphs, a hierarchical text structure is generated. This structure reflects the logical organization and content arrangement of the text. Different levels of text structure will correspond to different visual effects and display order in the subsequent video generation, with titles often having the highest weight.

[0055] Example Explanation: When processing financial annual reports, the system can automatically identify "Market Analysis" as the core theme and extract relevant paragraphs and charts. Simultaneously, the system will recognize headings such as "2023 Market Performance" and subheadings like "Stock Market" and "Bond Market," helping to construct the report's logical hierarchy. The generated video will then arrange visual content according to this hierarchical structure. For example, relevant charts will be displayed when playing "Stock Market Performance," and visual backgrounds and animation effects will switch when transitioning to "Bond Market Performance," ensuring the content's coherence and professionalism.

[0056] By analyzing core themes and recognizing text structure, we can not only improve the understanding of text content during video generation, but also ensure that the logical structure of the text is mapped onto the visual layout, thereby ensuring the logical consistency of the generated video content. The hierarchical structure of the text facilitates the orderly presentation of visual elements, ensuring that viewers can more clearly understand the core information of the text, especially in the financial field, ensuring the logical coherence and coherence of reports and analyses.

[0057] S30, perform semantic analysis based on the keywords, named entities, text structure and core themes, and generate semantic analysis results;

[0058] In this embodiment, semantic analysis results are generated by comprehensively analyzing keywords, named entities, text structure, and core themes in the text. This process is a crucial step in text understanding, ensuring that the system accurately extracts the semantic relationships within the text, enabling the mapping between semantics and visual elements in subsequent steps. Keywords are representative words extracted from the text using natural language processing techniques (such as TF-IDF and word frequency analysis). Named entities are specific categories of words extracted from the text using named entity recognition technology, such as names of people, places, and organizations.

[0059] Association analysis techniques are used to determine the semantic relationships between keywords and named entities. These relationships can be inferred through context or syntactic structure; for example, if keywords and named entities appear together in the same paragraph, it indicates that they may be semantically related.

[0060] The hierarchical relationships within text structures (such as headings, subheadings, and paragraphs) also play a crucial role in semantic analysis. The system needs to analyze the semantic logic between different parts of the text by utilizing the hierarchical information of the text structure. For example, a heading may express the main theme of the entire text, while paragraphs provide detailed explanations, and subheadings serve to break up the text into sections. Text structure analysis helps the system determine the importance of certain keywords and named entities, further enhancing the accuracy of semantic analysis.

[0061] The core theme is the main idea derived from thematic analysis of the text. Each paragraph and sentence may be associated with one or more core themes. During semantic analysis, it is necessary to determine the degree of association between each paragraph and the core themes and generate the semantic correlation between paragraphs and the core themes. This helps the system understand the semantic emphasis of each paragraph, thereby allowing for the appropriate arrangement of visual elements in the subsequent video generation process.

[0062] The semantics of named entities and keywords depend not only on their inherent meaning but also on their context. Contextual analysis, by considering the context, determines the semantic meaning of named entities and keywords within a specific context. This analysis ensures that the system can recognize subtle semantic differences.

[0063] The system first extracts keywords and named entities from the text. Then, through contextual and syntactic analysis, it determines their relationship within the text. For example, if the keyword "market" and the named entity "Shanghai" frequently appear in the same paragraphs, the system can infer that they are semantically related.

[0064] The system analyzes the hierarchical structure of the text (such as headings, subheadings, and paragraphs) to determine the semantic function of each part. For example, the heading can be considered the main semantic expression of the entire text, while subheadings and paragraphs serve to refine and expand upon it. The system further optimizes semantic analysis based on this hierarchical relationship.

[0065] Using models such as LDA or BERT, the system matches paragraph content with the core topic, generating a semantic correlation between the paragraph and the topic. This helps the system prioritize content that is more closely related to the core topic when processing paragraphs. The system uses contextual analysis techniques (such as dependency-based syntactic analysis) combined with the context of the sentences containing keywords and named entities to determine their specific semantics.

[0066] Example Explanation: In financial applications, analyzing the core themes and text structure of the text to be converted can effectively help financial institutions achieve intelligent content generation. Taking a financial report from a large commercial bank as an example, the text in the report may include a title, subtitles, and multiple paragraphs detailing the bank's financial situation. By performing word segmentation, keyword extraction, and named entity recognition on the text, core financial terms such as "loans," "deposits," and "net profit," as well as related named entities such as "[Bank Name]" and "regional branch," can be accurately identified.

[0067] Through comprehensive semantic analysis of keywords, named entities, text structure, and core themes, the system can accurately capture key information in the text and generate in-depth semantic analysis results. This semantic analysis helps the system generate scenes, backgrounds, and visual effects highly relevant to the text content during subsequent visual element matching, making the video content more coherent and logical. This is particularly important in the financial field, ensuring that data, charts, and text content can be appropriately matched and presented.

[0068] S40, Based on the semantic analysis results, establish a mapping relationship between semantics and visual elements, and retrieve images, video clips, sound effects, and background music that match the semantic analysis results from the material library;

[0069] In this embodiment, based on the results of semantic analysis, the system will associate keywords and named entities in the text with images, video clips, etc. in the visual material library through the mapping rules between semantics and visual elements.

[0070] The system associates keywords such as "net profit" with chart materials and named entities such as "a certain bank" with images of bank buildings. It uses preset mapping rules or automatically learns the association between semantics and visual elements through machine learning models. This step relies on semantic analysis results to ensure that each keyword and named entity has a corresponding visual material.

[0071] The system will automatically retrieve images and video clips from the media library based on extracted keywords. For text in the financial field, such as "loan," the system may retrieve financial charts or data displays related to loans.

[0072] For keywords in the text, such as "growth" and "profit," images of rising arrows or videos of financial growth charts will be matched. Based on the semantic information and context of the keywords, matching images and video clips are retrieved, ensuring that visual materials accurately convey the text content.

[0073] The system retrieves image or video clips matching financial institutions, place names, etc., based on named entity information. For example, when the system recognizes the named entity "[Bank Name]", it will retrieve relevant images or videos from the resource library, such as bank buildings and meeting scenes. For instance, the named entity "[Bank Name]" might correspond to an image of a bank building; "financial regulatory agency" might match a video of a formal meeting scene. Combining the category of the named entity (such as people or institutions) and the semantic meaning of the context, the system retrieves and selects matching visual materials.

[0074] The system retrieves visual materials that convey the logic of the text based on information such as titles, subheadings, and paragraph levels. For example, the main title can be displayed as a full-screen image, the subheading with a simple background image, and the paragraph content unfolds gradually. For instance, the title "Annual Financial Report" could be matched with a simple report cover image, while the paragraph content "Net Profit Analysis" might be matched with detailed financial tables or charts. Combining the hierarchical structure of the text, the system automatically selects appropriate visual materials for each level of content, ensuring that the visual presentation matches the text structure.

[0075] The system selects background music and sound effects that match the sentiment of the text based on the sentiment tone determined by semantic analysis. For example, positive emotions (such as financial growth) will be accompanied by cheerful music, while neutral or negative emotions (such as losses) will be accompanied by calming or somber background music. Similarly, financial growth might use upbeat music, while financial decline might use slow background music. The system determines the sentiment tone of the text through the sentiment analysis module and automatically selects matching materials from its sound effects and background music library.

[0076] By combining semantic analysis results with the mapping of visual elements, financial text content can be intelligently matched with images, video clips, sound effects, and background music in a visual material library, automatically generating a visual presentation that matches the text content. This not only improves the efficiency of financial video generation but also ensures the coherence and professionalism of the video content, reduces the need for human intervention, and effectively improves the quality and expressiveness of video production.

[0077] S50, Based on the semantic analysis results, the text content corresponding to the keywords and named entities is spatially arranged with the images and video clips to generate a preliminary visual layout framework;

[0078] In this embodiment, the system spatially arranges the text information corresponding to keywords and named entities with matching images and video clips based on semantic analysis results. This arrangement determines the relative position, order, and display format of the text and visual materials. For example, the keyword "net profit" may be placed in a prominent position on the screen and paired with related charts and images to create a contrasting display. Similarly, the text of the named entity "[Bank Name]" may be displayed side-by-side with an image of the bank building or related video clips, making the information clear and intuitive. Based on the principle of visual harmony, the system may left-align the text information and right-align the images or video clips to ensure the content is easy to understand and conforms to people's visual browsing habits.

[0079] The system first analyzes the semantic analysis results to determine the corresponding visual materials for each keyword and named entity. Next, the system spatially arranges the text and visual materials according to visual design rules. For example, important financial data may be prominently displayed in the center or upper left corner of the screen, while background images or videos serve as supplementary content, providing context or a sense of situation. Using layout algorithms, a reasonable display area is calculated based on the text content and image size to ensure that text and visual elements do not overlap and that information is clearly conveyed.

[0080] Through the above steps, the system generates a preliminary visual layout framework, which includes the spatial distribution of text information, images, and video clips on the screen. This framework forms the basis for subsequent animation effects generation and video production.

[0081] The layout plan will detail the specific position, size, and spacing of each element (such as keywords, named entities, images, and videos). The system will use the initial layout plan to determine the application of subsequent animation effects to ensure visual smoothness.

[0082] The system generates a complete layout scheme based on the aforementioned layout strategy and converts it into specific layout coordinates and arrangement rules. This scheme can be adjusted and optimized in subsequent steps to adapt to the needs of the video timeline, such as dynamically displaying text and images, and the temporal order of video clips.

[0083] By strategically arranging keywords, named entities, and visual elements in a logical spatial layout, it's possible to automatically generate layout schemes that conform to visual design principles, reducing manual intervention. Particularly in the financial sector, this allows for the intuitive presentation of financial data and corresponding charts, simplifying information display complexity. The resulting videos are both professional and efficient in conveying information, enhancing user experience and content quality.

[0084] S60, Analyze the layout content features of the preliminary visual layout framework using a deep learning model;

[0085] In this embodiment, a deep learning model is used to analyze the initially generated visual layout framework. Layout content features refer to the visual attributes of each visual element in the layout, such as spatial arrangement, size, color, hierarchy, and contrast. Through analysis by the deep learning model, the layout scheme can be evaluated in terms of visual aesthetics and information transmission efficiency, providing a basis for subsequent animation effect generation.

[0086] The previously generated preliminary visual layout framework is input into a deep learning model. This approach includes the spatial layout information of text content, images, video clips, etc., on the screen. A pre-trained deep learning model (such as a CNN convolutional neural network) is used to extract and analyze features from the visual layout framework. The model extracts important features from the layout, including but not limited to the contrast of visual elements, color matching, and information transmission paths.

[0087] The model evaluates the visual appeal of a layout by analyzing the spatial distribution and contrast of elements, and provides optimization suggestions. For example, it ensures that the visual focus is on key locations and avoids excessive element piling up that could negatively impact the user's viewing experience.

[0088] The model generates an analysis report, pointing out the advantages and disadvantages of the layout, and providing a basis and guidance for the next step of generating animation effects.

[0089] Example Explanation: In generating a video analyzing annual financial data, the system uses a deep learning model to analyze layout features, ensuring the logical arrangement of line charts and bar charts. With the model's help, the system automatically optimizes the color scheme of the charts, avoiding visual clashes in data visualization, while ensuring that text information (such as "earnings growth rate") is displayed in the optimal position. The model also suggests adjusting the chart size to make the core data clearer, greatly improving the video's communicative effect and visual appeal.

[0090] Layout analysis using deep learning models can effectively improve the visual appeal and information delivery efficiency of video content. The model can automatically identify unreasonable layout designs and provide improvement suggestions, avoiding tedious manual adjustments.

[0091] S70, generate animation effects based on the layout content features, and apply the animation effects to the text content and visual elements in the initial visual layout framework;

[0092] In this embodiment, suitable animation effects are generated based on the analyzed visual layout content features, and these effects are applied to the text content and visual elements in the layout. Animation effects may include fade-in / fade-out, scaling, movement, rotation, etc., of text or images. The purpose is to enhance the expressiveness and information delivery efficiency of the video through appropriate dynamic effects, making the video more vivid and attracting the viewer's attention.

[0093] Key information is extracted from the layout content features obtained through previous analysis using deep learning models, including visual features such as the size, position, and contrast of text, images, or video clips.

[0094] Based on the analyzed layout features, select appropriate animation effects. For example, important visual elements can be highlighted using animation effects such as zooming or fading, while background elements may only require a simple fade-in / fade-out effect.

[0095] Configure animation parameters for each visual element, including animation duration, start time, end time, opacity change, and scaling. The animation effects should be designed to match the layout features to ensure consistency and smoothness of the video content.

[0096] The generated animation effects will be applied to the text and visual elements within the initial visual layout framework. The animation effects will be embedded into the timeline to ensure that the visual elements are displayed in the set order and rhythm during playback.

[0097] After applying the animation, preview the video to check its smoothness and consistency. Fine-tune the animation as needed to ensure optimal video quality.

[0098] Example Explanation: In generating a video of an annual financial report, the system designed animation effects for charts and text based on the visual layout characteristics of the financial data. For example, the profit curve appears with a zoom animation when it enters the screen, while the net profit figure is highlighted through color changes and a sliding effect. Through these animation effects, the video not only showcases the dynamic changes in financial data but also helps viewers quickly understand the key points of the data, greatly enhancing the video's visual appeal and the clarity of information delivery.

[0099] By generating animation effects based on layout and content features, the system can automatically design the most suitable dynamic effects for each visual element. This not only enhances the visual appeal of the video but also guides the viewer's attention to key content, improving the accuracy of information delivery. Especially in financial data videos, the dynamic display of key data helps viewers quickly understand complex trend changes, thereby enhancing the video's professionalism and user experience.

[0100] S80, based on the initial visual layout framework with applied animation effects, background music and sound effects, constructs the video timeline;

[0101] In this embodiment, the animation effects, background music, and sound effects in the generated preliminary visual layout framework are integrated and arranged and synchronized reasonably based on the timeline. The timeline is the core structure of the entire video playback, defining the start time, duration, and end time of each element to ensure that all visual and auditory elements play in a predetermined time order, achieving a smooth presentation of the video content.

[0102] Create a video timeline frame to hold various elements from the video, including text, images, video clips, animations, background music, and sound effects. The length of the timeline is determined based on the total duration of the video.

[0103] Based on the initial visual layout framework, apply the animation effects of each visual element (such as text, images, and video clips) to the timeline, determining the start time, duration, and transition effects for each element. Text and images may require shorter display times, while video clips will need longer durations.

[0104] Embed background music into the timeline, adjust the start and end times of the background music and its synchronization with animation effects to ensure a smooth integration of background music and video content, and avoid conflicts or breaks between background music and visual elements.

[0105] Place sound effects at their corresponding positions on the timeline based on events in the video (such as the appearance of animations). The trigger time of the sound effect needs to coincide with the corresponding visual event to ensure audiovisual synchronization.

[0106] Preview the video's timeline to ensure all elements are timed logically and coherently. Adjust the timing as needed, such as shortening or lengthening the duration of certain animation effects, or rearranging the timing of sound effects.

[0107] Example Explanation: When generating a quarterly financial report video, the system first constructs a timeline based on the order in which the financial data is presented, synchronizing the animation of the line graph showing the company's revenue growth with background music. When key data such as net profit and gross profit are highlighted, sound effects are triggered during those times, allowing viewers to perceive key information simultaneously through sight and sound. This timeline construction ensures the video's smoothness, avoids misalignment or interference between elements, and ultimately achieves a professional and easy-to-understand presentation.

[0108] By constructing a video timeline based on an initial visual layout framework with applied animation effects, background music, and sound effects, the synchronization and coordination of all visual and auditory elements within the video are ensured, enhancing the video's coherence and viewing experience. This timeline construction not only ensures smooth playback of the video content but also allows for adjustments to the arrangement of background music and sound effects according to the rhythmic changes of different content, making the video more rhythmic and engaging, particularly suitable for the precise presentation of financial reports and market analysis videos.

[0109] S90, based on the video timeline, integrate the initial visual layout framework, background music, and sound effects to generate integrated complete video content.

[0110] In this embodiment, based on the constructed video timeline, all elements (text content, images, video clips, etc.), background music, and sound effects in the initial visual layout framework are finally integrated to ensure that all elements are arranged synchronously and perfectly blended according to the timeline. The goal of this step is to generate a smooth video content with coordinated and unified visual and auditory effects.

[0111] Based on the constructed video timeline, the text content, images, and video clips in the initial visual layout framework are arranged according to the order and duration of the timeline. This ensures that each visual element is displayed synchronously according to the timeline; for example, text content appears at a specific time point, while video clips and images play sequentially.

[0112] Integrate animation effects related to visual elements into the video timeline. Animation effects will be applied to text, images, video clips, etc., with the timing and duration of the animations arranged according to the timeline. For example, text slides in from left to right, a video clip fades in, and its start and end times are scheduled on the timeline.

[0113] Embed background music into the video timeline, ensuring the music starts synchronized with the video's beginning. Adjust the music's duration to match the video's rhythm, based on the timeline. If the video timeline exceeds the music's length, you can choose to loop the music or adjust its tempo and segment length.

[0114] Add sound effects to the timeline based on key events in the video (such as the triggering of animation effects or the switching of video segments). Ensure that the sound effects are perfectly synchronized with the corresponding visual events (such as image or video transitions) to enhance the presentation of important information in the video.

[0115] After integrating all elements, preview the entire video to check if each element plays in sync with the timeline. If any inconsistencies are found (such as delayed animations or out-of-sync audio), make minor adjustments according to the timeline to ensure smooth playback of the final integrated video.

[0116] Example Explanation: When generating an annual financial summary video, the system first integrates animated displays of the company's financial statements, background music, and textual content related to financial analysis, all arranged chronologically. The background music subtly enhances the video's atmosphere as the report data changes, while appropriate sound effects accompany the transitions of key performance indicators (such as revenue and net profit), allowing viewers to quickly grasp crucial information through a combination of visual and auditory elements. This integration method ensures the professionalism and fluency of the final video content, enabling it to effectively convey complex financial information within a short timeframe.

[0117] By integrating the initial visual layout framework, background music, and sound effects according to the video timeline, the final generated video content achieves a high degree of coordination and unity among all elements, avoiding playback issues caused by inconsistencies between different elements. This integration method makes the video presentation smoother and more professional, especially when displaying complex content such as financial data or industry analysis. The video can present a highly consistent visual and auditory experience, greatly improving the user's viewing experience and the efficiency of information comprehension.

[0118] This invention relates to the fields of artificial intelligence and fintech, and discloses a text-to-video method based on deep semantic analysis. The method extracts keywords and named entities from the text to be converted, analyzes the core theme and text structure, performs semantic analysis to generate semantic analysis results, retrieves matching images, video clips, sound effects, and background music based on the semantic analysis results, generates a preliminary visual layout framework through spatial layout, uses a deep learning model to generate animation effects and applies them to the layout scheme, constructs a video timeline, integrates the layout scheme, background music, and sound effects to generate complete video content, and finally outputs a video file. This invention accurately extracts keywords, named entities, core themes, and text structure from the text through semantic analysis, and intelligently matches images, video clips, sound effects, and background music that match the semantics, avoiding the problem of requiring a lot of manual intervention in existing technologies. Simultaneously, it generates adaptive animation effects through a deep learning model and applies them to the preliminary visual layout framework, ensuring the dynamic presentation of the video content and the coherence of the visual effects. After constructing the video timeline and integrating all elements, it automatically generates complete video content, significantly improving video production efficiency and ensuring professionalism and content consistency.

[0119] In one embodiment, S10 includes:

[0120] S101, Obtain the text to be converted, and use the natural language processing module to perform word segmentation on the text to be converted to generate word segmentation results;

[0121] S102, perform part-of-speech tagging on the word segmentation results to generate part-of-speech tagging results, so as to identify the part-of-speech category of each segmented word;

[0122] S103, Based on the part-of-speech tagging results, extract keywords from the word segmentation results;

[0123] S104, Analyze the relationship network between segmented words, determine the importance of segmented words based on their position and connectivity in the relationship network, and rank the keywords based on the importance.

[0124] S105, using the named entity recognition module, identify named entities including person names, place names and / or organization names from the text to be converted, and generate a list of named entities with entity category labels.

[0125] In this embodiment, the raw text data to be processed is acquired. This text can be a text file in formats such as financial reports or market analyses. The file input module can read the text to be converted via manual upload, API interface, or from a specified database. Multiple file formats are supported, including TXT, PPT, and PDF. The text preprocessing module will parse the file format, extract the text, and store it for subsequent operations.

[0126] Word segmentation is the process of dividing continuous text into the smallest semantic units. This is particularly crucial for understanding text in Chinese. Natural Language Processing (NLP) tools such as jieba or specialized tools tailored for the financial sector are used. During segmentation, the system identifies and segments each word in the text, generating a series of segmentation results. The system supports custom dictionaries to ensure the correct recognition of special terms such as financial terminology and company names.

[0127] Part-of-speech tagging (POS) adds a part-of-speech label (NOT), such as noun, verb, or adjective, to each segmented word. Natural language processing libraries (such as Stanford NLP and SpaCy) are used to perform POS tagging on the segmented results. The system uses a pre-trained model for tagging, generating a result table containing words and their corresponding POS tags. Specific financial terms and company names are accurately tagged using a dedicated POS tagging dictionary.

[0128] By analyzing the part-of-speech tagging results, the most representative and important keywords are extracted from the text. Keyword extraction algorithms (such as TF-IDF or TextRank) are used to extract keywords from the part-of-speech tagging results. Keyword extraction is based on the frequency of word occurrence in the text and its position in the lexical network. The system sorts the extracted keywords, ensuring that the most important keywords are listed first.

[0129] By constructing a semantic relationship network among segmented words, the strength and relationships between words are determined, and the importance of keywords is ranked. Graph algorithms (such as PageRank) are used to construct the lexical relationship network and analyze the connectivity of words in the segmentation results. The representativeness of words in the entire text is determined by calculating their importance within the relationship network. Keywords are then ranked by importance, with the most representative words being selected as the primary keywords.

[0130] Named Entity Recognition (NER) is used to identify specific categories of entities from text, such as person names, place names, company names, etc., and assign corresponding category labels to them. Use NER models (such as BERT, SpaCy, StanfordNER) to scan the text and identify the named entities in it. Determine the category of the entity (such as person name, company name, etc.) according to the context, and add the category label to it. Generate a list containing the named entities and their category labels, and the information in this list will be used in subsequent steps for semantic analysis and material matching.

[0131] In this embodiment, through steps of word segmentation,词性标注 (not clear in Chinese, please correct if wrong), keyword extraction, and named entity recognition, the system can accurately extract key information from financial text. This provides a semantic basis for subsequent semantic analysis and video generation, greatly improving the accuracy of text understanding, especially in the recognition of financial terms, reducing the workload of manual processing, and enhancing the efficiency and professionalism of video generation.

[0132] In one embodiment, the above S20 includes:

[0133] S201, preprocess the text to be converted, and the preprocessing includes removing stop words and punctuation marks;

[0134] S202, perform a theme analysis on the text to be converted based on a preset theme model, and extract key sentences and high-frequency words whose occurrence frequencies exceed a preset threshold in the text to be converted;

[0135] S203, generate one or more core themes based on the key sentences and high-frequency words;

[0136] S204, use a text structure analysis module to identify the title, subtitle, and / or paragraphs in the text to be converted;

[0137] S205, analyze the hierarchical relationship between the title, subtitle, and / or paragraphs, and generate a hierarchical text structure.

[0138] In this embodiment, before text analysis, the system will preprocess the text to be converted, clean up irrelevant content, such as stop words (such as "的", "了") and punctuation marks, which helps to improve the efficiency and accuracy of subsequent analysis. Use a predefined stop word list and regular expressions to clean the stop words and punctuation marks in the text. Save the cleaned text content for subsequent theme analysis and text structure analysis.

[0139] It should be noted that there seems to be an unclear part in the original text where "词性标注" is mentioned in Chinese. You may need to correct it if it's not accurate in the context.Through topic modeling, the system can identify the main topic of a text and extract representative sentences and high-frequency words, which help the system understand the core meaning of the text. Topic modeling algorithms, such as LDA (Latent Dirichlet Allocation) or LSA (Latent Semantic Analysis), are used to extract topics from the text content. Key sentences reflecting the topic are extracted from the text. Word frequency analysis algorithms are used to statistically analyze the words in the text and extract high-frequency words whose frequency exceeds a set threshold.

[0140] Based on key information extracted from topic analysis, the system can generate core themes describing the main content of the text. These themes provide a basis for subsequent visual content generation and matching. The system classifies and clusters the extracted key sentences and high-frequency words, and generates one or more core themes. The core themes will be tagged and saved for later use, such as in material library retrieval and visual element matching.

[0141] By analyzing text structure, the system can identify headings, subheadings, and paragraphs, helping to understand the text's hierarchical structure and providing logical organization for subsequent video generation. Using the text analysis module, structural elements in the text are identified, such as headings (H1-H3 levels), subheadings, and paragraph marks. The identified structural elements are then hierarchically labeled and saved as a tree structure.

[0142] By analyzing the hierarchical structure of the text, the system can organize the content into different levels, ensuring the logic and coherence of the content presentation during video generation. Based on the extracted titles, subtitles, and paragraphs, a hierarchical text structure is constructed. The system then associates these text structures with thematic information to provide a reference for the subsequent logical layout of the video content.

[0143] This embodiment, through text preprocessing, topic analysis, and structured hierarchical analysis, can accurately extract the core theme of the text and construct a clear text hierarchical structure by analyzing hierarchical information such as titles, subtitles, and paragraphs. This not only improves the accuracy of text semantic analysis but also provides logical organization for subsequent video content generation, ensuring that the generated video presents information more coherently and logically, avoiding the problems of content fragmentation and structural chaos in existing technologies.

[0144] In one embodiment, S30 includes:

[0145] S301, Analyze the association between the keywords and named entities, and determine the semantic relationship between the keywords and named entities based on the context and syntactic structure;

[0146] S302, Based on the hierarchical information of the text structure, analyze the semantic logic between the title, subtitle and / or paragraph;

[0147] S303, based on the matching analysis between the core theme and paragraph content, determines the degree of relevance between each paragraph and the core theme;

[0148] S304, Perform contextual analysis on named entities to determine the frequency of occurrence and contextual semantic meaning of named entities in the text to be converted;

[0149] S305. Based on the semantic relationship between the keywords and named entities, the degree of association between the paragraph and the core theme, and the frequency of occurrence and contextual semantic meaning of the named entities in the text to be converted, a comprehensive semantic analysis result is generated.

[0150] In this embodiment, semantic associations are established between keywords and named entities in the text to help the system understand the relationships between key concepts and entities. Context and syntactic structure are used to further enhance the accuracy of these associations. The system uses natural language processing techniques (such as dependency parsing and syntactic tree analysis) to analyze the position and structure of keywords and named entities in sentences. The system uses a context window to capture the co-occurrence frequency of keywords and named entities, as well as their part-of-speech relationships, to determine the semantic associations between them. A semantic graph is generated, displaying keywords and named entities in the form of a network graph, with connections between nodes representing semantic relationships.

[0151] By analyzing the relationships between headings, subheadings, and paragraphs, the system ensures it can accurately understand the logical structure of the text based on hierarchical information. Using a hierarchical text analysis module, headings, subheadings, and paragraphs are annotated and a tree structure is created. The system analyzes the contextual relationships between each level, identifying the connections between higher-level headings and lower-level content. Through continuity analysis between paragraphs, the system ensures logical consistency across all parts of the content.

[0152] To ensure that the content of paragraphs in the text is closely related to the core theme, the system needs to use semantic matching algorithms to determine the degree of relevance of each paragraph to the extracted core theme. Using a topic matching algorithm, the extracted core theme and paragraph content are analyzed for similarity. Based on the distribution and frequency of topic words, the system calculates the matching degree between paragraphs and the core theme, generating a paragraph-topic relevance score. The generated results are used to prioritize and arrange the content of paragraphs during video generation.

[0153] By analyzing the frequency of named entities in text and the context in which they appear, the system can understand the importance of named entities and the specific semantics they represent. The system scans the frequency of named entities throughout the text and calculates the specific semantic differences of entities in different contexts. Combining contextual information, such as keywords before and after the named entity and grammatical structure, it determines its specific meaning in different contexts. For example, "bank" can refer to "financial institution" or "riverbank" in different contexts. Based on the contextual analysis results, the system assigns a semantic label to each named entity and saves the results for subsequent visual element mapping.

[0154] Integrating all preceding semantic analysis results, a final comprehensive semantic analysis result is generated, ensuring that the semantic information of the text can be extracted and understood completely and accurately. The system summarizes the results from various analysis modules, including keywords, named entities, text structure, core themes, and paragraph matching. Through comprehensive semantic analysis, all analysis results are integrated into a detailed semantic graph, displaying the semantic relationships between elements such as keywords, named entities, paragraphs, and core themes. The generated comprehensive semantic analysis result provides a basis for subsequent visual element mapping.

[0155] This embodiment improves the accuracy of text semantic understanding through semantic association analysis of keywords and named entities, matching analysis of paragraphs and core themes, and contextual analysis of named entities. Especially when processing complex reports in the financial field, the system can accurately identify core themes, key statements, and important named entities, ensuring clear and coherent semantic logic in video generation and reducing the need for manual intervention. Simultaneously, the generated semantic analysis results provide high-quality data support for subsequent visual element matching, enhancing the professionalism and efficiency of video generation.

[0156] In one embodiment, S40 includes:

[0157] S401, based on the keyword information in the semantic analysis results, combined with the mapping rules between keywords and visual elements, retrieve images or video clips that match the keywords from the material library;

[0158] S402, Based on the named entity information in the semantic analysis results, and combined with the category and contextual semantic meaning of the named entity, retrieve images or video clips that match the named entity from the material library;

[0159] S403, based on the text structure hierarchy information in the semantic analysis results, retrieve image or video clips that match the text structure;

[0160] S404: Based on the core theme in the semantic analysis results, select a scene background image or scene background video clip that matches the core theme;

[0161] S405 selects background music and sound effects that match the emotional tone based on the semantic analysis results.

[0162] In this embodiment, the system identifies keywords (such as "financial market" and "stock price") in the text through semantic analysis. Based on pre-defined mapping rules between keywords and visual elements, it matches images or video clips semantically related to the keywords. This mapping can be achieved through keyword-visual element tags or topic relevance. The system extracts keywords from the semantic analysis results and combines them with mapping rules to match tagged visual elements in a media library. For example, the keyword "stock" might be associated with real-time stock price images from the financial market or video clips of exchange scenes. The retrieved images or video clips are sorted according to their relevance scores, with the most relevant elements selected first.

[0163] Named entities (such as "a bank" or "a company") represent different entity categories in different contexts. By combining the category of the named entity (such as financial institution or enterprise) with the semantic meaning of the context, the system can retrieve images or video clips that match the entity from the resource library. The system analyzes the category tags of the named entity, such as "institution," "company," and "person," and combines them with the semantic meaning of the context to filter out visual materials that match the category and context. For example, for the named entity "a bank," the system will retrieve images related to banking business, such as bank buildings and financial product advertisements. The selected images or video clips are prioritized to match the specific context in which the named entity exists.

[0164] The structure of text (such as titles, subheadings, and paragraphs) represents the hierarchy and logical relationships of the content. Based on this hierarchical information, the system retrieves images or video clips from the resource library that match each text level, used to display the text content in a hierarchical manner. The system performs resource retrieval based on the hierarchical information of the text structure. For example, for a title like "Market Analysis," the system can retrieve market dynamic charts or news clips related to that topic. The level of text structure hierarchy determines the importance and display order of visual materials. For example, an image matching the title might be displayed at the beginning of a video or during scene transitions, while paragraph content might be embedded in the middle of the video.

[0165] The core theme is the key content in the text, representing its central idea. The system selects appropriate background images or scene videos to emphasize this theme through semantic matching with the core theme. The system matches background materials from its resource library with the core theme. For example, if the core theme is "financial innovation," the system might choose a modern office environment or a technologically advanced scene as the background. The background image or video clip will be displayed as a primary visual element, enhancing the thematic coherence of the video.

[0166] The emotional tone of a text (such as positive, tense, or serious) can be identified through semantic analysis. Based on this tone, the system selects background music and sound effects from a resource library that match the emotional tone to create an audiovisual atmosphere that aligns with the text's emotional state. The system first identifies the text's emotional tone and then filters music or sound effects corresponding to that tone from the resource library. For example, for a "rigorous financial report," a somber and solemn background music might be chosen; while for positive content in "market analysis," upbeat sound effects might be selected. The selected music and sound effects are used throughout the entire video generation process, adjusted according to the emotional needs of different scenarios to enhance the viewer's emotional experience.

[0167] Example Explanation: Taking an annual financial market report as an example, the system first extracts the keywords "stock market volatility" and the named entity "a certain bank" from the text, and maps the keywords and named entities to corresponding visual elements, such as video clips of stock market trading scenes and images of bank buildings. Simultaneously, based on the core theme of the text, "market trends," the system matches relevant market trend charts as the main visual content of the report. Furthermore, through analysis of emotional tone, the system selects fast-paced background music to create a positive market outlook. This integrated process ensures that the generated video content accurately conveys the key information in the report while possessing high visual and auditory appeal.

[0168] This embodiment improves the accuracy and professionalism of text-to-video conversion by effectively mapping semantic analysis results with visual elements, sound effects, and background music. Through multi-level mapping of keywords, named entities, text structure, and core themes, the system can generate video content that better conforms to the semantic logic of the text, reducing manual user intervention and improving video generation efficiency. Simultaneously, by matching emotional tone with background music, the generated video can better convey the emotional atmosphere of the text, enhancing the viewing experience.

[0169] In one embodiment, S90 includes:

[0170] S901, based on the video timeline, synchronize the images, video clips, and animation effects in the preliminary visual layout framework with the video timeline to determine the display order, duration, and transition effects of the images, video clips, and animation effects;

[0171] S902, Adjust the start time and duration of the background music according to the video timeline;

[0172] S903, associate the sound effect with the corresponding animation effect, and set the trigger time and duration of the sound effect;

[0173] S904, integrate the visual layout framework, background music and sound effects to generate complete video content that has been synchronized.

[0174] In this embodiment, during video generation, the system coordinates and synchronizes the images, video clips, and animation effects in the initial visual layout framework according to a pre-constructed video timeline. This step ensures that each visual element is displayed in the correct order, time, and transition effects, avoiding conflicts or inconsistencies between visual elements. The timeline construction module defines the display time and order of each visual element based on the generated semantic analysis results and layout scheme. Images and video clips are played in the order set in the timeline, and the system automatically adjusts the display order according to the duration and transition effects of the elements. For example, image 1 is displayed at a certain moment, and then switched to video clip 1 at a later moment, with transition effects such as gradients applied during the switch. Animation effects are synchronized with images and video clips, such as overlaying text animation when an image appears, or adding special effects animation during video playback.

[0175] The playback time and duration of background music need to be adjusted according to the video timeline to ensure that the music is synchronized with the images and video clips, and to avoid mismatches between the background music and the video visuals. The background music is set to start playing at a specific point on the video timeline, and the system sets the duration of the background music based on the overall length and content of the video. The system automatically detects various scenes in the video and adjusts the timing of the music playback according to the emotional tone, ensuring that key content in the video is consistent with the rhythm of the music. For example, the rhythm of the background music will gradually increase when important images or scenes appear.

[0176] The system binds specific sound effects to animation effects in the video, ensuring that the sound effects trigger at the appropriate time and that their duration matches the length of the animation. Sound effect triggers set the playback timing based on the animation's timing. For example, at the start of an animation, a sound effect related to that animation is triggered. The system automatically adjusts the sound effect's duration to ensure consistency with the animation. For instance, an animation of an icon flying in is accompanied by a short sound effect, while there is no sound effect when the icon is stationary. The system selects a suitable sound effect library based on the characteristics of each animation effect and applies the relevant sound effects.

[0177] After all visual elements, background music, and sound effects have been timeline-synchronized, the system integrates all elements into a complete video, ensuring seamless transitions between different parts and ultimately outputting high-quality video content. The integration module, based on the video timeline, background music, and sound effect synchronization settings, combines all processed images, video clips, animations, background music, and sound effects into a continuous video stream. The system integrates the order and duration of all elements and synchronizes the audio and video content to ensure error-free and lag-free playback. The integrated, complete video content is output in a standard playable format, featuring full visual effects, animations, sound effects, and background music.

[0178] This embodiment, through the steps described above, achieves effective synchronization and integration of the initial visual layout framework, background music and sound effects with the video timeline, ensuring the coherence and high-quality presentation of the video content. It accurately controls the playback order and timing of images, video clips, animation effects, background music, and sound effects, resulting in a final video content with strong visual impact and auditory synchronization. This reduces the workload of manually adjusting the timeline and materials, significantly improving the efficiency and quality of video production.

[0179] In one embodiment, after S90 above, the following is also included:

[0180] S905 renders the integrated video content, combining all visual and audio elements frame by frame;

[0181] S906 uses a video encoding module to compress the rendered video content;

[0182] S907, generate the final video file and output the video file in a preset video format.

[0183] In this embodiment, rendering involves combining all time-synchronized visual and audio elements frame-by-frame to generate the final video content. During rendering, the system integrates all elements, including images, video clips, animation effects, background music, and sound effects, into a continuous video stream. The rendering module starts from the first point in time on the timeline and processes all visual and audio elements frame-by-frame, ensuring that each frame of video accurately matches its corresponding background music and sound effects. The system renders frame-by-frame according to the set resolution and frame rate, combining the images, videos, animation effects, text, background music, and sound effects contained in each frame. After rendering is complete, the output is a complete video file containing all visual and audio elements, which is used for subsequent compression and encoding.

[0184] After rendering, video files are typically quite large, so encoding and compression are necessary to reduce their size for storage and transmission. The video encoding module is responsible for compressing the rendered video content while maintaining video quality. The system inputs the rendered raw video content into the encoding module and compresses it according to a preset encoding standard (such as H.264, H.265, etc.). During encoding, the system analyzes the changes between frames in the video and reduces file size by removing redundant data. For example, for video segments with infrequent scene changes, the encoder can reduce redundant frames using inter-frame compression techniques. The compressed video will occupy less storage space while maintaining a certain level of image and audio quality, making it suitable for subsequent playback and distribution.

[0185] After compression, the system generates the final video file and outputs it in a user-specified preset video format. This step ensures the generated video file is compatible with various playback devices or platforms. The system encapsulates the compressed video file into a standard playable video format, such as MP4, AVI, MKV, etc., with the specific format selected based on user needs or playback platform requirements. During the encapsulation process, the system formats the video and audio to ensure the final output file contains complete audio and video streams and can be played normally on common media players. The output video file will be saved to the user-specified path or transferred to an external storage platform via an API interface for further use or distribution.

[0186] This embodiment effectively optimizes video production and playback performance by rendering, compressing, and outputting the integrated video content frame by frame into a preset format. The system can efficiently process a large number of visual and audio elements, ensuring smooth and high-quality generated video content. Simultaneously, through video encoding and compression, it significantly reduces the size of video files, improving storage and transmission efficiency. The final output video format is diverse, ensuring strong compatibility and adapting to the needs of various playback devices and platforms.

[0187] In one embodiment, a text-to-video device based on deep semantic analysis is provided, which corresponds one-to-one with the text-to-video method based on deep semantic analysis in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the text-to-video device based on deep semantic analysis of the present invention. The modules include: text processing module 10, text analysis module 20, semantic analysis module 30, material retrieval module 40, layout generation module 50, layout feature analysis module 60, animation effect generation module 70, timeline construction module 80, and video generation module 90. Detailed descriptions of each functional module are as follows:

[0188] Text processing module 10 is used to acquire the text to be converted, perform word segmentation and part-of-speech tagging on the text to be converted, and extract keywords and named entities from the text to be converted;

[0189] Text analysis module 20 is used to analyze the core theme and text structure of the text to be converted, the text structure including title, subtitle and / or paragraph;

[0190] Semantic analysis module 30 is used to perform semantic analysis based on the keywords, named entities, text structure and core topics, and generate semantic analysis results;

[0191] The material retrieval module 40 is used to establish a mapping relationship between semantics and visual elements based on the semantic analysis results, and to retrieve images, video clips, sound effects and background music that match the semantic analysis results from the material library;

[0192] The layout generation module 50 is used to spatially arrange the text content corresponding to the keywords and named entities with images and video clips based on the semantic analysis results, and generate a preliminary visual layout framework.

[0193] The layout feature analysis module 60 is used to analyze the layout content features of the preliminary visual layout framework using a deep learning model.

[0194] Animation effect generation module 70 is used to generate animation effects based on the layout content features and apply the animation effects to the text content and visual elements in the preliminary visual layout framework.

[0195] Timeline building module 80 is used to build a video timeline based on a preliminary visual layout framework with applied animation effects, background music, and sound effects;

[0196] The video generation module 90 is used to integrate the initial visual layout framework, background music and sound effects according to the video timeline to generate integrated complete video content.

[0197] In one embodiment, the text processing module 10 is specifically used for:

[0198] The text to be converted is obtained, and the natural language processing module is used to perform word segmentation on the text to be converted to generate word segmentation results;

[0199] The word segmentation results are then tagged with part-of-speech tags to generate part-of-speech tagging results, in order to identify the part-of-speech category of each segmented word;

[0200] Based on the part-of-speech tagging results, keywords are extracted from the word segmentation results;

[0201] The relationship network between segmented words is analyzed, and the importance of segmented words is determined based on their position and connectivity in the relationship network. Keywords are then ranked based on their importance.

[0202] The named entity recognition module is used to identify named entities, including person names, place names and / or organization names, from the text to be converted, and to generate a list of named entities with entity category labels.

[0203] In one embodiment, the text analysis module 20 is specifically used for:

[0204] The text to be converted is preprocessed, including the removal of stop words and punctuation marks;

[0205] The text to be converted is analyzed based on a preset topic model, and key sentences and high-frequency words with a frequency exceeding a preset threshold are extracted from the text to be converted.

[0206] Based on the key statements and high-frequency words, generate one or more core themes;

[0207] The text structure analysis module is used to identify the headings, subheadings, and / or paragraphs in the text to be converted.

[0208] Analyze the hierarchical relationships between the titles, subtitles, and / or paragraphs to generate a hierarchical text structure.

[0209] In one embodiment, the semantic analysis module 30 is specifically used for:

[0210] Analyze the association between the keywords and named entities, and determine the semantic relationship between the keywords and named entities based on the context and syntactic structure;

[0211] Based on the hierarchical information of the text structure, analyze the semantic logic between the title, subtitle and / or paragraph;

[0212] Based on the matching analysis between the core theme and paragraph content, the degree of relevance between each paragraph and the core theme is determined;

[0213] Contextual analysis of named entities is performed to determine the frequency of occurrence and contextual semantic meaning of named entities in the text to be converted;

[0214] Based on the semantic relationship between the keywords and named entities, the degree of association between the paragraph and the core topic, and the frequency of occurrence and contextual semantic meaning of the named entities in the text to be converted, a comprehensive semantic analysis result is generated.

[0215] In one embodiment, the material retrieval module 40 is specifically used for:

[0216] Based on the keyword information in the semantic analysis results, and combined with the mapping rules between keywords and visual elements, images or video clips that match the keywords are retrieved from the material library.

[0217] Based on the named entity information in the semantic analysis results, combined with the category and contextual semantic meaning of the named entity, images or video clips that match the named entity are retrieved from the material library;

[0218] Based on the hierarchical information of text structure in the semantic analysis results, retrieve image or video clips that match the text structure;

[0219] Based on the core themes in the semantic analysis results, select scene background images or scene background video clips that match the core themes;

[0220] Based on the emotional tone in the semantic analysis results, background music and sound effects that match the emotional tone are selected.

[0221] In one embodiment, the video generation module 90 is specifically used for:

[0222] Based on the video timeline, the images, video clips, and animation effects in the preliminary visual layout framework are synchronized with the video timeline to determine the display order, duration, and transition effects of the images, video clips, and animation effects.

[0223] Adjust the start time and duration of the background music according to the video timeline;

[0224] Associate the sound effects with the corresponding animation effects, and set the trigger time and duration of the sound effects;

[0225] The visual layout framework, background music, and sound effects are integrated to generate complete video content that has been synchronized.

[0226] In one embodiment, the video generation module 90 is specifically used for:

[0227] Render the integrated video content, combining all visual and audio elements frame by frame;

[0228] The rendered video content is compressed using a video encoding module.

[0229] Generate the final video file and output the video file in a preset video format.

[0230] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a text-to-video method based on deep semantic analysis on the server side.

[0231] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the user-side functions or steps of a text-to-video method based on deep semantic analysis.

[0232] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0233] Obtain the text to be converted, perform word segmentation and part-of-speech tagging on the text to be converted, and extract keywords and named entities from the text to be converted;

[0234] Analyze the core theme and text structure of the text to be converted, including headings, subheadings, and / or paragraphs;

[0235] Semantic analysis is performed based on the keywords, named entities, text structure, and core themes to generate semantic analysis results.

[0236] Based on the semantic analysis results, a mapping relationship between semantics and visual elements is established, and images, video clips, sound effects, and background music that match the semantic analysis results are retrieved from the material library.

[0237] Based on the semantic analysis results, the text content corresponding to the keywords and named entities is spatially arranged with images and video clips to generate a preliminary visual layout framework.

[0238] The layout content features of the preliminary visual layout framework are analyzed using a deep learning model.

[0239] Animation effects are generated based on the layout content features, and the animation effects are applied to the text content and visual elements in the initial visual layout framework.

[0240] Based on the initial visual layout framework with applied animation effects, background music, and sound effects, construct the video timeline;

[0241] Based on the video timeline, the initial visual layout framework, background music, and sound effects are integrated to generate the integrated complete video content.

[0242] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0243] Obtain the text to be converted, perform word segmentation and part-of-speech tagging on the text to be converted, and extract keywords and named entities from the text to be converted;

[0244] Analyze the core theme and text structure of the text to be converted, including headings, subheadings, and / or paragraphs;

[0245] Semantic analysis is performed based on the keywords, named entities, text structure, and core themes to generate semantic analysis results.

[0246] Based on the semantic analysis results, a mapping relationship between semantics and visual elements is established, and images, video clips, sound effects, and background music that match the semantic analysis results are retrieved from the material library.

[0247] Based on the semantic analysis results, the text content corresponding to the keywords and named entities is spatially arranged with images and video clips to generate a preliminary visual layout framework.

[0248] The layout content features of the preliminary visual layout framework are analyzed using a deep learning model.

[0249] Animation effects are generated based on the layout content features, and the animation effects are applied to the text content and visual elements in the initial visual layout framework.

[0250] Based on the initial visual layout framework with applied animation effects, background music, and sound effects, construct the video timeline;

[0251] Based on the video timeline, the initial visual layout framework, background music, and sound effects are integrated to generate the integrated complete video content.

[0252] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0253] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0254] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0255] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for converting text to video based on deep semantic analysis, characterized in that, The method comprises the following steps: obtaining a text to be converted, performing word segmentation processing and part-of-speech tagging on the text to be converted, and extracting keywords and named entities from the text to be converted; analyzing the core theme and the text structure of the text to be converted, wherein the text structure comprises a title, a subtitle and / or a paragraph; performing semantic analysis based on the keywords, the named entities, the text structure and the core theme to generate a semantic analysis result; establishing a mapping relationship between semantics and visual elements based on the semantic analysis result, and retrieving images, video clips, sound effects and background music matching the semantic analysis result from a material library; based on the semantic analysis result, performing spatial layout of the text content corresponding to the keywords and the named entities and the images and video clips to generate a preliminary visual layout framework; analyzing the layout content features of the preliminary visual layout framework by using a deep learning model; generating animation effects based on the layout content features, and applying the animation effects to the text content and the visual elements in the preliminary visual layout framework; based on the preliminary visual layout framework, the background music and the sound effects to which the animation effects have been applied, constructing a video timeline; integrating the preliminary visual layout framework, the background music and the sound effects according to the video timeline to generate a complete video content after integration.

2. The deep semantic analysis based text-to-video method of claim 1, wherein, The method comprises the following steps: obtaining a text to be converted, performing word segmentation processing and part-of-speech tagging on the text to be converted, and extracting keywords and named entities from the text to be converted, comprising: obtaining a text to be converted, and performing word segmentation processing on the text to be converted by using a natural language processing module to generate a word segmentation result; performing part-of-speech tagging on the word segmentation result to generate a part-of-speech tagging result, so as to identify the part-of-speech categories of each word segmentation vocabulary; based on the part-of-speech tagging result, extracting keywords from the word segmentation result; analyzing the relationship network among the word segmentation vocabularies, determining the importance of the word segmentation vocabularies based on their positions and connection degrees in the relationship network, and sorting the keywords based on the importance; 3. The deep semantic analysis based text-to-video method of claim 1, wherein, using a named entity recognition module to identify named entities containing names of persons, places and / or organizations from the text to be converted, and generating a named entity list with entity category labels. The method comprises the following steps: obtaining a text to be converted, performing word segmentation processing and part-of-speech tagging on the text to be converted, and extracting keywords and named entities from the text to be converted, comprising: performing preprocessing on the text to be converted, wherein the preprocessing comprises removing stop words and punctuation marks; performing theme analysis on the text to be converted based on a preset theme model, extracting key sentences in the text to be converted and high-frequency words with an occurrence frequency exceeding a preset threshold; based on the key sentences and the high-frequency words, generating one or more core themes; 4. The deep semantic analysis based text-to-video method of claim 1, wherein, using a text structure analysis module to identify a title, a subtitle and / or a paragraph in the text to be converted; analyzing the hierarchical relationship among the title, the subtitle and / or the paragraph to generate a hierarchical text structure. The method comprises the following steps: obtaining a text to be converted, performing word segmentation processing and part-of-speech tagging on the text to be converted, and extracting keywords and named entities from the text to be converted, comprising: performing semantic analysis based on the keywords, the named entities, the text structure and the core theme to generate a semantic analysis result, comprising: analyzing the association between the keywords and the named entities, and determining the semantic relationship between the keywords and the named entities based on context and syntactic structure; analyzing the semantic logic between the title, sub-title and / or paragraph based on the hierarchical information of the text structure; determining the degree of association between each paragraph and the core topic based on the matching analysis of the core topic and the paragraph content; performing context analysis on the named entities to determine the frequency of occurrence and the contextual semantic meaning of the named entities in the text to be converted; generating a comprehensive semantic analysis result based on the semantic relationship between the keywords and the named entities, the degree of association between the paragraphs and the core topic, and the frequency of occurrence and the contextual semantic meaning of the named entities in the text to be converted. 5.The deep semantic analysis based text-to-video method of claim 1, wherein, According to the semantic analysis result, a mapping relationship between semantics and visual elements is established, and images, video clips, sound effects and background music matching the semantic analysis result are retrieved from a material library, including: Based on the keyword information in the semantic analysis result, and combining the mapping rules of keywords and visual elements, images or video clips matching the keywords are retrieved from the material library; Based on the named entity information in the semantic analysis result, and combining the category and contextual semantic meaning of the named entity, images or video clips matching the named entity are retrieved from the material library; Based on the hierarchical information of the text structure in the semantic analysis result, images or video clips matching the text structure are retrieved; Based on the core topic in the semantic analysis result, scene background images or scene background video clips matching the core topic are selected; Based on the emotional tone in the semantic analysis result, background music and sound effects matching the emotional tone are selected.

6. The deep semantic analysis based text-to-video method as claimed in claim 1, wherein, According to the video timeline, the preliminary visual layout framework, background music and sound effects are integrated to generate a complete video content after integration, including: According to the video timeline, the images, video clips and animation effects in the preliminary visual layout framework are synchronized with the video timeline to determine the display order, duration and transition effect of the images, video clips and animation effects; According to the video timeline, the start time and duration of the background music are adjusted; The sound effects are associated with the corresponding animation effects, and the trigger time and duration of the sound effects are set; The visual layout framework, background music and sound effects are integrated to generate a complete video content after synchronization processing.

7. The deep semantic analysis based text-to-video method of claim 1, wherein, After integrating the preliminary visual layout framework, background music and sound effects according to the video timeline to generate a complete video content after integration, it further includes: Rendering the integrated video content, combining all visual elements and audio elements frame by frame; Using a video encoding module to compress the rendered video content; Generating a final video file and outputting the video file in a preset video format.

8. A text-to-video apparatus based on deep semantic analysis, characterized by, The text-to-video device based on deep semantic analysis includes: a text processing module for obtaining a text to be converted, performing word segmentation processing and part-of-speech tagging on the text to be converted, and extracting keywords and named entities from the text to be converted; The text analysis module is configured to analyze a core topic of the text to be converted and a text structure including a title, a subtitle, and / or a paragraph. The semantic analysis module is configured to perform semantic analysis based on the keywords, the named entities, the text structure, and the core topic, and generate a semantic analysis result. The material retrieval module is configured to establish a mapping relationship between semantics and visual elements based on the semantic analysis result, and retrieve images, video clips, sound effects, and background music matching the semantic analysis result from a material library. The layout generation module is configured to perform spatial layout of literal content corresponding to the keywords and the named entities and images and video clips based on the semantic analysis result, and generate a preliminary visual layout framework. The layout feature analysis module is configured to analyze layout content features of the preliminary visual layout framework by using a deep learning model. The animation effect generation module is configured to generate animation effects based on the layout content features, and apply the animation effects to literal content and visual elements in the preliminary visual layout framework. The timeline construction module is configured to construct a video timeline based on the preliminary visual layout framework to which the animation effects have been applied, background music, and sound effects. The video generation module is configured to integrate the preliminary visual layout framework, the background music, and the sound effects according to the video timeline, and generate complete video content after integration.

9. A computer device, comprising: The computer device includes a memory, a processor, and a deep semantic analysis based text-to-video program stored on the memory and executable on the processor. When the deep semantic analysis based text-to-video program is executed by the processor, the steps of the deep semantic analysis based text-to-video method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The storage medium stores a deep semantic analysis based text-to-video program. When the deep semantic analysis based text-to-video program is executed by the processor, the steps of the deep semantic analysis based text-to-video method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Video generation method, video generation device, electronic equipment and storage medium

    CN116744055A

  • Method and system for automatic generation of video from non-visual information

    WO2016157150A1