Rich media document assistance generation apparatus
By designing a rich media text generation auxiliary device, intelligent methods are used to extract, sort, and synthesize materials, solving the problem of complex generation processes in existing technologies and achieving efficient generation of rich media comprehensive texts.
Patent Information
- Application Number
- CN202211632138.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-12-19
AI Technical Summary
Existing technologies lack auxiliary means to improve the efficiency of rich media text generation, resulting in a complex and labor-intensive video information generation process.
Design a rich media text generation auxiliary device, including a material extraction module, a topic sorting module, a semantic retrieval module, a structured data text generation module, an illustration recommendation module, and a video synthesis module, which uses intelligent means to assist in the generation of rich media comprehensive texts.
It enables rapid and accurate comprehensive description of thematic events, improves the efficiency of rich media text generation, simplifies the material selection process, and generates rich media comprehensive texts with audio, streaming media, and subtitles.
Smart Images

Figure CN116227463B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural language processing technology, and in particular relates to a rich media text generation aid device. Background Technology
[0002] In today's society, various hot topics can emerge daily. Researching these thematic events is an important task, and more and more journalists and self-media creators are engaging in writing comprehensive articles. Traditional thematic event research primarily presents findings in text form, supplemented by images. However, humans are visual creatures; video information is far richer in meaning and more memorable than text or images.
[0003] However, converting text into rich media video scripts requires tasks such as collecting video materials, writing narration, recording narration, and video editing before it can be officially released. The process is very complex and the workload is huge. Existing technologies lack auxiliary means to improve the efficiency of rich media script generation. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art by providing a rich media text generation auxiliary device that uses intelligent means to help users efficiently generate high-quality rich media comprehensive texts and quickly and accurately describe the theme event in a comprehensive manner.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] A rich media text generation aid device, the device comprising:
[0007] The material extraction module is used to extract writing materials for generating rich media documents from the received original materials;
[0008] The topic sorting module is used to cluster the writing materials by topic, extract the keywords of the clustered writing materials as the topic of each cluster of articles, form a topic list, score the topic list using a pre-built user profile, and sort the topic list based on the scoring results.
[0009] The semantic retrieval module is used to obtain the semantic vector of the received text information, and retrieve semantically similar text segments in the writing material based on the semantic vector;
[0010] The structured data text generation module is used to convert structured data obtained through the intelligent analysis engine into natural language text.
[0011] The illustration recommendation module is used to recommend illustrations that match a threshold based on the semantic information of the input text.
[0012] The video compositing module is used to generate videos based on the input text.
[0013] Furthermore, the material extraction module includes a paragraph extraction submodule, a summary extraction submodule, a key sentence detection submodule, and a knowledge extraction submodule;
[0014] The paragraph extraction submodule clusters several neighboring paragraphs based on paragraph semantic similarity and uses a central sentence extraction model to extract the central sentence of paragraphs of the same category.
[0015] The abstract extraction submodule uses a generative abstracting model to generate a short text abstract for each original writing text.
[0016] The golden sentence detection submodule divides the text into sentences according to punctuation marks, and then scores the sentences using a preset scoring model. Sentences with scores exceeding a threshold are identified as golden sentences.
[0017] The knowledge extraction submodule extracts the knowledge contained in the text to form triples or knowledge graphs.
[0018] Furthermore, the key phrase detection submodule includes a binary classification model, which is obtained through supervised training using labeled positive and negative samples.
[0019] Furthermore, the material extraction module also includes an image description generation submodule and a video description generation submodule;
[0020] The image description generation submodule uses a multimodal image-text model to generate image description information based on image content;
[0021] The video description generation submodule generates descriptive information based on typical video features.
[0022] Furthermore, the extraction of keywords from the clustered writing materials as the theme of each clustered article specifically includes:
[0023] The abstract extraction model is used to extract the abstract content of the original writing material, and then the central sentence extraction model is used to extract the central sentence or central phrase of the abstract content, which is used as the theme of the article in a cluster.
[0024] Furthermore, the intelligent analysis engine includes a target analysis engine and an event analysis engine. The target analysis engine analyzes the statistical patterns of the target with the target as the center, while the intelligent analysis engine analyzes the background and development trend of the event with the event as the center. Finally, the intelligent analysis engine outputs the analysis conclusions in the form of structured data.
[0025] Furthermore, the device also includes a text continuation module, which uses a sequence model to continue a new text after the end of the original text. The sequence model is trained in an unsupervised manner, using a context masking method on the original text to predict the next text based on the preceding text, and is trained automatically.
[0026] Furthermore, the device also includes a text rewriting module, which rewrites the input text according to set style control variables using a sequence model.
[0027] Furthermore, the device also includes an intelligent summarization module, which generates a summary based on the semantic information of the input content.
[0028] Furthermore, the device also includes an intelligent detection module and an audit and evaluation module;
[0029] The intelligent detection module includes word and phrase proofreading, punctuation proofreading, syntax proofreading, common sense proofreading, and factual proofreading of the input manuscript;
[0030] The review and evaluation module includes quantitative scoring of the current document's fluency, common sense compliance, and factual relevance.
[0031] Furthermore, the video synthesis module includes a text transcription submodule, a narration speech synthesis submodule, and a subtitle synthesis submodule;
[0032] The text transcription submodule is used to transcribe text into a specific style according to a style-controllable sequence model;
[0033] The narration speech synthesis submodule is used to automatically generate audio files based on the input video narration;
[0034] The subtitle synthesis submodule is used to automatically break the video narration into sentences according to punctuation, and to control the duration of the subtitles in the video based on the length of each narration audio sentence.
[0035] Furthermore, the video synthesis module includes a semantic-level image retrieval submodule and a semantic-level video segment retrieval submodule;
[0036] The semantic-level image retrieval submodule is used to automatically retrieve the most matching image from the image library based on the semantic information of each input explanatory text.
[0037] The semantic-level video segment retrieval submodule is used to automatically retrieve the most matching video segment from the video library based on the semantic information of each input narration text.
[0038] Furthermore, the device also includes a tag generation module and a publishing channel recommendation module;
[0039] The tag generation module generates tags based on video features and semantic information of text;
[0040] The publishing channel recommendation module recommends publishing channels based on user profiles.
[0041] The beneficial effects of this invention are as follows:
[0042] (1) This invention can help users quickly process original writing materials that are not used much into writing materials such as golden sentences, knowledge, and multimedia description information that can be used directly, saving the tedious material selection process.
[0043] (2) The manuscript writing process of this invention proposes a structured data text generation method based on intelligent analysis engine. It can conduct in-depth analysis of the development context, current status and development trend of major events through various event-related intelligent analysis engines and output structured conclusions.
[0044] (3) The video synthesis module proposed in this invention uses multiple artificial intelligence models to synthesize rich media short videos with voice, streaming media and subtitles based on the completed comprehensive text, organically combining text, pictures and videos to generate rich media comprehensive texts and improve creation efficiency. Attached Figure Description
[0045] Figure 1 This is a structural block diagram of the rich media document auxiliary generation device provided in the embodiments of the present invention;
[0046] Figure 2 This is a schematic diagram of the rich media integrated document generation process according to an embodiment of the present invention;
[0047] Figure 3 This is a schematic diagram of the material import process according to an embodiment of the present invention;
[0048] Figure 4 This is a schematic diagram of the content planning process according to an embodiment of the present invention;
[0049] Figure 5 This is a schematic diagram of the document writing process according to an embodiment of the present invention;
[0050] Figure 6 This is a schematic diagram of the intelligent proofreading process according to an embodiment of the present invention;
[0051] Figure 7 This is a schematic diagram of the video synthesis process according to an embodiment of the present invention. Detailed Implementation
[0052] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0053] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Converting text into rich media video scripts requires tasks such as collecting video footage, writing narration, recording narration, and video editing before it can be officially released. The process is very complex and involves a huge amount of work.
[0055] To address the aforementioned technical problems, the following embodiments of a rich media text generation aid device of the present invention are proposed.
[0056] Example 1
[0057] Reference Figure 1 ,like Figure 1 The diagram shown is a structural block diagram of the rich media integrated document generation auxiliary device provided in this embodiment. The device specifically includes the following structures:
[0058] The material extraction module is used to extract writing materials for generating rich media documents from the received original materials.
[0059] The material extraction module includes a paragraph extraction submodule, a summary extraction submodule, a key sentence detection submodule, and a knowledge extraction submodule.
[0060] The paragraph extraction submodule clusters several neighboring paragraphs based on their semantic similarity and uses a central sentence extraction model to extract the central sentence of each paragraph in the same category. In this embodiment, the central sentence extraction model can be TextRank, BERTsum, etc. The input is a text, which is first segmented into sentences, then intelligently sorted according to the semantic features of the sentences, and finally outputs the sentence with the highest ranking as the central sentence of the text.
[0061] The abstract extraction submodule utilizes a generative summarization model to generate a short text summary for each original text. In this embodiment, the generative summarization model can be BART, UNILM, etc. The input text is first encoded, then the model outputs the summary encoding word by word in an autoregressive manner, and finally the decoder outputs the text content of the summary.
[0062] The "golden sentence detection" submodule segments the text into sentences according to punctuation marks, and then scores the sentences using a preset scoring model. Sentences with scores exceeding a threshold are identified as "golden sentences." In this embodiment, the scoring model can be a regression-based machine learning model, such as linear regression based on LMS or nonlinear regression based on Volterra series. Taking a text as input, the scoring first encodes the input text, and the encoded text is then processed by a corresponding neural network to output the score corresponding to the text. The specific calculation method of the neural network varies depending on the algorithm used.
[0063] The knowledge extraction submodule extracts knowledge contained in the text to form triples or knowledge graphs.
[0064] In this embodiment, the quote detection submodule includes a binary classification model, which is obtained through supervised training using labeled positive and negative samples.
[0065] As one implementation method, the material extraction module in this embodiment further includes an image description generation submodule and a video description generation submodule;
[0066] The image description generation submodule utilizes a multimodal image-text model to generate image description information based on image content. In this embodiment, the multimodal image-text model can be Cross-Stream, VILBERT, etc. Taking an image as input, the model performs object recognition, scene recognition, and other functions to understand the image, thereby outputting descriptive text for the image.
[0067] The video description generation submodule generates descriptive information based on typical video features.
[0068] The topic sorting module is used to cluster writing materials by topic, extract keywords from the clustered writing materials as the topic of each cluster of articles, form a topic list, score the topic list using pre-built user profiles, and sort the topic list based on the scoring results.
[0069] Specifically, a summary extraction model is used to extract the summary content of the original writing material, and then a central sentence extraction model is used to extract the central sentence or central phrase of the summary content, which serves as the theme of a clustered article.
[0070] The semantic retrieval module is used to obtain the semantic vector of the received text information and retrieve semantically similar text segments in the writing material based on the semantic vector.
[0071] The structured data text generation module is used to convert structured data obtained through the intelligent analysis engine into natural language text.
[0072] The illustration recommendation module is used to recommend illustrations that match a threshold based on the semantic information of the input text.
[0073] The video compositing module is used to generate videos based on the input text.
[0074] As one implementation method, the video synthesis module in this embodiment includes a semantic-level image retrieval submodule and a semantic-level video segment retrieval submodule.
[0075] The semantic-level image retrieval submodule is used to automatically retrieve the most matching image from the image library based on the semantic information of each input explanatory text.
[0076] The semantic-level video clip retrieval submodule is used to automatically retrieve the most matching video clips from the video library based on the semantic information of each input narration text.
[0077] In this embodiment, the video synthesis module utilizes multiple artificial intelligence models to synthesize rich media short videos with audio, streaming media, and subtitles based on the completed comprehensive text. It organically combines text, images, and videos to generate rich media comprehensive texts, thereby improving creation efficiency.
[0078] As one implementation method, the rich media text generation device in this embodiment further includes a tag generation module and a publishing channel recommendation module. The tag generation module generates tags based on video features and semantic information of the text; the publishing channel recommendation module recommends publishing channels based on user profiles.
[0079] As one implementation method, the rich media integrated text generation device provided in this embodiment also includes a text continuation module. The text continuation module uses a sequence model to continue a new text after the end of the original text. The sequence model is trained in an unsupervised manner and uses a context masking method on the original text to predict the context based on the preceding text and train automatically.
[0080] As one implementation method, the rich media text generation device in this embodiment also includes a text rewriting module, which rewrites the input text according to the set style control variables through a sequence model.
[0081] As one implementation method, the rich media text generation device in this embodiment further includes an intelligent summary module, which generates a summary based on the semantic information of the input content.
[0082] As one implementation method, the rich media text generation device in this embodiment further includes an intelligent detection module and an audit and evaluation module.
[0083] The intelligent detection module includes proofreading of the input text for words, punctuation, syntax, common sense, and facts. The review and evaluation module includes quantitative scoring of the current text's fluency, common sense compliance, and factual accuracy.
[0084] The rich media text generation device provided in this embodiment can help users quickly process previously underutilized writing materials into readily usable writing materials such as catchy phrases, knowledge, and multimedia descriptive information, eliminating the tedious material selection process.
[0085] Example 2
[0086] This embodiment provides an application example of the rich media document generation auxiliary device provided in the foregoing embodiments. For example... Figure 2 The diagram shown is a schematic of the rich media document generation process provided in this embodiment. The method specifically includes the following steps:
[0087] The first step is to import the materials, refer to Figure 3 ,like Figure 3 The diagram shown is a schematic of the material import process in this embodiment. The specific steps are as follows:
[0088] The writing materials include not only crawled news texts, but also paragraph head sentences, key phrases, knowledge, image descriptions, and video descriptions processed through natural language processing and multimodal data processing. The original writing materials can be crawled from designated websites using a web crawler to extract news, images, and videos. After the original writing text materials are imported into the system, they undergo natural language processing, including paragraph extraction, summary extraction, key phrase detection, and knowledge extraction modules. The summary extraction module uses a generative summarization model to generate a short summary for each piece of original writing text, storing the news text and its corresponding summary in a news summary database. The paragraph extraction module first uses paragraph clustering to cluster neighboring paragraphs based on semantic similarity, and then uses a head sentence extraction model to extract the head sentences of paragraphs in the same category, storing the clustered paragraphs and their corresponding head sentences in a paragraph database. The key phrase detection module first segments the text according to punctuation marks such as periods, exclamation marks, and question marks, then identifies sentences with complete descriptive elements, excellent descriptive style, and clear content as key phrases and stores them in a key phrase database. The key phrase detection uses a binary classification model, employing supervised training with labeled positive and negative samples. The knowledge extraction module extracts knowledge from text and stores it in a knowledge base in the form of triples or knowledge graphs. The image description generation module processes crawled images and image data from news articles, utilizing a multimodal graph-text model to generate a one-sentence description based on image content. This description serves as an important reference for subsequent writing material retrieval and image titles; the processed image and description are stored in an image library. The video description generation module processes crawled video data, generating a one-sentence description based on typical video features; the processed video data and its description are stored in a video library. Writing materials from the news summary library, paragraph library, key phrase library, knowledge base, image library, and video library are crucial inputs for generating rich media integrated texts.
[0089] The material import process comprehensively utilizes various natural language processing and multimedia material processing models to directly transform previously underutilized writing materials into readily usable quotes, knowledge, and multimedia descriptive information. All steps are completed in the background of material import, without interfering with the user's writing time. The content planning process, based on a vast amount of writing material, assists users in quickly identifying writing topics, constructing outlines, and clarifying the writing structure.
[0090] The second step is content planning, referring to... Figure 4 ,like Figure 4 The diagram shown is a schematic of the content planning process in this embodiment. The specific steps are as follows:
[0091] The topic selection process begins with unsupervised clustering of trending events based on the writing source materials. Using an unsupervised text clustering model, the original writing materials are clustered by topic. After calculating cluster centers, the writing materials located at the cluster centers are used as input. A summary extraction model extracts the summary content of the original writing materials, and a central sentence extraction model extracts the central sentence or phrase, which serves as the topic of an article in one cluster. This process is repeated to calculate the topics of all articles in all clusters, forming a topic list. Further, user profiles built based on user reading, editing, and writing behaviors are used to score the topics in the topic list. The scoring model is a regression model trained in a supervised manner. Based on the scoring results, the topic list is sorted and recommended to writing users, who can then select a writing topic based on the recommendations.
[0092] The input to the outline development process is a set of articles from clusters corresponding to the selected topic. First, based on this set, articles relevant to the current writing task are further filtered. Then, corresponding article summaries from a news summary database and topic sentences from a paragraph database are used to score the articles. The scoring model is a regression model trained in a supervised manner. Based on the high-scoring summaries and topic sentences, phrase-style outline titles are generated and recommended to the writing user. The user then selects from the recommended outline title list to develop their writing outline.
[0093] The third step is manuscript writing, referring to... Figure 5 ,like Figure 5 The diagram shown illustrates the document writing process in this embodiment. The specific steps are as follows:
[0094] The manuscript writing process comprises several parts: intelligent semantic retrieval, structured data text generation, inspiration recommendation, text continuation, text rewriting, illustration recommendation, image title generation, and intelligent summarization. These parts are not strictly ordered; each provides assistance to the writing user as needed. Intelligent semantic retrieval supports semantic-level multi-modal writing material retrieval. Using a text as input, it leverages a large-scale pre-trained language model to obtain the text's semantic vector. Through this semantic vector, it can retrieve semantically similar paragraphs, summaries, key phrases, knowledge, images, and other writing materials from the writing material library, assisting users in manuscript writing. Structured data text generation converts structured data into natural language text. The structured data comes from an intelligent analysis engine, which is divided into a target analysis engine and an event analysis engine. The target analysis engine focuses on the target and analyzes its statistical patterns, while the event analysis engine focuses on the event, analyzing its background and development trends. The intelligent analysis engine ultimately outputs its analysis conclusions in the form of structured data. The data2text module converts these structured analysis conclusions into natural language to further assist users in writing. Text continuation takes a text as input and uses a sequence-to-sequence (S2S) model to add a new text after the end of the original text. The S2S model is trained unsupervised, using a context mask on the original text, and predicts the following text based on the preceding text, thus training automatically. Text rewriting takes a text and style control variables as input and uses an S2S model to rewrite the input text according to a specified style. Illustration recommendation takes a text as input and recommends the most relevant illustrations to the writer based on the semantic information of the input text. Illustration recommendation is implemented based on a multimodal pre-trained model. Illustration title generation takes the illustration as input and generates a one-sentence text description as the illustration title based on a multimodal pre-trained model. Inspiration recommendation supports users when they encounter writing bottlenecks. It takes already written content as input and, based on the semantic information of the written content, integrates intelligent semantic retrieval, text continuation, text rewriting, and illustration recommendation functions to provide writing assistance. Intelligent summarization takes the user's written content as input and generates a summary based on the semantic information of the written content. The outputs of each module are combined to form a final comprehensive document.
[0095] The manuscript writing process proposes a structured data text generation method based on intelligent analysis engines. It can conduct in-depth analysis of the development context, current status and development trend of major events through various event-related intelligent analysis engines and output structured conclusions.
[0096] The fourth step is intelligent proofreading, referring to... Figure 6 ,like Figure 6 The diagram shown is a schematic of the intelligent verification process in this embodiment. The specific steps are as follows:
[0097] This embodiment mainly includes three steps: intelligent detection, review and evaluation, and revision confirmation. Intelligent detection includes functions such as word and phrase proofreading, punctuation proofreading, syntax proofreading, common sense proofreading, and factual proofreading. Word and phrase proofreading includes typo and sensitive word proofreading, which can identify typos and sensitive words in the initial draft through an intelligent model and provide revision suggestions. Punctuation proofreading can identify incorrect punctuation in the initial draft and provide revision suggestions. Syntax proofreading can identify grammatical errors and awkward sentences in the initial draft and provide semantically equivalent sentences as revision suggestions. Common sense proofreading can identify sentences in the initial draft that violate common sense based on a background common sense database and provide the corresponding correct common sense from the database. Factual proofreading can identify sentences in the initial draft that contradict facts based on a background fact database and prompt the writer to revise them. After intelligent detection, a comprehensive review and quantitative evaluation of the current manuscript's fluency, common sense compliance, and factual relevance are conducted. Based on the detection and evaluation results, the user makes revisions and confirms the final draft.
[0098] The intelligent proofreading process not only identifies and corrects traditional errors in words, punctuation, and syntax, but also conducts in-depth analysis of text content based on language models, identifying counterintuitive and counterfactual errors, effectively preventing the release of erroneous information.
[0099] Step 5: Video compositing, refer to... Figure 7 ,like Figure 7 The diagram shown is a schematic of the video synthesis process in this embodiment. The specific steps are as follows:
[0100] Video narration is generated based on the written manuscript using a text transcription model;
[0101] The video narration generates subtitles, and an audio track is generated based on a speech synthesis model;
[0102] The narration text is segmented into sentences based on punctuation marks;
[0103] Sentence segmentation-based semantic information retrieval of relevant image and video clips;
[0104] Images, videos, and text are sequentially combined to generate a video-text presentation, ultimately producing a rich media presentation with narration, video images, and subtitles.
[0105] This step utilizes intelligent methods to generate short videos based on the generated text. Video synthesis mainly consists of a text processing workflow and a multimedia processing workflow. The text processing workflow includes text transcription, narration speech synthesis, and subtitle synthesis. Text transcription adapts the text to a style suitable for video narration, based on a style-controllable S2S model. Narration speech synthesis automatically generates fluent audio files based on the written video narration. Subtitle synthesis automatically segments the video narration according to punctuation and controls the duration of subtitles in the video based on the length of each narration audio sentence. The multimedia processing workflow mainly consists of semantic-level image retrieval and semantic-level video clip retrieval. Based on the semantic information of each narration text, it automatically retrieves the most matching images or video clips from image and video libraries, with the playback length of the images or video clips automatically matching the length of the narration text. The rich media short video synthesis module merges the narration audio, video clips, images, and subtitles to generate a single rich media short video file.
[0106] In this embodiment, the video synthesis process utilizes multiple artificial intelligence models to synthesize rich media short videos with audio, streaming media, and subtitles based on the completed comprehensive text. This organically combines text, images, and videos to generate rich media comprehensive texts, improving the reader's information acquisition efficiency.
[0107] The sixth step, document publication, includes the following: short video tag generation, distribution channel recommendation, publication preview, and publication confirmation. Short video tag generation automatically generates matching tags based on the short video content and the comprehensive text document, allowing subscribers to quickly index content according to their reading needs. Distribution channel recommendation utilizes an intelligent recommendation model to provide targeted distribution channel recommendations based on user profiles. Publication preview allows previewing the generated comprehensive text document and rich media short video; publication can proceed only after confirmation.
[0108] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A rich media document assistance generation apparatus characterized by comprising: The device includes: The material extraction module is used to extract writing materials for generating rich media documents from the received original materials; The material extraction module includes a paragraph extraction submodule, a summary extraction submodule, a key sentence detection submodule, and a knowledge extraction submodule; The paragraph extraction submodule clusters several neighboring paragraphs based on paragraph semantic similarity and uses a central sentence extraction model to extract the central sentence of paragraphs of the same category. The abstract extraction submodule uses a generative abstracting model to generate a short text abstract for each original writing text. The golden sentence detection submodule divides the text into sentences according to punctuation marks, and then scores the sentences using a preset scoring model. Sentences with scores exceeding a threshold are identified as golden sentences. The knowledge extraction submodule extracts the knowledge contained in the text to form triples or knowledge graphs; The topic sorting module is used to cluster the writing materials by topic, extract keywords from the clustered writing materials as the topic of each clustered article, form a topic list, score the topic list using a pre-built user profile, and sort the topic list based on the scoring results; use a summary extraction sub-model to extract the summary content of the original writing materials, and then use a central sentence extraction model to extract the central sentence or central phrase of the summary content as the topic of a clustered article; The semantic retrieval module is used to obtain the semantic vector of the received text information, and retrieve semantically similar text segments in the writing material based on the semantic vector; The structured data text generation module is used to convert structured data obtained through the intelligent analysis engine into natural language text. The illustration recommendation module is used to recommend illustrations that match a threshold based on the semantic information of the input text. The video compositing module is used to generate videos based on the input text.
2. The rich media document auxiliary generation apparatus of claim 1, wherein The key phrase detection submodule includes a binary classification model, which is obtained through supervised training using labeled positive and negative samples.
3. The rich media document auxiliary generation apparatus of claim 1, wherein The material extraction module also includes an image description generation submodule and a video description generation submodule; The image description generation submodule uses a multimodal image-text model to generate image description information based on image content; The video description generation submodule generates descriptive information based on typical video features.
4. The rich media document auxiliary generation apparatus of claim 1, wherein The keywords extracted from the clustered writing materials, serving as the theme of each clustered article, specifically include: The abstract extraction model is used to extract the abstract content of the original writing material, and then the central sentence extraction model is used to extract the central sentence or central phrase of the abstract content, which is used as the theme of the article in a cluster.
5. The rich media document auxiliary generation apparatus of claim 1, wherein The intelligent analysis engine includes a target analysis engine and an event analysis engine. The target analysis engine focuses on the target to analyze its statistical patterns, while the intelligent analysis engine focuses on the event to analyze its background and development trend. Finally, the intelligent analysis engine outputs the analysis conclusions in the form of structured data.
6. The rich media text generation device as described in claim 1, characterized in that, The device also includes a text continuation module, which uses a sequence model to continue a new text after the end of the original text. The sequence model is trained in an unsupervised manner, using a context masking method on the original text to predict the next text based on the preceding text, and is trained automatically.
7. The rich media text generation device as described in claim 1, characterized in that, The device also includes a text rewriting module, which rewrites the input text according to set style control variables using a sequence model.
8. The rich media text generation aid as described in claim 1, characterized in that, The device also includes an intelligent summarization module, which generates a summary based on the semantic information of the input content.
9. The rich media text generation device as described in claim 1, characterized in that, The device also includes an intelligent detection module and an audit and evaluation module; The intelligent detection module includes word and phrase proofreading, punctuation proofreading, syntax proofreading, common sense proofreading, and factual proofreading of the input manuscript; The review and evaluation module includes quantitative scoring of the current document's fluency, common sense compliance, and factual relevance.
10. The rich media text generation aid as described in claim 1, characterized in that, The video synthesis module includes a text transcription submodule, a narration speech synthesis submodule, and a subtitle synthesis submodule; The text transcription submodule is used to transcribe text into a specific style according to a style-controllable sequence model; The narration speech synthesis submodule is used to automatically generate audio files based on the input video narration; The subtitle synthesis submodule is used to automatically break the video narration into sentences according to punctuation, and to control the duration of the subtitles in the video based on the length of each narration audio sentence.
11. The rich media text generation aid as described in claim 1, characterized in that, The video synthesis module includes a semantic-level image retrieval submodule and a semantic-level video segment retrieval submodule; The semantic-level image retrieval submodule is used to automatically retrieve the most matching image from the image library based on the semantic information of each input explanatory text. The semantic-level video segment retrieval submodule is used to automatically retrieve the most matching video segment from the video library based on the semantic information of each input narration text.
12. The rich media text generation aid as described in claim 1, characterized in that, The device also includes a tag generation module and a publishing channel recommendation module; The tag generation module generates tags based on video features and semantic information of text; The publishing channel recommendation module recommends publishing channels based on user profiles.
Citation Information
Patent Citations
Automatic text writing method and system
CN108563620A
Intelligent image-text to video conversion method and system based on video structured data
CN115272533A