Method for quickly generating video from text

Through text recognition and classification algorithms, combined with pre-stored video libraries and consistency assessment, the problem of inconsistency between video and text in existing technologies is solved, and the effect of high-quality and fast video generation is achieved.

CN120654108APending Publication Date: 2025-09-16JIANGSU YUNSHEN INTELLIGENT SYST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510761546.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies lack the extraction and classification of entities, contexts, relationships, and events in text data, resulting in inconsistencies between the generated video and the text data. They also lack image-text consistency assessment, resulting in deviations between the generated video and the text content.

Method used

Use text recognition algorithms to identify entities and context in text data, extract relationships and events, classify them using preset classifiers, query pre-stored video libraries, perform image and text consistency assessments, and splice video materials in the order of text data.

Benefits of technology

The consistency between the generated video and the original text data is improved, ensuring that the generated video conforms to the entities, background, relationships and events described in the text, and achieving rapid generation of high-quality videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654108A_ABST
    Figure CN120654108A_ABST
Patent Text Reader

Abstract

The invention discloses a method for quickly generating a video from a text, which comprises the following steps of: inputting text data of a video required to be generated through a text write-in port; identifying entities and backgrounds in the text data by adopting a text identification algorithm, and extracting relationships and events in the text data; performing classification in a preset classifier according to entities, backgrounds, relationships and events in the text data to obtain a text classification result; querying a video material of a corresponding classification in a pre-stored video library according to the text classification result according to the index relationship; performing image-text consistency evaluation on the abstract keywords corresponding to the obtained video material and entities, backgrounds, relationships and events in the text data; and splicing the video materials according to the sequence of the text data to obtain a corresponding video. According to the invention, through extraction and classification of entities, backgrounds, relationships and events of text data, a creation video is rapidly generated; through image-text consistency evaluation, the consistency of the generated spliced video and the original text data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of text-to-video generation, and in particular relates to a method for rapidly generating video from text. Background Art

[0002] With the rise in popularity of video sharing platforms, more and more video creators are turning to using their own text summaries to create videos. This is crucial for better understanding text data and mining entities, context, relationships, and events. With the advancement of artificial intelligence in natural language processing and computer vision, and the continuous advancement of computing power, the technology for generating videos or images from text is becoming increasingly mature.

[0003] In a Chinese authorized invention patent application numbered CN202410446159.X, a method for generating video from text based on image-guided video editing is disclosed. The method comprises: obtaining the base text for generating a target video, performing semantic analysis of the base text from multiple angles using a text analysis module to generate a multi-dimensional text feature vector; generating a base image using the text feature vector and a fine-tuned text-to-image model; generating a base video using the text feature vector and an existing text-to-video model; and generating a target video using the base image, base video, and base text as inputs through image- and text-guided video editing. This invention leverages the advantages of high quality and high resolution of text-generated images and the excellent temporal modeling of text-generated videos, achieving high-quality and high-resolution video generation from text through high-quality image and text-guided video editing.

[0004] The flaw in the existing patent is that while the patented technology uses high-quality images and text to guide video editing, enabling the generation of high-quality, high-resolution videos from text, it lacks the extraction and classification of entities, context, relationships, and events within the text data, as well as the steps for assessing image-text consistency. This results in the generated and spliced ​​images not being consistent with the text data content, resulting in certain deviations. Furthermore, the order of the text and video may be disrupted. Summary of the Invention

[0005] Existing methods for generating videos from text lack the extraction and classification of entities, backgrounds, relationships, and events in text data, and also lack the step of evaluating the consistency of images and text, resulting in the problem that the generated and spliced ​​images do not conform to the content of the text data. The present invention provides a method for quickly generating videos from text.

[0006] In order to achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0007] A method for quickly generating a video from text, comprising the steps of:

[0008] S1. Input the text data required to generate the video through the text write port;

[0009] S2. Use text recognition algorithms to identify entities and contexts in text data and extract relationships and events from text data;

[0010] S3. Classify the text data using a preset classifier based on entities, context, relationships, and events to obtain text classification results.

[0011] S4, searching the pre-stored video library for video materials of corresponding classification based on the text classification result according to the index relationship;

[0012] S5. Perform a text-image consistency assessment on the corresponding summary keywords of the acquired video materials and the entities, backgrounds, relationships, and events in the text data;

[0013] S6. Splice the video materials in the order of the text data to obtain the corresponding video.

[0014] Furthermore, in step S2, the text recognition algorithm for the text data entity is a named entity recognition method based on CRF. The specific recognition process of the text data entity includes:

[0015] Preprocess the collected text data;

[0016] Perform feature extraction on text data;

[0017] Input the extracted features into the pre-trained CRF-based named entity recognition model;

[0018] The output recognition result is the entity string and the location of the text data.

[0019] Furthermore, the collected text data is preprocessed including text data cleaning, word segmentation and tagging.

[0020] Furthermore, the specific recognition algorithm formula in the CRF-based named entity recognition model is:

[0021]

[0022] Among them, Y is the label sequence, X is the input sequence, and L k (Y,X) is the loss function for a specific category, λ k is the weight of the corresponding category, and Z(Y) is the normalization factor of the classifier.

[0023] Furthermore, the relationship extraction method in text data:

[0024] Locate the text data between two entities in a chapter, paragraph or sentence in the text data based on the entity recognition result and extract it;

[0025] Then, the text data between the two entities is extracted and input one by one according to the text data relationship fields pre-stored in the database for retrieval and matching;

[0026] If the pre-stored text data relationship field matches the same field in the text data between two entities, it is considered a relationship between the two entities;

[0027] If the pre-stored text data relationship field does not match the same field in the text data between the two entities, the matching corresponding relationship field is searched in the text data before and after the entity in the same sentence / paragraph / chapter text data;

[0028] If the query matches the same field, it is considered a relationship between the two entities;

[0029] If the query still fails to match the same fields, feedback will be given to the front-end user, and the relationship fields will be set according to the user.

[0030] Furthermore, the event extraction method in the text data adopts the K-means clustering algorithm. The specific steps of the event extraction method in the text data are:

[0031] Initialize the preset event corresponding to the entity, relationship and key field as the condensation point, and the condensation point is the event;

[0032] According to the principle of proximity, the corresponding entities, relationships and key fields with similar meanings are gathered to the condensation point;

[0033] Calculate the center position of each initial classification;

[0034] Re-cluster using the calculated center position;

[0035] The cycle is repeated until the position of the condensation point converges, that is, the condensation point is determined to be a specific event.

[0036] Furthermore, the context is the cultural context in which the event in which the entity occurs or the scene in which the entity is located. The context recognition algorithm adopts a key field query matching method;

[0037] The process of identifying background information using the key field query matching method is as follows:

[0038] Split the text data into paragraphs or sentences, and obtain the split paragraphs and sentences;

[0039] Perform data cleaning and word segmentation on the split paragraphs and sentences;

[0040] Import the obtained non-repeated segmented words into the background field storage database to query whether there is a corresponding background field;

[0041] If it exists, the location of the background field in the text data and the string composition of the background field are recorded;

[0042] If it does not exist, feedback will be sent to the terminal user to confirm whether the word segmentation is a background field; if it is not a background field, the next word segmentation will be imported into the background field storage database to query whether there is a corresponding background field. If it is a background field, the user will specify the background field in the background field storage database to correspond to the background image.

[0043] Furthermore, a decision tree classifier is selected as the classifier for text data;

[0044] The main node represents each decision tree category label, and the sub-nodes are the specific labels of entities, contexts, relationships, and events;

[0045] The branch nodes under the entity, context, relationship, and event labels are specific category description labels;

[0046] When the category description label of the branch node under the entity, background, relationship and event labels of the text data matches the video label corresponding to each video material stored in the database with the highest matching degree, it is regarded as the text classification result.

[0047] Furthermore, the detailed steps of step S5 include:

[0048] According to the labels of the corresponding summary keywords of the video material, the visual semantic similarity is calculated in reverse order among the entity, background, relationship and event database labels in the text data;

[0049] If the labels of the generated video material corresponding to the summary keywords are ranked in the top r labels in any of the entity, background, relationship and event database labels of the text data, then the image and text are considered consistent;

[0050] If the labels of the generated video material corresponding to the summary keywords are not ranked among the top r labels in any of the entity, background, relationship and event database labels of the text data, it is considered that the image and text are inconsistent.

[0051] Furthermore, the evaluation index for visual semantic similarity calculation is R-Precision;

[0052] The calculation formula of R-Precision is:

[0053] Where R is the total number of documents in the dataset that are relevant to the query;

[0054] Specifically, in addition to generating the true labels for the video material, additional labels are randomly sampled from the dataset. The cosine similarity between the true labels for the video material and the text embeddings of the entity, context, relation, and event database labels for each text data point is then calculated, and the labels are sorted in descending order of similarity. A success is considered if the true label for the generated image ranks among the top r labels.

[0055] Furthermore, the text data is spliced ​​into the video material in the order of the text data positions of entities, backgrounds, relationships and events located during the query and classification process (the splicing here can be completed according to the pixel points of entities, backgrounds, relationships and events), and the chapters, paragraphs or sentences in the text data are finally spliced ​​in sequence to obtain the final generated video.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] By extracting, classifying, and locating entities, context, relationships, and events in text data, the system obtains multiple labels for the text data. These labels are then entered into a pre-stored video library to search for video materials of corresponding categories, enabling rapid creation of creative videos. Furthermore, through image-text consistency assessment, the generated spliced ​​video is ensured to conform to the entities, context, relationships, and events described in the text data, further improving the consistency of the generated creative video with the original text data. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is an overall flow chart of a method for quickly generating videos from text in an embodiment of the present invention;

[0059] Figure 2 This is a flowchart of the specific identification of text data entities in an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and drawings. The contents mentioned in the embodiments are not intended to limit the present invention.

[0061] like Figure 1 As shown, this embodiment provides a method for quickly generating a video from text, including the steps of:

[0062] S1. Input the text data required to generate the video through the text write port;

[0063] S2. Use text recognition algorithms to identify entities and contexts in text data and extract relationships and events from text data;

[0064] S3. Classify the text data using a preset classifier based on entities, context, relationships, and events to obtain text classification results.

[0065] S4, searching the pre-stored video library for video materials of corresponding classification based on the text classification result according to the index relationship;

[0066] S5. Perform a text-image consistency assessment on the corresponding summary keywords of the acquired video materials and the entities, backgrounds, relationships, and events in the text data;

[0067] S6. Splice the video materials in the order of the text data to obtain the corresponding video.

[0068] like Figure 2 As shown, in step S2, the text recognition algorithm of the text data entity is a named entity recognition method based on CRF. The specific recognition process of the text data entity includes:

[0069] S2010, preprocessing the collected text data;

[0070] S2011, extracting features from text data;

[0071] S2012, inputting the extracted features into a pre-trained CRF-based named entity recognition model;

[0072] S2013. Output the recognition result as the entity string and the location of the text data.

[0073] Preprocessing the collected text data involves cleaning, word segmentation, and tagging. This tagging facilitates locating the chapter, paragraph, or sentence position of each character in the text data, including entities, context, relationships, and events. It also facilitates the subsequent splicing of video footage according to the order of entities, context, relationships, and events in the text data.

[0074] The specific recognition algorithm formula in the CRF-based named entity recognition model is:

[0075]

[0076] Among them, Y is the label sequence, X is the input sequence, and L k (Y,X) is the loss function for a specific category, λ k is the weight of the corresponding category, and Z(Y) is the normalization factor of the classifier.

[0077] Relation extraction methods in text data:

[0078] Locate the text data between two entities in a chapter, paragraph or sentence in the text data based on the entity recognition result and extract it;

[0079] Then, we input the extracted text data between the two entities one by one according to the pre-stored text data relationship fields in the database for search and matching. Big data analysis shows that the relationship between two entities in the text data is more likely to appear between the two entities than before or after them. Therefore, we prioritize determining whether the description field of the entity relationship exists between the two entities.

[0080] If the pre-stored text data relationship field matches the same field in the text data between two entities, it is considered a relationship between the two entities;

[0081] If the pre-stored text data relationship field does not match the same field in the text data between the two entities, the matching corresponding relationship field is searched in the text data before and after the entity in the same sentence / paragraph / chapter text data;

[0082] If the query matches the same field, it is considered a relationship between the two entities;

[0083] If the query still fails to match the same fields, feedback will be given to the front-end user, and the relationship fields will be set according to the user.

[0084] The event extraction method in text data adopts K-means clustering algorithm. The specific steps of the event extraction method in text data are as follows:

[0085] Initialize the preset event corresponding to the entity, relationship and key field as the condensation point, and the condensation point is the event;

[0086] According to the principle of proximity, the corresponding entities, relationships and key fields with similar meanings are gathered to the condensation point;

[0087] Calculate the center position of each initial classification;

[0088] Re-cluster using the calculated center position;

[0089] This process repeats until the clustering point converges, confirming that the clustering point is the event. The occurrence of a historical event or a real-time hot topic usually locates the corresponding entity person, entity object, and corresponding relationship. By confirming the occurrence of the entity person, entity object, and corresponding relationship, the K-means clustering algorithm is used to determine that the text data describes the event.

[0090] The context is the cultural background of the event in which the entity occurs or the scene in which the entity is located. The background recognition algorithm uses the key field query matching method;

[0091] The process of identifying background information using the key field query matching method is as follows:

[0092] Split the text data into paragraphs or sentences, and obtain the split paragraphs and sentences;

[0093] Perform data cleaning and word segmentation on the split paragraphs and sentences;

[0094] Import the obtained non-repeated segmented words into the background field storage database to query whether there is a corresponding background field;

[0095] If it exists, the location of the background field in the text data and the string composition of the background field are recorded;

[0096] If it doesn't exist, feedback is sent to the end user to confirm whether the segmented word is a background field. If not, the next segmented word is imported into the background field storage database to query whether a corresponding background field exists. If it is a background field, the user specifies the background field and stores the corresponding background image in the background field storage database. Obtaining the background field of the current entity facilitates switching image backgrounds when stitching together each image segment. This also enables stitching and scene switching of the same entity under different backgrounds.

[0097] Select the decision tree classifier as the classifier for text data;

[0098] The main node represents each decision tree category label, and the sub-nodes are the specific labels of entities, contexts, relationships, and events;

[0099] The branch nodes under the entity, context, relationship, and event labels are specific category description labels;

[0100] The text classification result is determined by the highest degree of match between the category description labels of the branch nodes under the entity, background, relationship, and event labels of the text data and the corresponding video labels of each video material stored in the database. Here, a deep learning algorithm can be used to train the dataset to obtain the optimal proportion of entity, background, relationship, and event labels. When the matching degrees of entity, background, relationship, and event labels differ, the matching degree is calculated by multiplying them by the optimal scaling factor and summing them. A comprehensive comparison is then made of the matching degrees of the decision tree category labels under different videos. This determines the selection of the matching video material that most closely matches the entity, background, relationship, and event labels of the current text data.

[0101] The detailed steps of step S5 include:

[0102] According to the labels of the corresponding summary keywords of the video material, the visual semantic similarity is calculated in reverse order among the entity, background, relationship and event database labels in the text data;

[0103] If the labels of the generated video material corresponding to the summary keywords are ranked in the top r labels in any of the entity, background, relationship and event database labels of the text data, then the image and text are considered consistent;

[0104] If the labels of the generated video material corresponding to the summary keywords are not ranked among the top r labels in any of the entity, background, relationship and event database labels of the text data, it is considered that the image and text are inconsistent.

[0105] The evaluation index for visual semantic similarity calculation is R-Precision;

[0106] The calculation formula of R-Precision is:

[0107] Where R is the total number of documents in the dataset that are relevant to the query;

[0108] Specifically, in addition to generating the true labels for the video material, additional labels are randomly sampled from the dataset. The cosine similarity between the true labels for the video material and the text embeddings of the entity, context, relation, and event database labels for each text data point is then calculated, and the labels are sorted in descending order of similarity. A success is considered if the true label for the generated image ranks among the top r labels.

[0109] The video materials are spliced ​​in the order of the text data positions of entities, backgrounds, relationships and events located during the query and classification process (the splicing here can be completed according to the pixel points of entities, backgrounds, relationships and events). At the same time, the chapters, paragraphs or sentences in the text data are finally spliced ​​in sequence to obtain the final generated video.

[0110] Compared with the prior art, the present invention has the following beneficial effects:

[0111] By extracting, classifying, and locating entities, context, relationships, and events in text data, the system obtains multiple labels for the text data. These labels are then entered into a pre-stored video library to search for video materials of corresponding categories, enabling rapid creation of creative videos. Furthermore, through image-text consistency assessment, the generated spliced ​​video is ensured to conform to the entities, context, relationships, and events described in the text data, further improving the consistency of the generated creative video with the original text data.

[0112] The above describes in detail a method for quickly generating a video from text provided by this application. The description of the specific embodiments is only intended to help understand the method and core concept of this application. It should be noted that for those skilled in the art, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A method for quickly generating video from text, characterized in that: Including steps: S1. Input the text data required to generate the video through the text write port; S2. Use text recognition algorithms to identify entities and contexts in text data and extract relationships and events from text data; S3. Classify the text data using a preset classifier based on entities, context, relationships, and events to obtain text classification results. S4, searching the pre-stored video library for video materials of corresponding classification based on the text classification result according to the index relationship; S5. Perform a text-image consistency assessment on the corresponding summary keywords of the acquired video materials and the entities, backgrounds, relationships, and events in the text data; S6. Splice the video materials in the order of the text data to obtain the corresponding video.

2. The method for quickly generating a video from text according to claim 1, characterized in that: In step S2, the text recognition algorithm for text data entities is a named entity recognition method based on CRF. The specific recognition process of text data entities includes: Preprocess the collected text data; Perform feature extraction on text data; Input the extracted features into the pre-trained CRF-based named entity recognition model; The output recognition result is the entity string and the location of the text data.

3. The method for quickly generating a video from text according to claim 2, characterized in that: The specific recognition algorithm formula in the CRF-based named entity recognition model is: Among them, Y is the label sequence, X is the input sequence, and L k (Y,X) is the loss function for a specific category, λ k is the weight of the corresponding category, and Z(Y) is the normalization factor of the classifier.

4. The method for quickly generating a video from text according to claim 3, characterized in that: Relation extraction methods in text data: Locate the text data between two entities in a chapter, paragraph or sentence in the text data based on the entity recognition result and extract it; Then, the text data between the two entities is extracted and input one by one according to the text data relationship fields pre-stored in the database for retrieval and matching; If the pre-stored text data relationship field matches the same field in the text data between two entities, it is considered a relationship between the two entities; If the pre-stored text data relationship field does not match the same field in the text data between the two entities, the matching corresponding relationship field is searched in the text data before and after the entity in the same sentence / paragraph / chapter text data; If the query matches the same field, it is considered a relationship between the two entities; If the query still fails to match the same fields, feedback will be given to the front-end user, and the relationship fields will be set according to the user.

5. The method for quickly generating a video from text according to claim 4, characterized in that: The event extraction method in text data adopts K-means clustering algorithm. The specific steps of the event extraction method in text data are as follows: Initialize the preset event corresponding to the entity, relationship and key field as the condensation point, and the condensation point is the event; According to the principle of proximity, the corresponding entities, relationships and key fields with similar meanings are gathered to the condensation point; Calculate the center position of each initial classification; Re-cluster using the calculated center position; Repeat the cycle until the condensation point position converges.

6. The method for quickly generating a video from text according to claim 5, characterized in that: The context is the cultural background of the event in which the entity occurs or the scene in which the entity is located. The background recognition algorithm uses the key field query matching method; The process of identifying background information using the key field query matching method is as follows: Split the text data into paragraphs or sentences, and obtain the split paragraphs and sentences; Perform data cleaning and word segmentation on the split paragraphs and sentences; Import the obtained non-repeated segmented words into the background field storage database to query whether there is a corresponding background field; If it exists, the location of the background field in the text data and the string composition of the background field are recorded; If it does not exist, feedback is sent to the end user to confirm whether the word segmentation is a background field; If it is not a background field, the next word segmentation is performed to import the background field storage database to query whether there is a corresponding background field. If it is a background field, the user specifies the background field in the background field storage database corresponding to the background image.

7. The method for quickly generating a video from text according to claim 6, characterized in that: Select the decision tree classifier as the classifier for text data; The main node represents each decision tree category label, and the sub-nodes are the specific labels of entities, contexts, relationships, and events; The branch nodes under the entity, context, relationship, and event labels are specific category description labels; When the category description label of the branch node under the entity, background, relationship and event labels of the text data matches the video label corresponding to each video material stored in the database with the highest matching degree, it is regarded as the text classification result.

8. The method for quickly generating a video from text according to claim 7, characterized in that: The detailed steps of step S5 include: According to the labels of the corresponding summary keywords of the video material, the visual semantic similarity is calculated in reverse order among the entity, background, relationship and event database labels in the text data; If the labels of the generated video material corresponding to the summary keywords are ranked in the top r labels in any of the entity, background, relationship and event database labels of the text data, then the image and text are considered consistent; If the labels of the generated video material corresponding to the summary keywords are not ranked among the top r labels in any of the entity, background, relationship and event database labels of the text data, it is considered that the image and text are inconsistent.

9. The method for quickly generating a video from text according to claim 8, characterized in that: The evaluation index for visual semantic similarity calculation is R-Precision; The calculation formula of R-Precision is: Where R is the total number of documents in the dataset that are relevant to the query.

10. The method for quickly generating a video from text according to claim 1, characterized in that: The video material is spliced ​​in the order of the text data locations of entities, backgrounds, relations and events located during the query and classification process. At the same time, the chapters, paragraphs or sentences in the text data are finally spliced ​​in sequence to obtain the final generated video.

Citation Information

Patent Citations

  • A text-to-video method based on image-guided video editing

    CN118037898B