Video generation method, device, electronic device, and storage medium
By using the target parameter library and video library to generate voice text and subtitle text and synthesize videos, the problem of semantic inconsistency in short video generation is solved, and cost-effective video generation is achieved.
Patent Information
- Application Number
- CN202310147489.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-02-09
AI Technical Summary
The existing short video generation technology is difficult to effectively combine voice text and video content, resulting in inconsistent semantics and high cost of generated videos.
By using the target parameter library to generate voice text and subtitle text, combining the target video library to generate initial video, and synthesis processing, we ensure the matching and consistency of audio information and video content.
The generated video semantic consistency and integrity are achieved, reducing the cost of video generation.
Smart Images

Figure CN116156248B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video processing technology, and in particular to the fields of natural language processing, text-to-speech conversion, and video synthesis technology. More specifically, the present disclosure provides a video generation method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] Short videos are a way of disseminating Internet content. With the popularization of mobile terminals and the development of Internet technology, short videos have gradually gained favor from major platforms due to their fragmentation and easy dissemination. Various applications developed based on short videos, such as short video marketing, are also widely used in e-commerce, new media operations, public relations activities and other scenarios. Summary of the Invention
[0003] The present disclosure provides a video generation method, apparatus, electronic device, storage medium, and program product.
[0004] According to one aspect of the present disclosure, a video generation method is provided, comprising: in response to receiving a video generation request, generating speech text and subtitle text respectively using a target parameter library, wherein the video generation request includes a target object entity word of a target object, and the target parameter library is related to the target object; converting the speech text to obtain audio information; generating an initial video based on the duration of the audio information using a target video library, wherein the target video library is related to the target object; and synthesizing the subtitle text, the audio information and the initial video to generate a target video.
[0005] According to another aspect of the present disclosure, a video generation device is provided, including: a first generation module for generating voice text and subtitle text respectively using a target parameter library in response to receiving a video generation request, wherein the above-mentioned video generation request includes a target object entity word of the target object, and the above-mentioned target parameter library is related to the above-mentioned target object; a conversion module for converting the above-mentioned voice text to obtain audio information; a second generation module for generating an initial video using a target video library according to the duration of the above-mentioned audio information, wherein the above-mentioned target video library is related to the above-mentioned target object; and a third generation module for synthesizing the above-mentioned subtitle text, the above-mentioned audio information and the above-mentioned initial video to generate a target video.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described above.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 Schematically illustrates an exemplary system architecture to which the video generation method and apparatus according to an embodiment of the present disclosure may be applied;
[0012] Figure 2 The flowchart of the video generation method according to the embodiment of the present disclosure is schematically shown;
[0013] Figure 3 The flowchart of the method for generating an initial video according to an embodiment of the present disclosure is schematically shown;
[0014] Figure 4 The following schematically shows an implementation flow chart of the time clipping method according to an embodiment of the present disclosure;
[0015] Figure 5 The following schematically shows an implementation flow chart of a time clipping method according to another embodiment of the present disclosure;
[0016] Figure 6 The flowchart of the target video generation method according to an embodiment of the present disclosure is schematically shown;
[0017] Figure 7 A flowchart of a subtitle sentence segmentation method according to an embodiment of the present disclosure is schematically shown;
[0018] Figure 8 A schematic diagram schematically illustrates a video generation method according to an embodiment of the present disclosure;
[0019] Figure 9 A block diagram schematically shows a video generating apparatus according to an embodiment of the present disclosure; and
[0020] Figure 10A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] Figure 1 An exemplary system architecture to which the video generation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0023] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the video generation method and apparatus may be applied may include a terminal device, but the terminal device may implement the video generation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0024] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0025] Users can use terminal devices 101 , 102 , 103 to interact with server 105 via network 104 to receive or send messages, etc.
[0026] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0027] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports video generation requests initiated by users using the terminal devices 101, 102, and 103. The background management server may generate a video based on the object entity words included in the received video generation request and feed the generated target video back to the terminal device.
[0028] It should be noted that the video generation method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the video generation device provided in the embodiment of the present disclosure can also be set in the terminal device 101, 102, or 103. Alternatively, the video generation method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the video generation device provided in the embodiment of the present disclosure can generally be set in the server 105. The video generation method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the video generation device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0029] For example, a user may initiate a video generation request through an input operation on the terminal device 101. After receiving the video generation request, the terminal device 101 may obtain a target parameter library and a target video library from the server 105 via the network 104, and then use the target parameter library to generate speech text and subtitle text respectively, convert the speech text to obtain audio information, use the target video library to generate an initial video, and then synthesize the subtitle text, audio information and initial video into a target video, and display it on the display interface of the terminal device 101. Alternatively, the terminal device 101 may forward the video generation request to the server 105 via the network 104, and the server 105 may execute the video generation method to generate the target video. Afterwards, the server 105 may return the target video to the terminal device 101 via the network 104, and the terminal device 101 may display the target video on the display interface.
[0030] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0031] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0032] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0033] Figure 2 The flowchart of the video generation method according to the embodiment of the present disclosure is schematically shown.
[0034] like Figure 2 As shown, the video generating method 200 may include operations S210 to S240.
[0035] In operation S210, in response to receiving a video generation request, a speech text and a subtitle text are respectively generated using a target parameter library.
[0036] In operation S220, the speech text is converted to obtain audio information.
[0037] In operation S230, an initial video is generated using a target video library according to the duration of the audio information.
[0038] In operation S240, the subtitle text, the audio information, and the original video are synthesized to generate a target video.
[0039] According to an embodiment of the present disclosure, a video generation request may be a request for generating a video related to a target object. The target object may be various commodities, such as mobile phones, steel gratings, etc., or may be a person, a place name, a natural landscape, etc., which are not limited here. The video generation request may include a target object entity word of the target object. The target object entity word may refer to a well-known name or abbreviation of the target object. For example, the target object may be a "household remote-controlled air circulation fan", then the target object entity word of the target object may be "electric fan".
[0040] According to an embodiment of the present disclosure, the target parameter library may be related to the target object. For example, the target parameter library may be a database constructed based on the dimension of the target object entity word of the target object. The database type of the target parameter library is not limited, for example, it may be mySQL, Redis, etc. The target parameter library may include multiple object parameters, and the object parameters may include parameter items and parameter values. The storage method of the parameter items and parameter values in the target parameter library is related to the database type of the target parameter library. For example, the database type of the target parameter library may be a KV database, then the parameter items in the object parameters may be used as primary keys, and the parameter values may be used as attribute values, and stored in the form of key-value pairs.
[0041] According to embodiments of the present disclosure, the speech text and subtitle text may include a string of characters consisting of natural language and punctuation marks. The language of the natural language contained in the string is not limited herein and may be, for example, Chinese or English. As an optional embodiment, the string may be composed of natural language in multiple languages.
[0042] According to the embodiments of the present disclosure, the conversion of speech text into audio information may be achieved by utilizing various Text to Speech (TTS) algorithms or tools, which are not limited herein.
[0043] According to an embodiment of the present disclosure, a target video library may be associated with a target object. Specifically, the target object may have subordinate objects, each of which may be an object that has ownership rights or other rights over the target object. The videos included in the target video library may all be filmed or produced by the subordinate objects. As an optional embodiment, the videos included in the target video library may be filmed or produced by another object, and the subordinate object has obtained permission to use the videos.
[0044] According to an embodiment of the present disclosure, there may be differences in data such as duration, frame rate, and size of each video frame between the videos included in the target video library. The process of using the target video library to generate the initial video may include operations such as video cropping, splicing, frame rate adjustment, and video frame filling and scaling.
[0045] According to an embodiment of the present disclosure, when synthesizing a target video, the subtitle text and audio information can be separately embedded into the original video. Specifically, since the duration of the original video is consistent with that of the audio information, the audio information can be directly imported into the audio track of the original video. For the subtitle text, the video frames corresponding to the individual text segments segmented within the subtitle text can be determined based on the time points of the audio information, and then the individual text segments can be rendered in their corresponding video frames based on the corresponding relationships.
[0046] According to an embodiment of the present disclosure, when generating a video, a priori knowledge base related to the target object, i.e., a target parameter library, can be used to generate speech text and subtitle text respectively. The speech text can be converted into audio information using technologies such as TTS. Then, based on the duration of the audio information, an initial video can be generated using a target video library related to the target object, thereby achieving audio and video matching. Afterwards, the audio information and subtitle text can be embedded in the initial video respectively to ultimately obtain a synthesized target video, which can effectively ensure the consistency and integrity of the video semantics and reduce the cost of video generation.
[0047] It can be understood that the method of the present disclosure is described in detail above, and the method provided by the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0048] According to an embodiment of the present disclosure, operation S210 may specifically include the following operations:
[0049] In response to receiving a video generation request, a text template related to the target object is obtained based on the target object entity word, wherein the text template includes multiple information items to be replaced; based on the target object entity word, a target parameter library is generated from the initial parameter library; multiple target object parameters corresponding to each of the multiple information items are determined from the target parameter library; the multiple information items are replaced respectively with the multiple target object parameters to obtain subtitle text; and according to the feature word detection results of each of the multiple target object parameters, the subtitle text is used to generate speech text.
[0050] According to an embodiment of the present disclosure, a text template can be associated with an object entity word. That is, different objects with the same object entity word can have the same corresponding text template. For example, if the object entity word for "household remote-controlled air circulation fan" and "tower fan" is "electric fan," then "household remote-controlled air circulation fan" and "tower fan" have the same corresponding text template.
[0051] According to an embodiment of the present disclosure, a text template can be pre-compiled by a developer. The text template can include multiple information items to be replaced. For example, the text template can be: ${object name} is a ${object category} with a ${parameter item} of ${parameter value}, and ${object name}, ${parameter item}, ${parameter value} and ${object category} can be the information items to be replaced in the text template. Each information item can be composed of an information item indicator and an information item content identifier. The information item indicator can be used to indicate the position of the information item. Taking the above example as an example, the information item indicator in the information item can be ${}. The information item content identifier can represent the category of the parameter used to replace the information item. Taking the above example as an example, the information item content identifier in the first information item can be an object name, that is, indicating that the category of the parameter used to replace the first information item is the object name. The multiple information items included in the text template can be determined by text recognition, regular expression matching, etc., which are not limited here.
[0052] According to an embodiment of the present disclosure, the initial parameter library may include multiple object parameter groups aggregated by object entity words as dimensions, and each object parameter group may include multiple object parameters. Generating the target parameter library based on the target object entity words may include extracting the target object parameter groups associated with the target object entity words from the initial parameter library, and then generating the target parameter library based on the target object parameter groups.
[0053] As an optional implementation, when generating a target parameter library based on a target object parameter group, the object parameters included in the target object parameter group can be filtered and classified. Object parameters can include parameter items and parameter values. Specifically, filtering can include filtering based on frequency thresholds for parameter values within the same dimension, filtering based on a parameter item blacklist, filtering based on regular expressions, filtering based on inclusion relationships, or filtering based on similarity. For example, object parameters with a parameter value frequency of less than 5 can be directly filtered. For another example, object parameters with a parameter item such as "company name" that hits a blacklist can be directly filtered. For another example, for two object parameters with parameter values of "easy to maintain" and "easy to maintain," a semantic similarity model can be used to determine that the two object parameters have similar meanings, and therefore one of the two object parameters can be filtered. Classification can combine regular expressions and text classification models to classify the object parameters included in the target object parameter group into different categories for storage and use. Object parameter categories can include, for example, functional parameters, usage parameters, advantage parameters, professional parameters, brand parameters, etc. For example, if the parameter item of an object parameter is "main function", the object parameter can be classified as a functional parameter. For another example, if the parameter value of an object parameter is "high temperature resistance", the object parameter can be classified as an advantage parameter.
[0054] According to an embodiment of the present disclosure, determining multiple target object parameters corresponding to each of the multiple information items from the target parameter library can be performed based on the information item content identifiers of the multiple information items, that is, for the corresponding information items and target object parameters, the information item content identifier of the information item is consistent with the category to which the target object parameter belongs.
[0055] According to an embodiment of the present disclosure, replacing multiple information items with multiple target object parameters respectively can be performed by determining the position information of each information item based on the information item indicator of each information item, and then replacing the multiple information items with multiple target object parameters corresponding to each of the multiple information items based on the position information. For example, for the target object "hot-dip galvanized steel grating", the text template can be: ${object name} is a ${category} with a ${professional parameter value} of ${professional parameter item}, which can be used in scenarios such as ${purpose parameter value}. After determining the target object parameters and replacing them with the target object parameters, the subtitle text obtained can be: Hot-dip galvanized steel grating is a building material with a service life of 10 years, which can be used in scenarios such as platforms in fields such as municipal engineering.
[0056] As an optional implementation, the target object parameters required by some information items in the text template may not be directly stored in the target parameter library. For example, if the target object parameter category required by the information item is industrial belt data, this industrial belt data can be obtained by statistically analyzing the object parameters with the parameter category of production in the target parameter library.
[0057] According to an embodiment of the present disclosure, the feature word detection of the target object parameter may be performed by text recognition, regular expression matching, etc. to determine whether the target object parameter includes the feature word.
[0058] According to the embodiments of the present disclosure, by utilizing text templates and target parameter libraries to generate language text and subtitle text, the knowledge richness and personalization of the text can be effectively improved and the text generation cost can be reduced.
[0059] According to an embodiment of the present disclosure, the initial parameter library used in operation S210 may be generated by the following method:
[0060] Entity recognition is performed on the object title included in the object information to obtain object entity words related to the object information; and based on the object entity words, object parameters included in the object information are written into an initial parameter library.
[0061] According to embodiments of the present disclosure, object information can come from a second object. The second object can provide the object information by uploading information, or the object information can be obtained from the internet after obtaining the second object's permission. For example, if the object information is a product and the second object is a merchant, the object information can be derived from information such as the title, category, and parameters entered by the merchant when listing the product on the platform.
[0062] According to an embodiment of the present disclosure, when performing entity recognition on an object title, functional terms included in the object title can be deleted to obtain the object entity word. For example, if the object title can be: Platform hot-dip galvanized steel grating, strong bearing capacity, beautiful appearance, durable and easy to clean, then the object entity word obtained by entity recognition can be hot-dip galvanized steel grating.
[0063] According to an embodiment of the present disclosure, object parameters included in object information can be aggregated into object parameter groups using object entity words as a dimension. As an optional implementation, object categories can also be determined based on object entity words, and object parameter groups can be aggregated using the object entity words and object categories as aggregation dimensions.
[0064] According to an embodiment of the present disclosure, by enriching the initial parameter library based on object information from a second object, the accuracy of the object parameters in the initial parameter library can be improved, thereby enhancing the knowledge richness and personalization of the generated text.
[0065] According to an embodiment of the present disclosure, generating speech text using subtitle text according to the feature word detection results of each of the multiple target object parameters may specifically include the following operations:
[0066] For each target object parameter, feature word detection is performed on the target object parameter to obtain a second detection result; when the second detection result indicates that a first feature word exists in the target object parameter, the first feature word included in the target object parameter is converted using a mapping table to obtain a second feature word; and the first feature word in the subtitle text is replaced with the second feature word to generate a speech text.
[0067] According to an embodiment of the present disclosure, the first feature word may include various characters or words that do not conform to language expression habits, and the language expression habits may be related to the main language in the subtitle text. For example, if the main language in the subtitle text is Chinese, then the language expression habits should conform to the language expression habits of Chinese. Specifically, the first feature word may include various English abbreviations, Japanese characters, etc. with Chinese meanings. For example, in the case where the object is a commodity, many parameters of the commodity will have English units, such as mm (millimeter), db (decibel), MHz (megahertz), etc. Correspondingly, the second feature word may be a character or word that has the same meaning as the first feature word and conforms to the language expression habits. For example, the target object parameter may be 20db. When performing feature word detection, it can be detected that the target object parameter includes the first feature word db. If the target object parameter is directly converted into audio information before conversion, the first feature word will be converted into the pronunciation of English letters, thereby affecting the professionalism and comprehensibility of the audio. Using the mapping table, it can be determined that the second feature word corresponding to the first feature word db can be decibel, that is, the target object parameter can be converted from 20db to 20 decibels. The audio converted by the converted target object parameter can be easier to understand, thereby improving the professionalism and accuracy of the audio information obtained by speech-to-text conversion.
[0068] According to the embodiments of the present disclosure, the mapping table may be completed by manual annotation or obtained by annotation using various artificial intelligence algorithms, which is not limited here.
[0069] According to an embodiment of the present disclosure, multiple target object parameters may not contain the first feature word, that is, when the second detection results of the multiple target object parameters all indicate that the first feature word does not exist in the target object parameters, the subtitle text can be determined to be speech text.
[0070] According to an embodiment of the present disclosure, an initial video can be generated using a target video library according to the duration of the audio information. The generation of the initial video can be achieved, for example, by an initial video generation method.
[0071] Figure 3The flowchart of the initial video generation method according to the embodiment of the present disclosure is schematically shown.
[0072] like Figure 3 As shown, the initial video generation method may include operations S231 to S233.
[0073] In operation S231 , an expected size of a video frame included in a target video is determined based on size information of each video frame of each of the plurality of first videos.
[0074] In operation S232 , the plurality of first videos are pre-processed based on an expected size of a video frame included in the target video to obtain a plurality of second videos.
[0075] In operation S233 , the plurality of second videos are cropped and spliced according to the duration of the audio information to obtain an initial video.
[0076] According to an embodiment of the present disclosure, multiple first videos may be included in a target video library. The first video may be composed of multiple video frames, and the frame rates of the multiple first videos are not limited herein. The size information of the video frame of the first video may represent the size of the video frame. For example, if the size of the video frame is 3.5 cm × 2 cm, the size information may be represented as 3.5 cm × 2 cm. The multiple video frames included in the first video may have the same size.
[0077] According to an embodiment of the present disclosure, based on the size information of each video frame of each of the multiple first videos, determining the expected size of the video frames included in the target video can be done by dividing the first videos with the same size information into the same group based on the size information, and counting the number of first videos included in each group. The size information corresponding to the group with the highest number of first videos can be determined as the expected size of the video frames included in the target video. Alternatively, a target first video can be randomly selected from the multiple first videos, and the size information of the video frames in the target first video can be determined as the expected size of the video frames included in the target video.
[0078] According to an embodiment of the present disclosure, preprocessing the first video may include padding, scaling, and other processing of each video frame of the first video. For example, when the video frame size indicated by the size information of the video frame in the first video is smaller than the expected size, the video frame of the first video may be padded with black edges. For another example, when the video frame size indicated by the size information of the video frame in the first video is larger than the expected size, the video frame of the first video may be scaled to the length of the video frame of the target video based on its long side, and then the black edge padding process may be performed.
[0079] According to embodiments of the present disclosure, the total duration of the multiple second videos can be greater than the duration of the audio information. When generating an initial video using the multiple second videos, the multiple second videos can be first spliced into a single video, and then the initial video can be obtained by cropping the single video based on the duration of the audio information. Alternatively, partial video clips can be captured from each of the multiple second videos, and then the multiple video clips can be spliced into the initial video.
[0080] According to the embodiments of the present disclosure, by using multiple first videos included in the target video library to generate the initial video, the content consistency among the multiple first videos can be maintained because they are all related to the target object entity word. Furthermore, by generating the initial video based on the duration of the audio information, the semantic consistency and integrity of the video can be effectively guaranteed.
[0081] According to an embodiment of the present disclosure, a target video library may be generated based on an initial video library after receiving a video generation request. Specifically, the generation of the target video library may include the following operations:
[0082] In response to receiving a video generation request, multiple third videos are extracted from an initial video library based on the target object entity word; video quality filtering is performed on the multiple third videos to obtain multiple fourth videos; and based on object attributes associated with the target object, multiple first videos are determined from the multiple fourth videos.
[0083] According to an embodiment of the present disclosure, multiple third videos may all be related to the target object entity word. Specifically, for each video in the initial video library, the target object entity word and the title of the video may be subjected to a semantic similarity calculation to determine whether the video is a related video, i.e., a third video. Alternatively, the titles of the videos in the initial video library may be sequence-labeled, and keywords therein may be identified through weight calculation, and then whether the video is a third video may be determined by determining whether the keywords include the target object entity word.
[0084] According to embodiments of the present disclosure, video quality filtering may include, but is not limited to, risk control filtering and image quality filtering. For example, risk control filtering may filter videos containing pornography, gambling, drugs, or other illegal content. Image quality filtering may filter videos with low resolution, watermarks, subtitles, or flash images.
[0085] According to an embodiment of the present disclosure, the object attributes associated with the target object may be represented as information such as copyright and ownership. Taking product A as an example, the target object may include videos of product A and similar products from different manufacturers. The first video obtained by screening may be a video of product A and similar products from a single manufacturer.
[0086] According to the embodiments of the present disclosure, by using the above video screening method, the use risk of the synthesized initial video can be reduced and the rights and interests of the video demander can be protected.
[0087] According to an embodiment of the present disclosure, the initial video library can be generated by the following method:
[0088] Performing information detection on the description information related to the fifth video to obtain a first detection result; and adding the fifth video to the initial video library if the first detection result indicates that the description information meets accuracy and richness conditions.
[0089] According to an embodiment of the present disclosure, the fifth video may come from the first object. Specifically, the fifth video may be uploaded to the platform by the first object, or the fifth video may be obtained from the Internet after obtaining permission from the first object.
[0090] According to an embodiment of the present disclosure, the description information related to the fifth video may be information provided by the first object when the fifth video is made public on the Internet. The description information may include an object title, object parameters, and the like.
[0091] According to an embodiment of the present disclosure, the description information satisfying the accuracy condition may mean that the accuracy of the description information determined by the machine learning model using accuracy recognition is greater than a preset threshold. Correspondingly, the description information satisfying the richness condition may mean that the accuracy of the description information determined by the machine learning model using richness recognition is greater than a preset threshold.
[0092] As an optional implementation, when the fifth video is actively uploaded by the first object, if the first detection result indicates that the description information does not meet the accuracy and richness conditions, feedback information can also be sent to the first object to prompt the first object to re-fill in the description information of the fifth video.
[0093] According to an embodiment of the present disclosure, multiple second videos are cropped and spliced according to the duration of the audio information to obtain an initial video, which can be achieved by, for example, a time cropping method.
[0094] Figure 4 The following schematically shows an implementation flow chart of the time clipping method according to an embodiment of the present disclosure.
[0095] like Figure 4 As shown, the implementation process of the time clipping method may include operations S23301 to S23309.
[0096] In operation S23301, a target duration is determined based on the duration of the audio information and the number of second videos.
[0097] In operation S23302, the duration of each of the plurality of second videos and the sum of the target duration and the preset duration are compared to obtain a plurality of comparison results.
[0098] In operation S23303, it is determined whether a comparison result indicates that the duration of the second video is less than the sum of the target duration and the preset duration. If it is determined that at least one target comparison result indicates that the duration of the second video is less than the sum of the target duration and the preset duration, operation S23304 is performed. If it is determined that multiple comparison results all indicate that the duration of the second video is greater than or equal to the sum of the target duration and the preset duration, operation S23308 is performed.
[0099] In operation S23304, the plurality of second videos are divided into a first video set and a second video set based on the plurality of comparison results.
[0100] In operation S23305, a remaining duration is determined based on the respective durations of the at least one target second video and the duration of the audio information.
[0101] In operation S23306, a new target duration is determined based on the remaining duration and the number of second videos included in the second video set.
[0102] In operation S23307, the duration of the second video included in the second video set is compared with the target duration and the preset duration to obtain a new comparison result. After completing operation S23307, the process returns to operation S23303.
[0103] In operation S23308, the plurality of second videos are divided into a plurality of video segments using a sliding window with a window size of a target duration and a preset duration as a step length.
[0104] In operation S23309, an initial video is obtained by splicing based on at least one target second video and / or multiple video segments.
[0105] According to an embodiment of the present disclosure, based on the duration of the audio information and the number of the second video, the target duration may be determined as shown in formula (1):
[0106]
[0107] In formula (1), T may represent the duration of the audio information, n may represent the number of second videos, and t may represent the target duration.
[0108] According to an embodiment of the present disclosure, the preset duration can be set according to a specific application scenario. For example, the preset duration can be set to 1s, 0.5s, etc., which is not limited here.
[0109] According to an embodiment of the present disclosure, at least one target comparison result may belong to multiple comparison results. Accordingly, at least one target comparison result may correspond to at least one target second video in the multiple second videos. Dividing the multiple second videos into the first video set and the second video set based on the multiple comparison results may be dividing the at least one target second video into the first video set, and dividing the second videos in the multiple second videos other than the at least one target second video into the second video set.
[0110] According to an embodiment of the present disclosure, based on the duration of each of the at least one target second video and the duration of the audio information, the remaining duration may be determined as shown in formula (2):
[0111]
[0112] In formula (2), T′ may represent the remaining duration, m may represent the number of target second videos, and t2(i) may represent the duration of the i-th target second video.
[0113] According to an embodiment of the present disclosure, determining a new target duration can be achieved using formula (1) as described above, by replacing the duration of the audio information in formula (1) with the remaining duration, and replacing the number of second videos with the number of second videos included in the second video set.
[0114] According to an embodiment of the present disclosure, the duration of the video segments obtained by segmentation may be the target duration.
[0115] According to the embodiments of the present disclosure, when performing video splicing, the target second videos and the video segments can be spliced in any order, and the number of initial videos that can be spliced can be the same as the number of permutations and combinations of the target second videos and the video segments. Therefore, through the temporal cropping method, more initial videos can be spliced based on the same number of first videos, thereby fully utilizing existing videos to create new videos, effectively reducing the cost of video generation.
[0116] Figure 5 The following schematically shows a flow chart of an implementation of a time clipping method according to another embodiment of the present disclosure.
[0117] like Figure 5 As shown, the implementation process of the time clipping method includes operations S23311 to S23319.
[0118] In operation S23311, a target duration is determined based on the duration of the audio information and the number of second videos.
[0119] In operation S23312, it is determined whether there is a second video with a duration less than the target duration. If it is determined that there is a second video with a duration less than the target duration, operation S23313 is performed. If it is determined that the duration of each of the plurality of second videos is greater than or equal to the target duration, operation S23316 is performed.
[0120] In operation S23313, at least one target second video having a duration shorter than a target duration is determined.
[0121] In operation S23314, a remaining duration is determined based on the respective durations of at least one target second video and the duration of the audio information.
[0122] In operation S23315, based on the remaining duration and the number of second videos included in the second video set, a new target duration is determined, and at least one target second video is removed from the plurality of second videos. After completing operation S23315, the process may return to operation S23312.
[0123] In operation S23316, it is determined whether there is a second video whose duration is less than the sum of the target duration and the preset duration. If it is determined that there is a second video whose duration is less than the sum of the target duration and the preset duration, operation S23317 is performed. If it is determined that the duration of each of the plurality of second videos is greater than or equal to the sum of the target duration and the preset duration, operation S23318 is performed.
[0124] In operation S23317, the target second video is obtained by cropping the second video whose duration is less than the sum of the target duration and the preset duration. After operation S23317 is completed, operation S23314 is performed.
[0125] In operation S23318, the plurality of second videos are divided into a plurality of video segments using a sliding window with a window size of a target duration and a preset duration as a step length.
[0126] In operation S23319, an initial video is obtained by splicing based on at least one target second video and / or multiple video clips.
[0127] According to an embodiment of the present disclosure, the duration of the second video selected in operation S23317 can be between [t, t+step), where t can represent a target duration and step can represent a preset duration. The target second video captured from the selected second video can be any video clip of duration t captured from the selected second video.
[0128] According to an embodiment of the present disclosure, by comparing the duration of the second video with the target duration and the sum of the target duration and the preset duration in sequence, existing video resources can be more fully utilized to further reduce video generation costs.
[0129] According to an embodiment of the present disclosure, subtitle text, audio information and an initial video may be synthesized to generate a target video. The generation of the target video may be achieved, for example, by a target video generation method.
[0130] Figure 6 The flowchart of the target video generation method according to the embodiment of the present disclosure is schematically shown.
[0131] like Figure 6 As shown, the target video generating method may include operations S241 to S242.
[0132] In operation S241, the subtitle text is segmented into sentences to obtain a subtitle file.
[0133] In operation S242, the subtitle file and the audio information are respectively embedded into the original video to generate a target video.
[0134] According to an embodiment of the present disclosure, sentence segmentation of subtitle text can be performed by setting multiple segmentation points based on constraints, thereby segmenting the subtitle text into multiple text paragraphs, and the multiple text paragraphs constitute a subtitle file. Constraints may include font size constraints, line break constraints, single line control constraints, etc., which are not limited here.
[0135] According to the embodiments of the present disclosure, by first segmenting the subtitle text and then generating the target video, the subtitles on the target video will not block the video content and the semantic integrity of the subtitles can be guaranteed, thereby effectively improving the viewing experience and marketing effect of the target video.
[0136] According to an embodiment of the present disclosure, sentence segmentation of a subtitle text to obtain a subtitle file can be achieved, for example, by a subtitle sentence segmentation method.
[0137] Figure 7 The flowchart of the subtitle sentence segmentation method according to an embodiment of the present disclosure is schematically shown.
[0138] like Figure 7 As shown, the subtitle sentence segmentation method may include operations S2411 to S2416.
[0139] In operation S2411, a first segmentation point is determined based on positions of preset symbols included in the subtitle text.
[0140] In operation S2412, a plurality of first text paragraphs included in the subtitle text are determined based on the first segmentation point.
[0141] In operation S2413, it is determined whether the first text paragraph satisfies the display constraint condition. If it is determined that the first text paragraph does not satisfy the display constraint condition, operation S2414 is performed. If it is determined that the first text paragraph satisfies the display constraint condition, operation S2416 is performed.
[0142] In operation S2414 , sequence tagging is performed on the first text paragraph to obtain second segmentation points.
[0143] In operation S2415 , a third dividing point is determined from the second dividing points based on a preset strategy.
[0144] In operation S2416 , sentence processing is performed on the subtitle text using the first segmentation point and third segmentation points corresponding to each of the plurality of first text paragraphs to obtain a subtitle file.
[0145] According to an embodiment of the present disclosure, the preset symbols may include punctuation marks such as a period, a question mark, and a semicolon.
[0146] According to an embodiment of the present disclosure, the first text paragraph meeting the display constraint condition may mean that all characters of the first text paragraph can be displayed on the same video frame.
[0147] According to an embodiment of the present disclosure, a display constraint may include multiple constraints. For example, if the pixel width of the target video and the pixel width of each character have been determined, the maximum number of characters that can be accommodated in each line of the target video can be determined, and then the maximum number of characters in the text paragraph can be determined based on the number of subtitle lines limited to each video frame. For another example, proper nouns in the subtitle text need to be controlled as isolated lines. Therefore, when a proper noun is included in the first text paragraph, it is necessary to evaluate whether the proper noun can be fully displayed in a single line to form a single-line control constraint.
[0148] According to an embodiment of the present disclosure, for the first text paragraph, there can be multiple second segmentation points determined after sequence annotation. When selecting the third segmentation point, the preset strategy can be expressed as first segmenting the first text paragraph at the second segmentation point closest to the center, and judging whether the sub-paragraphs obtained by segmentation all meet the display constraint conditions. If so, the second segmentation point closest to the center can be determined as the third segmentation point. If not, a greedy number system can be used to first determine the first third segmentation point from the first text paragraph. The first third segmentation point can be one of the sub-paragraphs obtained by segmentation that meets the display constraint conditions. For the other sub-paragraph, the preset strategy is continued to be used for subsequent segmentation actions.
[0149] According to an embodiment of the present disclosure, the start and end display times of each paragraph in a subtitle file can be controlled by a timestamp.
[0150] According to the embodiments of the present disclosure, through the above subtitle sentence segmentation solution, the subtitles can be adaptively adapted to the size of the target video, thereby improving the display effect of the subtitles and reducing the generation cost of the subtitle file.
[0151] Figure 8 The diagram schematically shows a video generation method according to an embodiment of the present disclosure.
[0152] like Figure 8 As shown, after receiving a video generation request initiated by user 801, data can be recalled from the initial video library 803 and initial parameter library 804 in parallel based on the target object entity word 802 included in the video generation request to obtain a target video library 805 and a target parameter library 806, respectively. A text template 807 can also be determined based on the target object entity word 802. Using the information items included in text template 807, corresponding target object parameters can be obtained from the target parameter library 806. These target object parameters are then used to replace the information items to obtain subtitle text 808. Feature words in subtitle text 808 are replaced to obtain audio text 809. Audio text 809 can be converted using text-to-speech (TTS) to obtain audio information 810. Based on the duration of audio information 810, a time cropping method can be used to crop and splice the multiple videos included in target video library 805 to obtain an initial video 811. Simultaneously, based on preset constraints, the text in subtitle text 808 can be segmented into multiple text paragraphs to obtain a subtitle file 812. Afterwards, the audio information 810 , the initial video 811 and the subtitle file 812 may be synthesized to generate a target video 813 .
[0153] According to an embodiment of the present disclosure, when generating a video, a priori knowledge base related to the target object, i.e., a target parameter library, can be used to generate speech text and subtitle text respectively. The speech text can be converted into audio information using technologies such as TTS. Then, based on the duration of the audio information, an initial video can be generated using a target video library related to the target object, thereby achieving audio and video matching. Afterwards, the audio information and subtitle text can be embedded in the initial video respectively to ultimately obtain a synthesized target video, which can effectively ensure the consistency and integrity of the video semantics and reduce the cost of video generation.
[0154] Figure 9 The block diagram schematically shows a video generating device according to an embodiment of the present disclosure.
[0155] like Figure 9 As shown, the video generating apparatus 900 may include a first generating module 910 , a conversion module 920 , a second generating module 930 and a third generating module 940 .
[0156] The first generation module 910 is configured to generate speech text and subtitle text respectively using a target parameter library in response to receiving a video generation request, wherein the video generation request includes a target object entity word of a target object, and the target parameter library is related to the target object.
[0157] The conversion module 920 is used to convert the speech text to obtain audio information.
[0158] The second generating module 930 is configured to generate an initial video using a target video library according to the duration of the audio information, wherein the target video library is related to the target object.
[0159] The third generation module 940 is used to synthesize the subtitle text, audio information and the initial video to generate a target video.
[0160] According to an embodiment of the present disclosure, when generating a video, a priori knowledge base related to the target object, i.e., a target parameter library, can be used to generate speech text and subtitle text respectively. The speech text can be converted into audio information using technologies such as TTS. Then, based on the duration of the audio information, an initial video can be generated using a target video library related to the target object, thereby achieving audio and video matching. Afterwards, the audio information and subtitle text can be embedded in the initial video respectively to ultimately obtain a synthesized target video, which can effectively ensure the consistency and integrity of the video semantics and reduce the cost of video generation.
[0161] According to an embodiment of the present disclosure, the target video library includes a plurality of first videos.
[0162] According to an embodiment of the present disclosure, the second generation module 910 includes:
[0163] The first generating submodule is configured to determine an expected size of a video frame included in a target video based on size information of each video frame of each of the plurality of first videos.
[0164] The second generating submodule is configured to pre-process the plurality of first videos based on an expected size of video frames included in the target video to obtain a plurality of second videos.
[0165] The third generation submodule is used to crop and splice the multiple second videos according to the duration of the audio information to obtain an initial video.
[0166] According to an embodiment of the present disclosure, the third generation submodule includes:
[0167] The first generating unit is configured to determine a target duration based on the duration of the audio information and the number of the second videos.
[0168] The second generating unit is used to compare the duration of each of the plurality of second videos and the sum of the target duration and the preset duration respectively to obtain a plurality of comparison results.
[0169] The third generation unit is used to divide the multiple second videos into multiple video segments using a sliding window with a window size of the target duration and a preset duration, with a preset duration as a step size. When multiple comparison results all indicate that the duration of the second video is greater than or equal to the sum of the target duration and the preset duration, the multiple second videos are divided into multiple video segments using a sliding window with a window size of the target duration.
[0170] The fourth generating unit is configured to splice the multiple video clips to obtain an initial video.
[0171] According to an embodiment of the present disclosure, the third generating unit further includes:
[0172] The fifth generation unit is used to divide multiple second videos into a first video set and a second video set based on multiple comparison results when at least one target comparison result indicates that the duration of the second video is less than the sum of the target duration and the preset duration, wherein at least one target comparison result belongs to multiple comparison results, the first video set includes at least one target second video, and the at least one target second video corresponds to each of the at least one target comparison result.
[0173] The sixth generating unit is configured to determine a remaining duration based on the duration of each of the at least one target second video and the duration of the audio information.
[0174] The seventh generating unit is configured to determine a new target duration based on the remaining duration and the number of second videos included in the second video set.
[0175] The eighth generating unit is configured to compare the duration of the second video included in the second video set with the target duration and the preset duration to obtain a new comparison result.
[0176] According to an embodiment of the present disclosure, the fourth generating unit includes:
[0177] The first generating subunit is configured to splice an initial video based on at least one target second video and a plurality of video segments.
[0178] According to an embodiment of the present disclosure, the video generating apparatus 900 may further include:
[0179] The extraction module is used to extract multiple third videos from the initial video library based on the target object entity words in response to receiving the video generation request.
[0180] The processing module is used to perform video quality filtering on the multiple third videos to obtain multiple fourth videos.
[0181] The determining module is configured to determine a plurality of first videos from a plurality of fourth videos based on an object attribute associated with the target object.
[0182] According to an embodiment of the present disclosure, the video generating apparatus 900 may further include:
[0183] The detection module is configured to perform information detection on the description information related to the fifth video to obtain a first detection result, wherein the fifth video comes from the first object.
[0184] The adding module is used to add the fifth video to the initial video library when the first detection result indicates that the description information meets the accuracy and richness conditions.
[0185] According to an embodiment of the present disclosure, the third generation module includes:
[0186] The fourth generation submodule is used to perform sentence processing on the subtitle text to obtain a subtitle file.
[0187] The fifth generation submodule is used to embed the subtitle file and audio information into the initial video respectively to generate the target video.
[0188] According to an embodiment of the present disclosure, the fourth generation submodule includes:
[0189] The ninth generating unit is configured to determine a first segmentation point based on positions of preset symbols included in the subtitle text.
[0190] The tenth generating unit is configured to determine a plurality of first text paragraphs included in the subtitle text based on the first segmentation point.
[0191] The eleventh generating unit is configured to perform sequence annotation on each first text paragraph to obtain a second segmentation point when the first text paragraph does not satisfy the display constraint condition.
[0192] The twelfth generating unit is configured to determine a third dividing point from the second dividing points based on a preset strategy.
[0193] The thirteenth generating unit is configured to perform sentence segmentation processing on the subtitle text by using the first segmentation point and third segmentation points corresponding to each of the plurality of first text paragraphs to obtain a subtitle file.
[0194] According to an embodiment of the present disclosure, the first generation module includes:
[0195] The sixth generation submodule is configured to, in response to receiving a video generation request, obtain a text template related to the target object based on the target object entity word, wherein the text template includes a plurality of information items to be replaced.
[0196] The seventh generation submodule is used to generate a target parameter library from the initial parameter library based on the target object entity word.
[0197] The eighth generating submodule is configured to determine, from the target parameter library, a plurality of target object parameters corresponding to the plurality of information items.
[0198] The ninth generation submodule is used to replace the multiple information items with the multiple target object parameters respectively to obtain the subtitle text.
[0199] The tenth generation submodule is used to generate speech text using the subtitle text according to the feature word detection results of each of the multiple target object parameters.
[0200] According to an embodiment of the present disclosure, the tenth generation submodule includes:
[0201] The fourteenth generating unit is configured to perform feature word detection on each target object parameter to obtain a second detection result.
[0202] The fifteenth generating unit is configured to, when the second detection result indicates that the first feature word exists in the target object parameter, convert the first feature word included in the target object parameter by using the mapping table to obtain the second feature word.
[0203] The sixteenth generating unit is configured to replace the first feature word in the subtitle text with the second feature word to generate a speech text.
[0204] According to an embodiment of the present disclosure, the tenth generation submodule further includes:
[0205] The seventeenth generating unit is configured to determine that the subtitle text is a speech text when the second detection results of the plurality of target object parameters all indicate that the first feature word does not exist in the target object parameters.
[0206] According to an embodiment of the present disclosure, the video generating apparatus 900 may further include:
[0207] The recognition module is used to perform entity recognition on the object title included in the object information to obtain object entity words related to the object information, wherein the object information comes from the second object.
[0208] The writing module is used to write the object parameters included in the object information into the initial parameter library based on the object entity words.
[0209] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0210] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0211] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0212] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0213] Figure 10 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. Electronic device 1000 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0214] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0215] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0216] The computing unit 1001 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the video generation method by any other appropriate means (e.g., by means of firmware).
[0217] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0218] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0219] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0220] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0221] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0222] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0223] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0224] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A video generation method, comprising: In response to receiving a video generation request, generating speech text and subtitle text respectively using a target parameter library, wherein the video generation request includes a target object entity word of a target object, and the target parameter library is related to the target object; Converting the speech text to obtain audio information; generating an initial video using a target video library according to a duration of the audio information, wherein the target video library is related to the target object; and Performing synthesis processing on the subtitle text, the audio information and the initial video to generate a target video; The step of generating an initial video using a target video library according to the duration of the audio information includes: The plurality of second videos are cropped and spliced according to the duration of the audio information to obtain the initial video, including: Determining a target duration based on the duration of the audio information and the number of the second videos; Comparing the duration of each of the plurality of second videos with the sum of the target duration and the preset duration, respectively, to obtain a plurality of comparison results; If at least one target comparison result indicates that the duration of the second video is less than the sum of the target duration and the preset duration, dividing the plurality of second videos into a first video set and a second video set based on the plurality of comparison results, wherein the at least one target comparison result belongs to the plurality of comparison results, the first video set includes at least one target second video, and the at least one target second video corresponds to each of the at least one target comparison result; determining a remaining duration based on a duration of each of the at least one target second video and a duration of the audio information; Determining a new target duration based on the remaining duration and the number of the second videos included in the second video set; Comparing the duration of the second video included in the second video set with the target duration and the preset duration to obtain a new comparison result; and The initial video is obtained by splicing based on the at least one target second video.
2. The method according to claim 1, wherein The target video library includes a plurality of first videos; The method further comprises: Determining an expected size of a video frame included in the target video based on size information of each video frame of each of the plurality of first videos; Based on the expected size of the video frames included in the target video, the multiple first videos are pre-processed to obtain the multiple second videos.
3. The method according to claim 2, further comprising: If the comparison results all indicate that the duration of the second video is greater than or equal to the sum of the target duration and the preset duration, dividing the plurality of second videos into a plurality of video segments using a sliding window with a window size of the target duration and a step size of the preset duration; as well as The initial video is obtained by splicing the multiple video clips and the at least one target second video.
4. The method according to claim 2, further comprising: In response to receiving the video generation request, extracting a plurality of third videos from an initial video library based on the target object entity word; Performing video quality filtering on the multiple third videos respectively to obtain multiple fourth videos; as well as The plurality of first videos are determined from the plurality of fourth videos based on an object property associated with the target object.
5. The method according to claim 4, further comprising: Performing information detection on the description information related to the fifth video to obtain a first detection result, wherein the fifth video is from the first object; and When the first detection result indicates that the description information meets the accuracy and richness conditions, the fifth video is added to the initial video library.
6. The method according to claim 1, wherein The synthesizing process of the subtitle text, the audio information and the initial video to generate a target video includes: Sentence processing is performed on the subtitle text to obtain a subtitle file; and The subtitle file and the audio information are respectively embedded into the original video to generate the target video.
7. The method according to claim 6, wherein: The sentence-segmenting process of the subtitle text to obtain a subtitle file includes: Determining a first segmentation point based on positions of preset symbols included in the subtitle text; Based on the first segmentation point, determining a plurality of first text paragraphs included in the subtitle text; For each of the first text paragraphs, if the first text paragraph does not meet the display constraint condition, performing sequence tagging on the first text paragraph to obtain a second segmentation point; Based on a preset strategy, determining a third split point from the second split points; and The subtitle text is segmented using the first segmentation point and third segmentation points corresponding to each of the plurality of first text paragraphs to obtain the subtitle file.
8. The method according to claim 1, wherein The method of generating a speech text and a subtitle text respectively by using a target parameter library in response to receiving a video generation request includes: In response to receiving the video generation request, obtaining a text template related to the target object based on the target object entity word, wherein the text template includes a plurality of information items to be replaced; Based on the target object entity word, generating the target parameter library from the initial parameter library; determining, from the target parameter library, a plurality of target object parameters corresponding to each of the plurality of information items; Replacing the multiple information items with the multiple target object parameters respectively to obtain the subtitle text; and The speech text is generated using the subtitle text according to the feature word detection results of each of the multiple target object parameters.
9. The method according to claim 8, wherein The step of generating the speech text by using the subtitle text according to the feature word detection results of each of the plurality of target object parameters includes: For each target object parameter, perform feature word detection on the target object parameter to obtain a second detection result; In a case where the second detection result indicates that a first feature word exists in the target object parameter, converting the first feature word included in the target object parameter using a mapping table to obtain a second feature word; and The first feature word in the subtitle text is replaced with the second feature word to generate the speech text.
10. The method according to claim 9, further comprising: When the second detection results of each of the plurality of target object parameters indicate that the first feature word does not exist in the target object parameter, the subtitle text is determined to be the speech text.
11. The method according to claim 8, further comprising: Performing entity recognition on the object title included in the object information to obtain object entity words related to the object information, wherein the object information comes from the second object; and Based on the object entity word, the object parameters included in the object information are written into the initial parameter library.
12. A video generating device, comprising: A first generating module is configured to, in response to receiving a video generation request, generate a speech text and a subtitle text using a target parameter library, wherein the video generation request includes a target object entity word of a target object, and the target parameter library is related to the target object, and includes: after generating the subtitle text, replacing characters or words in the subtitle text that do not conform to language expression habits with characters or words that conform to language expression habits to obtain the speech text; A conversion module, used to convert the speech text to obtain audio information; a second generating module, configured to generate an initial video using a target video library according to a duration of the audio information, wherein the target video library is related to the target object; and A third generating module is used to synthesize the subtitle text, the audio information and the initial video to generate a target video; The second generation module includes: The third generation submodule is configured to crop and splice the plurality of second videos according to the duration of the audio information to obtain the initial video, including: a first generating unit, configured to determine a target duration based on the duration of the audio information and the number of the second videos; A second generating unit is configured to compare the duration of each of the plurality of second videos with the sum of the target duration and the preset duration, respectively, to obtain a plurality of comparison results; a fifth generating unit, configured to, if at least one target comparison result indicates that a duration of the second video is less than a sum of the target duration and the preset duration, divide the plurality of second videos into a first video set and a second video set based on the plurality of comparison results, wherein the at least one target comparison result belongs to the plurality of comparison results, the first video set includes at least one target second video, and the at least one target second video corresponds to each of the at least one target comparison result; a sixth generating unit, configured to determine a remaining duration based on a duration of each of the at least one target second video and a duration of the audio information; a seventh generating unit, configured to determine a new target duration based on the remaining duration and the number of the second videos included in the second video set; an eighth generating unit, configured to compare the duration of the second video included in the second video set with the target duration and the preset duration to obtain a new comparison result; and The first generating subunit is configured to splice the at least one target second video to obtain the initial video.
13. The device according to claim 12, wherein The target video library includes a plurality of first videos; Wherein, the second generation module further includes: A first generating submodule, configured to determine an expected size of a video frame included in the target video based on size information of each video frame of each of the plurality of first videos; The second generating submodule is configured to pre-process the plurality of first videos based on an expected size of video frames included in the target video to obtain the plurality of second videos.
14. The device according to claim 13, wherein The third generation submodule further includes: a third generating unit, configured to, if the plurality of comparison results all indicate that the duration of the second video is greater than or equal to the sum of the target duration and the preset duration, segment the plurality of second videos into a plurality of video segments using a sliding window with a window size of the target duration and a step size of the preset duration; and The fourth generating unit is configured to splice the initial video based on the multiple video clips and the at least one target second video.
15. The apparatus according to claim 13, further comprising: an extraction module, configured to extract a plurality of third videos from an initial video library based on the target object entity word in response to receiving the video generation request; a processing module, configured to perform video quality filtering on the plurality of third videos respectively to obtain a plurality of fourth videos; as well as A determination module is configured to determine the plurality of first videos from the plurality of fourth videos based on an object attribute associated with the target object.
16. The apparatus according to claim 15, further comprising: a detection module, configured to perform information detection on description information related to a fifth video to obtain a first detection result, wherein the fifth video is from the first object; and An adding module is used to add the fifth video to the initial video library when the first detection result indicates that the description information meets the accuracy and richness conditions.
17. The device according to claim 12, wherein The third generation module includes: A fourth generating submodule is used to perform sentence processing on the subtitle text to obtain a subtitle file; and The fifth generating submodule is used to embed the subtitle file and the audio information into the initial video respectively to generate the target video.
18. The device according to claim 17, wherein The fourth generation submodule includes: a ninth generating unit, configured to determine a first segmentation point based on a position of a preset symbol included in the subtitle text; a tenth generating unit, configured to determine a plurality of first text paragraphs included in the subtitle text based on the first segmentation points; an eleventh generating unit, configured to, for each first text paragraph, perform sequence tagging on the first text paragraph to obtain a second segmentation point if the first text paragraph does not satisfy a display constraint condition; a twelfth generating unit, configured to determine a third dividing point from the second dividing points based on a preset strategy; and The thirteenth generating unit is configured to perform sentence segmentation processing on the subtitle text by using the first segmentation points and third segmentation points corresponding to each of the plurality of first text paragraphs to obtain the subtitle file.
19. The device according to claim 12, wherein The first generating module includes: a sixth generating submodule, configured to, in response to receiving the video generation request, acquire a text template related to the target object based on the target object entity word, wherein the text template includes a plurality of information items to be replaced; a seventh generating submodule, configured to generate the target parameter library from the initial parameter library based on the target object entity word; an eighth generating submodule, configured to determine, from the target parameter library, a plurality of target object parameters corresponding to each of the plurality of information items; a ninth generating submodule, configured to replace the plurality of information items with the plurality of target object parameters respectively to obtain the subtitle text; and The tenth generating submodule is configured to generate the speech text using the subtitle text according to the feature word detection results of each of the plurality of target object parameters.
20. The device according to claim 19, wherein The tenth generation submodule comprises: a fourteenth generating unit, configured to perform feature word detection on each target object parameter to obtain a second detection result; a fifteenth generating unit, configured to, when the second detection result indicates that the first feature word exists in the target object parameter, convert the first feature word included in the target object parameter using a mapping table to obtain a second feature word; and A sixteenth generating unit is configured to replace the first feature word in the subtitle text with the second feature word to generate the speech text.
21. The apparatus according to claim 20, further comprising: A seventeenth generating unit is configured to determine that the subtitle text is the speech text when the second detection results of each of the plurality of target object parameters indicate that the first feature word does not exist in the target object parameters.
22. The apparatus of claim 19, further comprising: a recognition module, configured to perform entity recognition on an object title included in the object information to obtain object entity words related to the object information, wherein the object information comes from the second object; and A writing module is used to write the object parameters included in the object information into the initial parameter library based on the object entity word.
23. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
25. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for generating video on basis of images, equipment and storage medium
CN107948730A
Video generation method and device, computer equipment and storage medium
CN114513706A