Video generation method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202310079103.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-01-18
AI Technical Summary
但该种方法生成的视频,质量不可控,也不够灵活
[0027] In this embodiment of the disclosure, based on the analysis of the description information, the 3D scene, sub-scene, camera movement method of the sub-scene, and shot switching method required for video recording are determined, thereby generating a 3D video description file so that the 3D rendering engine can generate the video corresponding to the description information. Therefore, this embodiment of the disclosure can automatically match the content of the description information to generate high-quality video based on the understanding of the description information, without being limited to the content of the image material.
Smart Images

Figure CN116389849B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to the field of artificial intelligence technology, specifically deep learning, natural language processing, computer vision and other related technologies. Background Technology
[0002] With the increasing popularity of mobile devices such as smartphones and tablets, users' interest in watching videos has grown significantly. Manually recorded videos require investment of manpower, resources, and time.
[0003] Besides manual recording, given image materials such as video clips and pictures, the video clips and pictures can be stitched together to obtain a video. However, the quality of the video generated by this method is unpredictable and lacks flexibility. Summary of the Invention
[0004] This disclosure provides a video generation method, apparatus, device, and storage medium.
[0005] According to one aspect of this disclosure, a video generation method is provided, comprising:
[0006] Obtain descriptive information used to describe the video content;
[0007] Identify the target 3D (Three Dimensions) scene that matches the description information;
[0008] Among the multiple sub-scenes included in the target 3D scene, identify the sub-scenes that match each description fragment in the description information;
[0009] Based on the semantic analysis results of each descriptive fragment, the camera movement method for each sub-scene is determined; and,
[0010] Based on the semantic analysis results of adjacent description fragments, the camera switching method between sub-scenes of adjacent description fragments is determined;
[0011] Based on the order of each sub-scene in the description information, the camera movement of each sub-scene, the camera switching between sub-scenes, and the description information, a 3D video description file is generated.
[0012] The 3D video description file is processed using a 3D rendering engine to generate a video corresponding to the description information.
[0013] According to another aspect of this disclosure, a video generation apparatus is provided, comprising:
[0014] The acquisition module is used to acquire descriptive information that describes the video content;
[0015] The first matching module is used to determine the target 3D scene that matches the description information;
[0016] The second matching module is used to identify the sub-scenes that match each description fragment in the description information among multiple sub-scenes included in the target 3D scene.
[0017] The camera movement determination module is used to determine the camera movement method for each sub-scene based on the semantic analysis results of each descriptive fragment;
[0018] The shot switching determination module is used to determine the shot switching method between sub-scenes of adjacent description segments based on the semantic analysis results of adjacent description segments;
[0019] The file generation module is used to generate 3D video description files based on the order of each sub-scene in the description information, the camera movement of each sub-scene, the camera switching method between sub-scenes, and the description information.
[0020] The video generation module is used to process 3D video description files based on the 3D rendering engine and generate videos corresponding to the description information.
[0021] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0022] At least one processor; and
[0023] The memory is communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0025] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0026] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0027] In this embodiment of the disclosure, based on the analysis of the description information, the 3D scene, sub-scene, camera movement method of the sub-scene, and shot switching method required for video recording are determined, thereby generating a 3D video description file so that the 3D rendering engine can generate the video corresponding to the description information. Therefore, this embodiment of the disclosure can automatically match the content of the description information to generate high-quality video based on the understanding of the description information, without being limited to the content of the image material.
[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0029] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0030] Figure 1(a) is a schematic diagram of an application scenario of a video generation method according to an embodiment of the present disclosure;
[0031] Figure 1(b) is a flowchart illustrating a video generation method according to an embodiment of the present disclosure;
[0032] Figure 2 This is a flowchart illustrating the process of determining a sub-scenario according to an embodiment of this disclosure;
[0033] Figure 3 This is a schematic diagram showing the positions of the anchor point and the current analysis point according to one embodiment of this disclosure;
[0034] Figure 4(a) is a schematic diagram of a video generation method according to another embodiment of the present disclosure;
[0035] Figure 4(b) is a schematic diagram of a video generation method according to another embodiment of the present disclosure;
[0036] Figure 4(c) is a schematic diagram of the overall process of a video generation method according to an embodiment of the present disclosure;
[0037] Figure 5 This is a schematic diagram illustrating scene configuration of sub-scenes based on scene elements according to an embodiment of this disclosure;
[0038] Figure 6 This is a schematic diagram of the video generation process according to another embodiment of the present disclosure;
[0039] Figure 7 This is a schematic diagram of the structure of a video generation apparatus according to an embodiment of the present disclosure;
[0040] Figure 8 This is a block diagram of an electronic device used to implement the video generation method of the embodiments of this disclosure. Detailed Implementation
[0041] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0042] Besides manually filming, related technologies also include video generation by splicing together various materials. For example, video clips and images can be combined to create a video. However, videos generated using this method have uncontrollable quality and lack flexibility.
[0043] In view of this, this disclosure proposes a video generation method, as shown in Figure 1(a), which illustrates a scenario in which the method is applied. Figure 1(a) includes a server 11 and a terminal device 12.
[0044] The terminal device 12 and the server 11 are connected via a wireless or wired network. The terminal device 12 includes, but is not limited to, electronic devices such as desktop computers, mobile phones, portable computers, tablets, media players, smart wearable devices, and smart TVs. The server 11 can be a single server, a server cluster consisting of several servers, or a cloud computing center. The server 11 can be an independent physical server, a server cluster consisting of multiple physical servers, or a distributed system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0045] In this embodiment of the disclosure, the user can upload text to the server 11 based on the terminal device 12, and the server 11 can generate a corresponding video based on the text and send the video to the terminal device 12 via the network for the user to watch.
[0046] Figure 1(b) shows a flowchart of the video generation method proposed in this embodiment, including:
[0047] S101, Obtain description information used to describe the video content.
[0048] The descriptive information can be text, document files, audio, Uniform Resource Locators (URLs), etc. A URL can be understood as a string used to retrieve information resources on the network, primarily used in various terminal applications and server-side programs. Using URLs allows for a unified format to retrieve various information resources, including files, videos, images, and directories. The text-based descriptive information can be obtained from creative works such as news articles, novels, blogs, and academic papers, or it can be created by the user; this disclosure does not limit this aspect.
[0049] The description information can also be in audio format. Audio description information can be obtained from publicly available works with audio, such as movies, TV dramas, and radio dramas, or it can be recorded by the user. For example, recording a voice message can be used to generate a video. This embodiment of the disclosure does not limit the method of obtaining audio.
[0050] Descriptive information can also be composite resources, which include both text and image materials. For example, a travel log may include textual descriptions of the trip and scenery, as well as pictures and video clips. Based on composite resources, multimodal content understanding can be performed, thereby generating videos.
[0051] S102, Determine the target 3D scene that matches the description information.
[0052] During implementation, a 3D scene resource library can be provided, from which a target 3D scene matching the description information can be obtained.
[0053] S103, among the multiple sub-scenes included in the target 3D scene, identify the sub-scenes that match each description fragment in the description information.
[0054] The target 3D scene can be a large scene. For example, the target 3D scene can be a city, and the city's 3D scene can include sub-scenes such as theaters, concert halls, schools, residential areas, streets, and cinemas.
[0055] S104. Based on the semantic analysis results of each descriptive fragment, determine the camera movement method for each sub-scene.
[0056] The camera movement method can be determined based on the content and description of the segment.
[0057] A camera control library can be pre-built, which includes various camera movement methods. For example, when describing a lyrical story, the camera movement of the sub-scene can be a slow zoom-in; when describing a thriller, the camera movement of the sub-scene can be a sudden rise or a sudden fall, etc. It should be noted that the camera movement methods in the camera control library are not limited to those exemplified in the embodiments of this disclosure.
[0058] In implementation, training samples with descriptive information can be obtained, and the standard camera movement corresponding to the training sample can be determined as the training label. The training samples can be input into a first neural network model to be trained for semantic understanding, and this model then determines the camera movement to be compared. The output camera movement to be compared is compared with the training label to determine the loss value. The model parameters of the first neural network model to be trained are then adjusted based on the loss value until the model meets the training convergence condition, at which point training ends, resulting in a first neural network model capable of determining the camera movement of a sub-scene. This convergence condition can be that the loss value is less than a preset threshold or that a preset number of iterations is met.
[0059] In other embodiments, the camera movement within a sub-scene is not limited to a single method. The camera movement of each sub-segment can be determined based on the semantic analysis results of different sub-segments within the descriptive segment. For example, the descriptive segment can be analyzed to identify the descriptive objects, and different sub-segments can be created for each descriptive object. For instance, continuously describing the same person might form one sub-segment, while continuously describing the same cat might form another. Thus, sub-segments describing different objects can be semantically analyzed independently to obtain appropriate camera movement methods.
[0060] Camera movement techniques may include at least one of the following: slow zoom in, slow zoom out, slow panning, slow movement, slow lifting and lowering, fast zoom in, fast panning, fast movement, fast lifting and lowering, and combined movements. The camera movement technique can be determined according to actual needs during implementation.
[0061] S105, Based on the semantic analysis results of adjacent description segments, determine the camera switching method between sub-scenes of adjacent description segments.
[0062] The camera control library may also include camera switching methods for different sub-scenes. For example, when the semantic analysis result of adjacent descriptive segments indicates a switch from one location to another, the switching method between the sub-scenes can be a smooth switch; when the semantic analysis result of adjacent descriptive segments indicates a flashback, the switching method between the sub-scenes can be a blocking shot transition. It should be noted that the camera switching methods in the camera control library are not limited to the camera switching methods proposed in the embodiments of this disclosure.
[0063] Similar to determining the camera movement method of a sub-scene, in this embodiment of the disclosure, a second neural network capable of providing the camera switching method for a sub-scene can be obtained by training a second-generation training neural network.
[0064] In implementation, training samples of adjacent descriptive segments can be obtained, and the shot transition methods between standard sub-scenes corresponding to these adjacent descriptive segments can be determined as training labels. The training samples can be input into a second neural network model to be trained for semantic understanding, and this second neural network model will then determine the shot transition methods to be compared. The output shot transition methods to be compared are compared with the training labels to determine the loss value. The model parameters of the second neural network model to be trained are then adjusted based on the loss value until the second neural network model meets the training convergence condition, at which point training ends, resulting in a second neural network model capable of determining the shot transition methods between sub-scenes. The convergence condition can be that the loss value is less than a preset threshold or that a preset number of iterations is met.
[0065] The method of camera transitions needs to consider the relationship between different shots and the technical means of transition. Camera transitions can be divided into two types: technical transitions and non-technical transitions. Technical transitions include: fade-in / fade-out, overlay, freeze frame, wipe, flip page, multi-screen split, switching between reality and illusion, flick-in / flick-out, computer-generated special effects, etc. Non-technical transitions utilize the semantic relationship between the preceding and following shots to shift time and space and connect scenes, thus making the camera transitions natural and smooth, without any trace of added techniques. In implementation, the camera transition method can be determined according to actual needs.
[0066] S106. Based on the order of each sub-scene in the description information, the camera movement of each sub-scene, the camera switching method between sub-scenes, and the description information, a 3D video description file is generated.
[0067] A 3D video description file is a three-dimensional scene description file, which a 3D rendering engine can use to generate a video.
[0068] S107 processes 3D video description files based on a 3D rendering engine to generate videos corresponding to the description information.
[0069] The 3D rendering engine can be, for example, Unreal Engine, and any 3D game engine capable of generating video based on a 3D scene description file is applicable to the embodiments disclosed herein.
[0070] In this embodiment, after determining the 3D target scene of the description information, corresponding suitable sub-scenes can be determined based on the description fragments. Furthermore, the camera movement method of the sub-scenes and the shot switching method between adjacent sub-scenes can be determined based on the understanding of the description fragments. This method can automatically match the scene and camera recording method that matches the description information to generate a video, resulting in a higher degree of matching between the generated video and the description information. Since the video is generated using a 3D video description file and a 3D rendering engine, videos that meet quality requirements can be generated, and the video quality is not limited by the original material; therefore, the quality of the generated video is controllable. Moreover, even if similar description information uses the same target scene, the camera movement method and shot switching method are determined based on the semantics of the description information, making the video generation method more flexible and resulting in differences in the generated video content.
[0071] In some embodiments, if the total length of the original descriptive information exceeds a length threshold, the original descriptive information is compressed to obtain the new descriptive information. Thus, when the descriptive information is long but a shorter video needs to be generated, compression can be used to obtain descriptive information that meets the length requirements, thereby laying the data foundation for subsequent video generation.
[0072] In this embodiment of the disclosure, the original descriptive information is compressed to retain as much of the core content as possible. Key content from the original descriptive information can be extracted using semantic understanding to generate new descriptive information. For example, if the original descriptive information is a long text, it can be rewritten using a text paraphrasing model to obtain shorter text content.
[0073] In other embodiments, when the original description information is text, a text summary of the original description information is extracted to obtain the description information.
[0074] For example, if the original description information is text and the length threshold is 500 characters, and the original description information contains 1000 characters, then a summary can be extracted. Of course, the length threshold can also be a duration threshold. For instance, if the original description information is text, and an audio recording corresponding to the original description information is generated, and the length of the audio recording exceeds the duration threshold, then the total length of the original description information is determined to be greater than the length threshold.
[0075] Regardless of whether the original descriptive information is text, audio, or a composite resource, compression can be achieved by extracting a summary. Taking text as an example, in practice, a summary model can be used to extract a summary from the original descriptive information. The summary model, also known as a text compression model, is used to compress long texts into shorter texts without losing the main information, allowing readers to understand the main content of longer texts even when browsing shorter ones. Based on the different methods of generating summaries, summary models can be divided into two categories: extractive summarization models and abstractive summarization models.
[0076] Extractive summarization models refer to models that select target sentences from long texts to obtain short texts. Among them, the target sentences need to meet two requirements: (1) informativeness, which means that they contain all the important information in the long text and are logically consistent with it; (2) redundancy, which means that they have the minimum redundancy, that is, sentences containing similar information should not appear in the short text at the same time, and sentences containing unimportant information should not appear in the short text.
[0077] Extractive summarization is often regarded as a sequence labeling task. In an extractive summarization model, the encoder performs binary classification on each clause of a long text. When decoding each clause, it combines the semantic information of the clause, the decoding state of the previous clause, and the global semantic information of the long text to determine whether the current clause should be selected as the target clause.
[0078] Generative summarization models can automatically generate short texts from key information in long texts using an encoder. Similar to extractive summarization models, generative summarization still needs to meet the two requirements mentioned above: informativeness and redundancy. Generative summarization models typically employ a seq2seq (sequence-to-sequence) framework, which, based on the semantic information of long texts, adaptively selects effective contextual information through an attention mechanism to generate coherent and consistent summaries word by word.
[0079] In this embodiment of the disclosure, when the original description information is longer than the length threshold and is text, the original description information can be compressed based on the extraction of a summary. This allows the generated video to adaptively adjust its length based on the compressed text while ensuring that the semantics of the description information remain unchanged. At the same time, the generated video can also express the core content of the original description information.
[0080] In other embodiments, when the original description information is audio, the text corresponding to the audio can be obtained; a text summary can be extracted from the text corresponding to the audio to obtain the description information.
[0081] When the description information is audio, speech recognition technology can be used to convert the audio into text, and then the original description information can be compressed based on the aforementioned method of extracting text summaries to obtain the compressed description information.
[0082] In this embodiment, since the original audio data may be too long, converting it to text would exceed the length threshold. Therefore, the descriptive information can be compressed based on summary extraction. This yields descriptive information that expresses the key content of the original audio, thereby generating a video of controllable length, while still conveying the content of the original audio.
[0083] Regarding the composite resources mentioned above, it can be determined whether the length of the text in the composite resource exceeds the length threshold. If it exceeds the length threshold, the method of extracting the summary described above can be used to compress it. The image elements in the composite resource can be used as scene materials required to generate the video.
[0084] In some embodiments, the generated video requires audio corresponding to the text description information. When the description information is text, audio of the description information can be generated; wherein the playback duration of the generated video matches the playback duration of the audio. That is, in this embodiment of the disclosure, the playback duration of the generated video is adapted to the duration of the audio of the text description information.
[0085] In the case where the descriptive information is text, audio corresponding to the descriptive information can be generated based on text-to-speech (TTS) technology to constrain the playback length of the generated video.
[0086] In this embodiment of the disclosure, by generating audio corresponding to the description information, the correspondence between audio and video can be realized, providing audio data support for video generation.
[0087] In some embodiments, since the source and format of the original description information are not limited, the content of the original description information may contain some noise. For example, the original description information may contain some advertisements. In order to ensure the quality of the generated video, in this embodiment of the disclosure, the advertisement content in the original description information can be removed.
[0088] Advertising content may include advertising images, watermarks of information sources, and advertising QR codes unrelated to the description of the information.
[0089] By removing advertisements from the original description information, the description information becomes more accurate in describing the video, free from interference from other factors, thus laying a data foundation for generating high-quality videos.
[0090] In implementation, ad recognition technology and watermark detection technology can be used to detect ad-related information in the description information, performing ad recognition and watermark detection separately. The order of ad recognition and watermark detection is not important. Ad recognition can detect keywords in the description information. For example, if keywords such as "product" or "goods" are detected in a clause of the description information, semantic understanding based on the context of the keywords can be performed to determine whether the clause is an advertisement. If the clause is an advertisement, it will be deleted.
[0091] In this embodiment of the disclosure, the watermark can refer to the source of text or a copyright mark, such as the manufacturer's logo, which can be understood as a watermark to be detected. After determining the location of the watermark area in the description information, a rectangular target area containing the watermark is determined. The vertices of the rectangular target area are related to the coordinates of the watermark. The watermark pixels within the rectangular target area are identified, and the pixel values of the watermark pixels within the area are adjusted to the background color pixel values. Based on this method, the watermark can be effectively removed.
[0092] Understandably, when the text portion of the description does not contain a watermark area, i.e., there is no situation where the watermark obscures the description, there is no need to perform the watermark removal process.
[0093] Regarding the selection of 3D scenes, this embodiment of the disclosure provides a 3D scene resource library, which includes multiple candidate scenes. To facilitate the selection of target 3D scenes suitable for the description information from the 3D scene resource library, text tags can be added to each candidate scene in the 3D scene resource library. Based on the obtained description information, determining the target 3D scene that matches the description information can be implemented as follows: determining the similarity between the description information and the text tags of each candidate scene; selecting the candidate scene with the highest similarity as the target 3D scene that matches the description information.
[0094] Taking descriptive information as an example, a Natural Language Processing (NLP) model can be used to process the descriptive information, thereby obtaining the similarity between the descriptive information and the text labels of each candidate scene. If the similarity is greater than a preset threshold, then the candidate scene corresponding to the semantic label can be used as the target 3D scene of the descriptive information.
[0095] When the description information is audio, the corresponding text information can be obtained, and the corresponding text description information can be obtained. Then, NLP technology can be used to obtain the matching target 3D scene.
[0096] When the descriptive information is a composite resource, a 3D target scene can be obtained based on multimodal content understanding. For example, text features matching the text in the resource and image features matching the image material in the resource can be extracted separately. Then, feature fusion can be performed on these two features to obtain fused features. Finally, matching processing can be performed based on the fused features and the text labels of candidate scenes to obtain the corresponding 3D target scene.
[0097] In this embodiment of the disclosure, the target 3D scene suitable for the description information can be filtered out by determining the similarity between the text tags of the description information and the candidate scene. The filtered target 3D scene meets the requirements of the description information in terms of semantics, thereby accurately determining the target 3D scene suitable for video recording.
[0098] For example, descriptive information is used to describe the content of financial interviews. The target 3D scene can be defined as a city street interview scene to facilitate the adaptation of financial descriptive information.
[0099] For example, if the descriptive information is a literary story, and the environment is described, a corresponding environmental scene can be matched. For instance, when describing the structure of a historical building, the interior of a relevant 3D model of that building can be matched as the scene for recording the video. Thus, by matching a suitable scene, the target 3D scene and the descriptive information are made compatible in content, improving the quality of the generated video.
[0100] Based on the obtained target 3D scene, since the target 3D scene includes multiple sub-scenes, and the description information also describes different content based on different description fragments, it is necessary to determine the sub-scenes that match each description fragment in the description information. This can be implemented as follows: Figure 2 As shown:
[0101] S201, Obtain multiple anchor points for description information.
[0102] The videos generated from the descriptive information can be categorized into long videos and short videos. For example, videos with a length of no more than n minutes can be called short videos, while videos with a length of more than n minutes can be called long videos, where n is a positive integer. In practice, n can be 2 or 5, and can be configured according to actual needs; this disclosure does not limit the value of n.
[0103] When the generated video is a short video, the anchor point can be a sentence anchor point. For example, the descriptive information can be divided into multiple sentences, with the beginning of each sentence serving as the sentence anchor point. In practice, multiple sentences can be divided based on punctuation marks.
[0104] Considering that frequent scene switching can lead to a decline in viewing experience, in this embodiment of the disclosure, all anchor points obtained in S201 are key anchor points, regardless of whether it is a long video or a short video.
[0105] For example, given sentence anchors, a preset filtering strategy is used to select key anchors from multiple sentence anchors as the multiple anchors required in S201.
[0106] For example, for each sentence anchor, sentences within the range of m sentences of that sentence anchor can be considered as neighboring sentences of that sentence anchor. In implementation, the semantics of neighboring sentences can be analyzed, and sentences with semantic similarity greater than a preset threshold can be merged to obtain a merged sentence set. The sentence anchor at the middle position of the merged sentence set can be taken as the anchor required in S201. Alternatively, the global semantics of the merged sentence set can be determined as the reference semantics, and the semantic similarity between each sentence in the merged sentence set and the reference semantics can be determined. The mean of the semantic similarity between each sentence and the reference semantics can be calculated, as shown in expression (1). Then, from the merged sentence set, sentences with semantic similarity to the reference semantics that is the mean can be determined, and the sentence anchor of the sentence can be used as the anchor required in S201. Of course, the sentence with the highest semantic similarity to the reference semantics can also be selected, and the sentence anchor of the sentence can be used as the anchor required in S201. Based on this method, relatively sparse key anchors can be obtained, thereby avoiding the waste of resources and avoiding frequent and excessive switching of scenarios.
[0107]
[0108] In expression (1), O i Let ai represent the mean semantic similarity of the i-th sentence in the merged sentence set, a1 represent the semantic similarity between the i-th sentence and the reference semantics, and a2 represent the semantic similarity between the i-th sentence and the reference semantics. n Let represent the semantic similarity between the i-th sentence and the reference semantics, and n represent the number of sentences in the merged sentence set.
[0109] S202, extract the first segment of each anchor point from the description information.
[0110] In some embodiments, since there are multiple anchor points in the description information, for each anchor point, content of a preset radius length is extracted from the description information with the anchor point as the center to obtain the first segment of the anchor point.
[0111] For example, the preset radius length can be measured by the number of sentences. For instance, if the preset radius length is k sentences, where k is a positive integer, then the k sentences before the anchor point and the k sentences after the anchor point, along with the sentence containing the anchor point, can be used as the first segment.
[0112] The preset radius length can also be described by duration. For example, if the anchor point is located at the 5th second and its preset radius length is 3 seconds, then the content within 2 to 8 seconds of the description information is extracted as the first segment of the anchor point.
[0113] It should be noted that the preset radius length can be set based on actual conditions, and this embodiment does not limit it. If the content before or after the anchor point is less than the preset radius length, all content less than the preset radius length is acquired. For example, if the anchor point is at position 5s, its preset radius length is 3s, and only 2s of content remain after the anchor point, then that 2s of content is acquired.
[0114] In this embodiment, the range of the first segment is determined by taking the anchor point as the center and combining the context. This operation is simple, requires no complex calculations, and can reduce the consumption of computer resources.
[0115] S203, perform semantic analysis on the first segment of each anchor point to obtain sub-scenes that match each anchor point respectively.
[0116] Since the target 3D scene includes multiple sub-scenes, in order to facilitate the selection of sub-scenes that match each anchor point from the target 3D scene, text labels can be added to each sub-scene of the target 3D scene.
[0117] In some embodiments, Natural Language Processing (NLP) techniques may be used to determine the sub-scene corresponding to the text tag that best matches the description fragment.
[0118] For example, if the target 3D scene is a panoramic view of a zoo, then the sub-scenes included within this target 3D scene could include a lion display sub-scene with the text tag "lion," a tiger display sub-scene with the text tag "tiger," a panda display sub-scene with the text tag "panda," and so on. If the first segment corresponding to anchor point A is "In the zoo, we saw a fierce tiger," then the semantic tag of this first segment can be determined as "tiger," and thus the tiger display sub-scene can be determined as a sub-scene matching anchor point A.
[0119] Since the first segment may not fully interpret the content of the same sub-scene, in order to find the accurate switching points between different sub-scenes in this embodiment of the disclosure, in S204, natural language processing technology is used to analyze the semantic turning points in each second segment between adjacent anchor points in the description information.
[0120] This semantic inflection point, the point where the semantics change, corresponds to different sub-scenes. Therefore, the semantic inflection point can be used to separate the descriptive information between two anchor points, enabling the switching of sub-scenes.
[0121] In some embodiments, the content before the first anchor point in the description information can be rendered using the sub-scene corresponding to the first anchor point, and the content after the last anchor point in the description information can be rendered using the sub-scene corresponding to the last anchor point. For the content between adjacent anchor points, the switching position of the sub-scene needs to be clearly defined. For the second segment between each adjacent anchor point in the description information, natural language processing techniques are used to analyze the semantic turning points in each second segment, which can be implemented as follows:
[0122] Step A1: For each second segment, obtain the current analysis point in the second segment and the content of a specified length before the current analysis point from the description information to obtain the reference information of the current analysis point.
[0123] The current analysis point is any point within the second segment. This can be understood as a point in time during video playback, or as any sentence position within the text.
[0124] The specified length is similar to the preset radius length of the first segment obtained above, and can be either the duration or the number of sentences. This embodiment does not limit this.
[0125] Step A2: Determine the semantic similarity between the reference information and the first and second anchor points in the second segment, respectively. The positions of the first anchor point, the second anchor point, and the current analysis point are as follows: Figure 3 As shown, in the description information, the first anchor point is located before the second anchor point.
[0126] Step A3: If the semantic similarity between the current analysis point and the first anchor point is lower than the semantic similarity between the current analysis point and the second anchor point, and the semantic similarity between the previous analysis point and the first anchor point is higher than the semantic similarity between the current analysis point and the second anchor point, then determine the current analysis point as a semantic turning point in the second segment.
[0127] In short, if the semantic similarity between any analysis point and anchor point A is higher than that of another anchor point, then the analysis point is applicable to the sub-scene of anchor point A. If the sub-scene applicable to the current analysis point is different from that of the previous analysis point, then the current analysis point is a semantic turning point.
[0128] Taking any second segment as an example, with the current analysis point at 8s and a specified length of 3s, the content from 5 to 8s can be determined as reference information for the current analysis point. Since the second segment is defined by two adjacent anchor points, assuming the first anchor point is 3s and the second anchor point is 10s, the semantic similarity between the reference information and the first and second anchor points in the second segment can be determined. NLP techniques can be used to semantically understand the reference information, obtain its corresponding sub-scenes, and determine the semantic similarity with the sub-scenes. If the semantic similarity between the current analysis point and the first anchor point is 50%, the semantic similarity between the current analysis point and the second anchor point is 90%, and the semantic similarity between the previous analysis point and the first anchor point is 98%, then the current analysis point is determined to be a semantic turning point in the second segment.
[0129] In this embodiment of the disclosure, semantic turning points can be accurately found based on the second segment and semantic similarity, thereby laying a data foundation for accurately segmenting the descriptive segments used for different sub-scenes, which can improve the quality of the generated video.
[0130] In other embodiments, the maximum semantic similarity between the content in the second segment and the first anchor point can also be determined. If the semantic similarity between the current analysis point in the second segment and the first anchor point is lower than this maximum value, and the difference between the two values is greater than a threshold, then a semantic inflection point at the current analysis point can be determined.
[0131] In addition to using the first anchor point as a benchmark, a second anchor point can also be used to identify semantic inflection points. Furthermore, the maximum semantic similarity between the content in the second segment and the second anchor point can be determined. If the semantic similarity between the current analysis point and the second anchor point in the second segment is lower than this maximum value, and the difference between the two values exceeds a threshold, then the semantic inflection point at the current analysis point can be identified.
[0132] S205, the content between two semantic inflection points adjacent to the same anchor point in the description information is determined as a description fragment.
[0133] S206, the sub-scene that matches the anchor points included in the description fragment is determined as a sub-scene of the description fragment.
[0134] In this embodiment of the disclosure, the descriptive information is divided into multiple descriptive segments based on semantic inflection points, and then the sub-scene corresponding to each descriptive segment is determined based on each descriptive segment. This division method can accurately understand the initial position and end position of each sub-scene, making the generated video closer to the descriptive information.
[0135] In some embodiments, when the sub-scene corresponding to the description fragment is determined, the layout of the sub-scene can also be adjusted based on the description fragment. This can be implemented as follows: for each description fragment, obtain the preset questions of the sub-scene of the description fragment; perform machine reading comprehension on the description fragment to obtain the answers corresponding to each preset question; in the 3D scene element set, select the elements of the sub-scene and the attribute information of the elements of the sub-scene based on the answers to each preset question.
[0136] Taking the target 3D scene as a panoramic view of a zoo as an example, the first segment corresponding to anchor point B is "The teacher led the students to visit the zoo. The teacher was first in line, then Xiaoming was second, and behind Xiaoming was Xiaohong. Xiaoming was wearing blue clothes, and Xiaohong was wearing red clothes." Therefore, the semantic label of this first segment can be determined as "queueing up to play" within the zoo. Consequently, the sub-scene of "queueing up to play" can be determined as the sub-scene matching anchor point B. The students in this scene can be either male or female, and their clothing can also differ. Therefore, the pre-set questions for this scene can be "Who are the characters included in the scene?" and "What are the physical characteristics of the characters?" Based on these pre-set questions, machine reading and comprehension (MRC) technology can be used to semantically understand the descriptive segment and find the answers, thus identifying the characters as students queuing up, and the student characteristics as the colors of each student.
[0137] It should be noted that the scene elements within a sub-scene can be known elements from a 3D resource library, and these elements can be pre-defined 3D images. Each sub-scene is associated with some elements, and the preset question for each sub-scene is a question posed to its associated elements. This allows the required scene elements within the sub-scene to be analyzed based on the description fragments.
[0138] For example, when broadcasting a weather forecast, a preset question could be: "What is the gender of the anchor?" This allows you to select a suitable anchor to set in a sub-scene.
[0139] For example, in a cultural performance scene, the pre-set questions can ask about the required musical instruments and the characteristics of the characters who use the instruments. Then, the 3D elements of the corresponding musical instruments are laid out in the sub-scene, and the characters who operate the instruments are also laid out.
[0140] When scene elements in a 3D resource library cannot accurately describe a descriptive segment, scene elements can be generated using generative techniques, such as Generative Adversarial Networks (GANs). The resulting video files can be more vivid and better match the description in the descriptive file, without being limited by the materials in the 3D resource library.
[0141] Furthermore, users can specify whether to use materials from a 3D resource library for the description information. For example, some literary works have similar descriptions, and scene materials can be obtained entirely through a generative approach, thus allowing videos with as much stylistic difference as possible to be produced from different descriptions.
[0142] For example, multiple shots can be placed in the same sub-scene. Based on the understanding of the descriptive information, the sub-scene can be shot from different camera angles, which will also produce different visual effects.
[0143] In this embodiment of the disclosure, based on a preset question, the features of scene elements in the description information can be obtained, thereby enabling an accurate understanding of the scene elements corresponding to the layout of the description information, and thus providing support for generating high-quality video resources.
[0144] Since the video generated based on the description information in this embodiment is a 3D scene, lighting is required regardless of whether the 3D scene is indoor or outdoor. In order to use appropriate lighting, in this embodiment, for each description segment, when the lighting conditions are not explicitly recorded in the description segment, sentiment analysis is performed on the description segment to obtain the sentiment analysis result; from the lighting set, the lighting conditions that match the sentiment analysis result are selected as the lighting conditions used for the sub-scene of the description segment.
[0145] Taking the target 3D scene as a panoramic view of Financial Street as an example, the sub-scenes are street interview A and Financial Street data display scene B. The semantic analysis result of the descriptive fragment is the financial report of Company C, and the sub-scene corresponding to this descriptive information is Financial Street data display scene B. In this sub-scene, the street sign at the street corner can be replaced with "Company C's financial report," and Financial Street data display scene B will display a bar chart of Company C's recent revenue data. Based on NLP, a sentiment analysis is performed on the descriptive fragment to obtain the sentiment analysis results. Given that the sentiment analysis results indicate optimism, and since many people are happier on sunny days and somewhat distressed on cloudy days, the weather in sub-scene B can be set to sunny, and the time to noon.
[0146] In this embodiment of the disclosure, the emotional tendency of the text is determined based on the descriptive fragment, and then the lighting conditions are set so that the generated video is closer to the content expressed by the descriptive information.
[0147] In some embodiments, intro, outro, intro theme, outro theme, and background music can be added to the generated video. When it is necessary to beautify it, filters can be applied to parts of the video.
[0148] In summary, the video generation method proposed in this disclosure can realize AI (Artificial Intelligence) directing, determine scenes, characters, and recording strategies (i.e., 3D video description files), and then generate videos. The overall process can be divided into two steps: "getting into position" and "shooting." The "getting into position" process includes selecting a target scene from a 3D scene resource library and determining sub-scenes. Then, for each sub-scene, scene elements are determined, and scene construction is performed in three steps. A schematic diagram is shown in Figure 4(a). Semantic understanding of the description information is performed to determine the target scene. Since the target scene includes multiple sub-scenes, semantic understanding of each description fragment in the description information is required to determine the sub-scene matching each description fragment. After determining the sub-scenes, scene elements in the sub-scenes can be selected and arranged based on the semantic information of the description fragments, and then scene construction is completed based on the scene elements and their arrangement.
[0149] The “shooting” process includes two steps: “blueprint planning” and “shooting”. As shown in Figure 4(b), the “blueprint planning” realizes that the timeline in the audio describing the information connects the sub-scene shots of each frame and includes the spatial positioning of the camera position. That is, it determines the camera movement method of the sub-scene based on the switching method between different sub-scenes.
[0150] The above is the preliminary work preparation, which will eventually yield a 3D video description file.
[0151] For example, suppose the target 3D scene corresponding to the 3D video description file (sequence) is "city", sub-scene A is captured by camera_a, and sub-scene B is captured by camera_b. If the camera's attribute adds an "output" label, it means that the camera's output content is used to generate video. The entire process is wrapped by multiple keyframes, each keyframe representing a series of object attribute values (such as position, rotation angle, etc.) at that point in time. One of the camera's attributes is "transition," which indicates the transition method (such as fly-in, direct cut, etc.). In this example, the 3D video description file (sequence) format is as follows:
[0152] <sequence scene="city"> / / Target 3D scene is a city
[0153] <key-frame time=”0”> / / Start time of sub-scene A
[0154] <object output id="camera_a"poistion="-63164.158798,39656.181052,73.000096"rotation="0"transition="">< / object> / / The camera used in sub-scene A is camera_a, which has attributes such as camera position, angle, and transition method;
[0155] <object id="host"poistion="-63162.198708,12315.181137,0.000096"rotation="0"transition="">< / object> / / Information such as the host's position and angle in sub-scene A;
[0156]
[0157] <key-frame time="5"> / / 5s time in sub-scene A
[0158] <object output id="camera_a"poistion="-63164.158798,39656.181052,73.000096"rotation="-27"transition="fly-in">< / object> / / Subscene A uses camera properties at time 5 seconds
[0159]
[0160] <key-frame time="20"> / / Start time of sub-scene B
[0161] <object output id="camera_b"poistion="-39142.009,812364.829102,28.1902"rotation="0"transition="cut-in">< / object> / / Subscene B uses camera properties
[0162]
[0163]
[0164] Similarly, through a series of descriptions, the camera and shooting angles used in each sub-scene are recorded in sequence, and the set-up sub-scenes are photographed. The resulting 3D video description file can clearly express the visuals, content, and methods of shooting, thus forming a clear recording strategy.
[0165] After the preliminary work was completed, Unreal Engine (UE) simulated the movement of a real photographer and the behavior of a director based on 3D video description files, and shot between different camera positions, as shown in Figure 4(b). Four key anchor points are shown on the timeline for example. Each anchor point corresponds to its own sub-scene.
[0166] Anchor point 1 corresponds to sub-scene camera position A, which is the host operating an interactive screen; anchor point 2 corresponds to sub-scene camera position B, which is a large screen with a digital human figure; anchor point 4 corresponds to sub-scene camera position C, which displays 3D data statistics; anchor point 3 corresponds to sub-scene camera position D, which is three 3D billboards. Switching between these four camera positions completes the scene movement and transitions, ultimately generating the video. For example, sub-scene camera position A records the host operating an interactive screen displaying World Cup-related reports, allowing the host to select which players to report on. Then, the camera switches to sub-scene camera position B. Camera position B focuses on the large screen with the digital human figure, displaying the players' performances in the World Cup, their resumes, and their development records. The camera then switches to the billboards in sub-scene camera position D, showcasing World Cup sponsors, and finally to sub-scene camera position C, displaying World Cup statistical data analysis. Thus, multiple sub-scenes are interconnected, generating corresponding content based on the descriptive information.
[0167] To gain a more detailed understanding of the video generation method proposed in this disclosure, taking text as the descriptive information as an example, its overall flowchart is shown in Figure 4(c):
[0168] S401, Obtain the input text and remove the advertising content from the input file.
[0169] S402, determine the target scene based on the input text and the scene resource library.
[0170] S403, determine the sub-scene corresponding to the text fragment based on the text fragment of the input text.
[0171] S404, determine scene elements in sub-scenes based on question-answering models.
[0172] S405, configures sub-scenes based on scene elements.
[0173] In the sub-scene that includes screen display, such as the large screen in camera position B in Figure 4(b). Figure 5 The materials displayed on this screen can come from an existing material library, be generated based on text, or originate from other material sources. When multiple materials exist on the large screen, these materials are sorted and distributed according to a timeline to obtain the sub-scene for scene configuration.
[0174] For example, in a scenario involving an interview with a famous soccer player, their goal highlights could be displayed on a large screen. However, due to potential copyright issues with some images, the materials displayed on the screen could be based on… Figure 5 The materials are obtained in the manner shown, and then sorted and distributed according to the timeline to obtain the configuration of the large screen in the sub-scene.
[0175] Even within the same scene, due to differences in camera angles and other factors, adjusting the scene layout and camera angles in conjunction with the text content can generate richer content, rather than simply copying existing content, thus reducing reliance on existing materials. Furthermore, based on the embodiments of this disclosure, content at different resolutions can be generated without depending on the resolution of the original content itself.
[0176] S406: For each sub-scene, generate the corresponding camera movement method based on the description fragment of that sub-scene.
[0177] S407, for adjacent sub-scenes, determine the camera switching method between sub-scenes.
[0178] S408 generates a 3D video description file.
[0179] S409 uses Unreal Engine to render 3D video description files and generate videos.
[0180] In summary, the overall video generation process can be considered a 3D film planning and shooting process led by AI (Artificial Intelligence). This process begins with scene selection and prop placement based on an understanding of the descriptive information, fully utilizing spatial layout and set design for scene arrangement. During implementation, not only can existing scene elements be selected, but scene elements can also be generated based on specific needs, ultimately achieving customization of scene elements to match the descriptive information. A timeline connects the transitions between spatial sub-scenes, using different camera angles within each sub-scene to achieve a 3D "live-action" shooting effect. Existing images and videos can be used to generate necessary clips for the scene, which can still be used in the film, and these existing clips can be displayed more naturally within the scene elements.
[0181] Taking the input description information as an article as an example, the entire video generation process is as follows: Figure 6As shown. After inputting an article, users select a corresponding scene from the scene library based on the article's content. For example, financial and social news articles can choose a city scene, sports reports can choose a studio scene, and essays can choose a natural landscape scene. After selecting a main scene, assuming a city scene is chosen, sub-scenes corresponding to different descriptive fragments can be selected based on the understanding of the article's content. In the scene library, sub-scenes exist in the form of templates. For example, financial articles can choose the Financial Street template, social news articles can choose the residential area template, and future-related articles can automatically generate scene templates based on the content if no suitable template is found. Of course, descriptive fragments that do not match a suitable sub-scene can generate corresponding sub-scenes based on the content of the descriptive fragment. Assuming the selected template is the Financial Street template, the descriptive fragments can be used to ask and answer questions to arrange the scene. For example, it can ask what kind of anchor is used, the emotional tone of the text, the time of the event, and for financial reports, it can summarize and analyze the profit situation and generate corresponding charts from the analyzed data for display in the sub-scenes. Thus, different sub-scenes are arranged and can be pieced together in chronological order. For the image materials needed for sub-scenes, not only can we rely on a material library, but we can also use existing image materials from the article, or even generate videos based on a video compositing platform, or generate materials using text-to-image conversion. For example, the introduction of a sports star can automatically generate a video based on publicly available content and display it on a large screen. According to the understanding of the article, the corresponding elements are placed and displayed in appropriate positions. After determining the camera movement methods for different sub-scenes and the transitions between sub-scenes, a 3D video description file can be generated. The 3D rendering engine renders the 3D video description file and saves the output video, thus completing the video recording. The recorded video can be sent to users via a link for download and viewing.
[0182] In terms of post-production, users can add intros and outros to the generated videos and apply filter effects. However, the overall video generation process is completed automatically by the AI director, and users only need to make minor adjustments and optimizations according to their own needs.
[0183] Based on the same technical concept, this disclosure provides a video generation apparatus 700, such as... Figure 7 The following are included:
[0184] The acquisition module 701 is used to acquire descriptive information used to describe the video content;
[0185] The first matching module 702 is used to determine the target 3D scene that matches the description information;
[0186] The second matching module 703 is used to determine the sub-scene that matches each description fragment in the description information among multiple sub-scenes included in the target 3D scene.
[0187] The camera movement determination module 704 is used to determine the camera movement method of each sub-scene based on the semantic analysis results of each description fragment;
[0188] The camera switching determination module 705 is used to determine the camera switching method between sub-scenes of adjacent description segments based on the semantic analysis results of adjacent description segments;
[0189] The file generation module 706 is used to generate a 3D video description file based on the order of each sub-scene in the description information, the camera movement of each sub-scene, the camera switching method between sub-scenes, and the description information.
[0190] The video generation module 707 is used to process 3D video description files based on the 3D rendering engine and generate videos corresponding to the description information.
[0191] In some embodiments, the second matching module includes:
[0192] The acquisition unit is used to acquire multiple anchor points for descriptive information;
[0193] The first segment determination unit is used to extract the first segment of each anchor point from the description information;
[0194] The first matching unit is used to perform semantic analysis on the first segment of each anchor point to obtain the sub-scenes that match each anchor point respectively.
[0195] The turning point determination unit is used to analyze the semantic turning points in each second segment between adjacent anchor points in the descriptive information using natural language processing techniques.
[0196] The description fragment determination unit is used to determine the content between two semantic inflection points adjacent to the same anchor point in the description information as a description fragment;
[0197] The second matching unit is used to identify sub-scenes that match the anchor points included in the description fragment as sub-scenes of the description fragment.
[0198] In some embodiments, a scene element determination module is further included, for:
[0199] For each description fragment, obtain the preset questions for the sub-scenes of the description fragment;
[0200] The machine reads and understands the descriptive passage to obtain the answers to each preset question.
[0201] In the 3D scene element set, the elements of the sub-scene and the attribute information of the elements in the sub-scene are selected based on the answers to each preset question.
[0202] In some embodiments, the inflection point determination unit is configured to:
[0203] For each second segment, the current analysis point in the second segment and the content of a specified length before the current analysis point are obtained from the description information to obtain the reference information of the current analysis point;
[0204] Determine the semantic similarity between the reference information and the first anchor point and the second anchor point in the second segment, respectively, wherein the first anchor point is located before the second anchor point in the description information;
[0205] If the semantic similarity between the current analysis point and the first anchor point is lower than the semantic similarity between the current analysis point and the second anchor point, and the semantic similarity between the previous analysis point and the first anchor point is higher than the semantic similarity between the current analysis point and the second anchor point, then the current analysis point is determined to be a semantic turning point in the second segment.
[0206] In some embodiments, the first segment determination unit is configured to:
[0207] For each anchor point, extract the content of a preset radius length from the description information centered on the anchor point to obtain the first segment of the anchor point.
[0208] In some embodiments, the acquisition module is configured to:
[0209] If the total length of the original description information exceeds a length threshold, the original description information is compressed to obtain the new description information.
[0210] In some embodiments, the acquisition module is further configured to:
[0211] When the original description information is text, extract the text summary of the original description information to obtain the description information.
[0212] In some embodiments, the acquisition module is further configured to:
[0213] If the original description information is audio, obtain the text corresponding to the audio.
[0214] Extract text summaries from the text corresponding to the audio to obtain descriptive information.
[0215] In some embodiments, an audio determination module is further included, for:
[0216] If the description information is text, an audio version of the description information is generated; wherein the playback duration of the generated video matches the playback duration of the audio.
[0217] In some embodiments, the first matching module is configured to:
[0218] Determine the similarity between the descriptive information and the text labels of each candidate scene;
[0219] The candidate scene with the highest similarity is selected as the target 3D scene that matches the description information.
[0220] In some embodiments, an ad removal module is also included, for:
[0221] Remove advertising content from the original description information.
[0222] In some embodiments, an illumination determination module is further included, for:
[0223] For each descriptive segment, in the absence of explicit recording of lighting conditions, sentiment analysis was performed on the descriptive segment to obtain the sentiment analysis results;
[0224] From the set of lighting conditions, lighting conditions that match the sentiment analysis results are selected as the lighting conditions used for the sub-scenes of the description segment.
[0225] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0226] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0227] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0228] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0229] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0230] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0231] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as video generation methods. For example, in some embodiments, the video generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the video generation method by any other suitable means (e.g., by means of firmware).
[0232] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0233] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0234] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0235] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0236] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0237] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0238] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0239] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video generation method, comprising: Obtain descriptive information used to describe the video content; Identify the target 3D scene that matches the description information; Among the multiple sub-scenes included in the target 3D scene, the sub-scenes that match each description fragment in the description information are identified, including: Obtain multiple anchor points of the description information; Extract the first segment of each anchor point from the description information; Semantic analysis is performed on the first segment of each anchor point to obtain the sub-scenes that match each anchor point respectively; For the second segment between each adjacent anchor point in the description information, natural language processing technology is used to analyze the semantic turning points in each second segment; The content between two semantic inflection points adjacent to the same anchor point in the description information is determined as a description fragment; Sub-scenes that match the anchor points included in the description fragment are identified as sub-scenes of the description fragment; Based on the semantic analysis results of each descriptive fragment, the camera movement method for each sub-scene is determined; and, Based on the semantic analysis results of adjacent description segments, the camera switching method between the sub-scenes of the adjacent description segments is determined; Based on the order of each sub-scene in the description information, the camera movement of each sub-scene, the camera switching method between sub-scenes, and the description information, a 3D video description file is generated. The 3D video description file is processed using a 3D rendering engine to generate the video corresponding to the description information.
2. The method according to claim 1, further comprising: For each description segment, obtain the preset questions for the sub-scene of the description segment; The description fragment is subjected to machine reading comprehension to obtain the answers to each preset question; In the 3D scene element set, the elements of the sub-scene and the attribute information of the elements of the sub-scene are selected based on the answers to each preset question.
3. The method according to claim 1, wherein, For the second segment between each adjacent anchor point in the description information, natural language processing techniques are used to analyze the semantic transition points in each of the second segments, including: For each second segment, the current analysis point in the second segment and the content of a specified length before the current analysis point are obtained from the description information to obtain the reference information of the current analysis point; Determine the semantic similarity between the reference information and the first anchor point and the second anchor point in the second segment, respectively, wherein the first anchor point is located before the second anchor point in the description information; If the semantic similarity between the current analysis point and the first anchor point is lower than the semantic similarity between the current analysis point and the second anchor point, and the semantic similarity between the previous analysis point of the current analysis point and the first anchor point is higher than the semantic similarity between the current analysis point and the second anchor point, then the current analysis point is determined to be the semantic turning point in the second segment.
4. The method according to claim 1, wherein, Extract the first segment of each anchor point from the description information, including: For each anchor point, with the anchor point as the center, extract content of a preset radius length from the description information to obtain the first segment of the anchor point.
5. The method according to claim 1, further comprising: If the total length of the original description information is greater than a length threshold, the original description information is compressed to obtain the new description information.
6. The method according to claim 5, wherein, The original description information is compressed, including: If the original description information is text, extract the text summary of the original description information to obtain the description information.
7. The method according to claim 5, wherein, The original description information is compressed, including: If the original description information is audio, obtain the text corresponding to the audio. The text summary corresponding to the audio is extracted to obtain the descriptive information.
8. The method according to claim 1, further comprising: If the description information is text, an audio version of the description information is generated; wherein the playback duration of the generated video matches the playback duration of the audio.
9. The method according to claim 1, wherein, Determining a target 3D scene that matches the description information includes: Determine the similarity between the description information and the text tags of each candidate scene; The candidate scene with the highest similarity is selected as the target 3D scene that matches the description information.
10. The method according to claim 1, further comprising: Remove the advertising content from the original description information.
11. The method according to any one of claims 1-10, further comprising: For each descriptive segment, in the absence of explicitly recording lighting conditions, sentiment analysis is performed on the descriptive segment to obtain the sentiment analysis results; From the set of lighting conditions, lighting conditions that match the sentiment analysis results are selected as the lighting conditions used for the sub-scene of the description segment.
12. A video generation apparatus, comprising: The acquisition module is used to acquire descriptive information that describes the video content; The first matching module is used to determine the target 3D scene that matches the description information; The second matching module is used to determine, among the multiple sub-scenes included in the target 3D scene, the sub-scenes that match each description fragment in the description information, including: The acquisition unit is used to acquire multiple anchor points of the description information; The first segment determination unit is used to extract the first segment of each anchor point from the description information; The first matching unit is used to perform semantic analysis on the first segment of each anchor point to obtain the sub-scenes that match each anchor point respectively. The turning point determination unit is used to analyze the semantic turning points in each of the second segments between each adjacent anchor point in the description information using natural language processing technology. The description fragment determination unit is used to determine the content between two semantic inflection points adjacent to the same anchor point in the description information as a description fragment; The second matching unit is used to determine the sub-scenes that match the anchor points included in the description fragment as sub-scenes of the description fragment; The camera movement determination module is used to determine the camera movement method for each sub-scene based on the semantic analysis results of each descriptive fragment; The camera switching determination module is used to determine the camera switching method between sub-scenes of adjacent description segments based on the semantic analysis results of adjacent description segments; The file generation module is used to generate a 3D video description file based on the order of each sub-scene in the description information, the camera movement method of each sub-scene, the camera switching method between sub-scenes, and the description information. The video generation module is used to process the 3D video description file based on the 3D rendering engine and generate the video corresponding to the description information.
13. The apparatus according to claim 12, further comprising a scene element determination module, configured to: For each description segment, obtain the preset questions for the sub-scene of the description segment; The description fragment is subjected to machine reading comprehension to obtain the answers to each preset question; In the 3D scene element set, the elements of the sub-scene and the attribute information of the elements of the sub-scene are selected based on the answers to each preset question.
14. The apparatus according to claim 12, wherein, The turning point determination unit is used for: For each second segment, the current analysis point in the second segment and the content of a specified length before the current analysis point are obtained from the description information to obtain the reference information of the current analysis point; Determine the semantic similarity between the reference information and the first anchor point and the second anchor point in the second segment, respectively, wherein the first anchor point is located before the second anchor point in the description information; If the semantic similarity between the current analysis point and the first anchor point is lower than the semantic similarity between the current analysis point and the second anchor point, and the semantic similarity between the previous analysis point of the current analysis point and the first anchor point is higher than the semantic similarity between the current analysis point and the second anchor point, then the current analysis point is determined to be the semantic turning point in the second segment.
15. The apparatus according to claim 12, wherein, The first segment determination unit is used for: For each anchor point, with the anchor point as the center, extract content of a preset radius length from the description information to obtain the first segment of the anchor point.
16. The apparatus according to claim 12, wherein the acquiring module is configured to: If the total length of the original description information is greater than a length threshold, the original description information is compressed to obtain the new description information.
17. The apparatus according to claim 16, wherein, The acquisition module is also used for: If the original description information is text, extract the text summary of the original description information to obtain the description information.
18. The apparatus according to claim 16, wherein, The acquisition module is also used for: If the original description information is audio, obtain the text corresponding to the audio. The text summary corresponding to the audio is extracted to obtain the descriptive information.
19. The apparatus of claim 12, further comprising an audio determination module, configured to: When the description information is text, audio of the description information is generated; wherein, The playback duration of the generated video matches the playback duration of the audio.
20. The apparatus according to claim 12, wherein, The first matching module is used for: Determine the similarity between the description information and the text tags of each candidate scene; The candidate scene with the highest similarity is selected as the target 3D scene that matches the description information.
21. The apparatus of claim 12, further comprising an ad removal module, configured to: Remove the advertising content from the original description information.
22. The apparatus according to any one of claims 12-21, further comprising an illumination determination module, for: For each descriptive segment, in the absence of explicitly recording lighting conditions, sentiment analysis is performed on the descriptive segment to obtain the sentiment analysis results; From the set of lighting conditions, lighting conditions that match the sentiment analysis results are selected as the lighting conditions used for the sub-scene of the description segment.
23. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
25. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Method for quickly generating short video based on 3D scene and related device
CN114286197A
Video generation method and device, electronic equipment and storage medium
CN114567819A
Method for generating digital human, model training method, device, equipment and medium
CN115082602A