Video generation method and apparatus, electronic device, and storage medium
By segmenting the original video and generating narration, the problems of low efficiency and high cost in generating accessible videos are solved, and reasonable and accurate narration is generated automatically and efficiently.
Patent Information
- Application Number
- CN202310020838.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-01-06
AI Technical Summary
Existing technologies for generating accessible videos are inefficient and costly, mainly because they require manual recording of narration and voice-over.
By acquiring the original video, dividing it into multiple video segments based on the video content, and using image recognition and natural language models to generate narration, the target video is automatically generated.
It enables the automatic generation of accessible videos, improves generation efficiency, reduces production costs, and ensures the rationality and accuracy of the narration.
Smart Images

Figure CN115955585B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of image processing, and particularly relate to a video generation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] An accessible video refers to a video provided for visually impaired people. Since visually impaired people cannot normally watch image content in a video, a large number of commentary voices are added to the video in the prior art to produce a target video for visually impaired people, so as to help the visually impaired people to "understand" the video. This technology also has wide application in other related scenarios.
[0003] In the prior art, the production process of such a video is complex, and usually requires manual recording of commentary dubbing in the video, resulting in low video generation efficiency and high production cost. SUMMARY
[0004] Embodiments of the present disclosure provide a video generation method and device, an electronic device, and a storage medium to overcome the problems of low video generation efficiency and high production cost.
[0005] In a first aspect, embodiments of the present disclosure provide a video generation method, comprising:
[0006] obtaining an original video, and based on video content in the original video, cutting the original video to obtain at least two video clips; generating commentary voices corresponding to the video clips, the commentary voices being used to describe picture content of the video clips; and generating a target video according to the video clips and the corresponding commentary voices.
[0007] In a second aspect, embodiments of the present disclosure provide a video generation device, comprising:
[0008] an obtaining module configured to obtain an original video, and based on video content in the original video, cut the original video to obtain at least two video clips;
[0009] a commentary module configured to generate commentary voices corresponding to the video clips, the commentary voices being used to describe picture content of the video clips;
[0010] a generating module configured to generate a target video according to the video clips and the corresponding commentary voices.
[0011] In a third aspect, embodiments of the present disclosure provide an electronic device, comprising:
[0012] a processor, and a memory in communication connection with the processor;
[0013] the memory stores computer execution instructions;
[0014] The processor executes computer-executed instructions stored in the memory to implement the video generation method according to the first aspect and various possible designs of the first aspect.
[0015] In a fourth aspect, the embodiments of the present disclosure provide a computer-readable storage medium, and the computer-readable storage medium stores computer-executed instructions. When a processor executes the computer-executed instructions, the video generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0016] In a fifth aspect, the embodiments of the present disclosure provide a computer program product, which includes a computer program. When a processor executes the computer program, the video generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0017] The video generation method, device, electronic device, and storage medium provided by the embodiments of the present disclosure obtain an original video, split the original video into at least two video segments based on video content in the original video, generate commentary voice corresponding to the video segments, the commentary voice is used to describe the picture content of the video segment, and generate a target video according to the video segment and the corresponding commentary voice. After the original video is obtained, the original video is first split into multiple video segments according to the video content, the video segments are analyzed based on the video segments to generate corresponding commentary voice, and then the video segments and the commentary voice are synthesized to obtain the target video, so that the automatic generation from the original video to the target video is realized, the rationality and accuracy of the commentary voice in the target video are ensured, and the video generation efficiency is improved and the production cost is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor.
[0019] Figure 1 An application scenario diagram of the video generation method provided by the embodiments of the present disclosure is shown in the figure.
[0020] Figure 2 A flowchart of the video generation method provided by the embodiments of the present disclosure is shown in the figure. Figure 1 ;
[0021] Figure 3 A specific implementation step flowchart of step S101 in the embodiment shown in the figure is shown in the figure. Figure 2
[0022] Figure 4 For Figure 2 The specific implementation step flow chart of step S102 in the embodiment is shown in the figure;
[0023] Figure 5 A schematic diagram of generating a target video provided by an embodiment of the present disclosure is shown in the figure;
[0024] Figure 6 The flowchart of the video generation method provided by the embodiment of the present disclosure is shown in the figure Figure 2 ;
[0025] Figure 7 A process schematic diagram for obtaining feature similarity provided by an embodiment of the present disclosure is shown in the figure;
[0026] Figure 8 The specific implementation step flow chart of generating a corresponding explanation text in step S205 is shown in the figure;
[0027] Figure 9 The structural block diagram of the video generation apparatus provided by the embodiment of the present disclosure is shown in the figure;
[0028] Figure 10 The structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown in the figure;
[0029] Figure 11 The hardware structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown in the figure. DETAILED DESCRIPTION
[0030] To make the purposes, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.
[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0032] The application scenarios of the embodiments of the present disclosure are explained as follows:
[0033] Figure 1An application scenario diagram of the video generation method provided by the embodiments of the present disclosure, the video generation method provided by the embodiments of the present disclosure can be applied to an application scenario of online generation and playing of barrier-free videos. Specifically, the method provided by the embodiments of the present disclosure can be applied to an electronic device such as a server or a terminal device, wherein the server or the terminal device processes an ordinary video to generate a target video containing commentary voice after receiving an instruction. Specifically, as shown in Figure 1 illustrated, for example, a video resource is stored in a platform server of a video platform. After receiving a video playing request sent by a terminal device to the platform server, the platform server converts an original video in the video resource to generate a barrier-free video, and sends a video stream of the barrier-free video to the terminal device to achieve the purpose of playing the target video of the barrier-free version in the terminal device. In another possible case, the method provided by the embodiments of the present disclosure can also be applied to a terminal device, that is, the terminal device performs the above steps to complete the process of converting the original video into the barrier-free video, which will not be described here.
[0034] In the prior art, for example, the production process of such barrier-free videos containing commentary voice is complex. Unlike generating corresponding subtitles according to the voice in the video, the commentary voice in the video is the understanding result of the image picture in the video, and needs to be played at the right time to achieve the purpose of making the visually impaired clearly understand the picture content in the video through the commentary voice. Therefore, in the prior art, it is usually necessary to manually record and insert the commentary voice in the video, resulting in low video generation efficiency, high production cost and other problems.
[0035] The embodiments of the present disclosure provide a video generation method to solve the above problems.
[0036] Reference Figure 2 , Figure 2 Flowchart of the video generation method provided by the embodiments of the present disclosure Figure 1 The method of the present embodiment can be applied in a terminal device or a server, and the video generation method comprises the following steps:
[0037] Step S101: obtaining an original video, and based on the video content in the original video, the original video is cut to obtain at least two video clips.
[0038] Exemplarily, the execution subject of the embodiment can be a terminal device, such as a smart phone, or a platform server. The embodiment is described by taking the platform server as an example. Specifically, the original video is a common video. The original video can be a video resource existing in the platform server or other storage medium. The platform server acquires the original video in response to a user instruction, performs image recognition on the original video to obtain video information representing the video content of the original video, and then splits the original video based on the video information to obtain at least two video clips (i.e., at least one splitting is performed).
[0039] The video information representing the video content of the original video can include at least two video content identifiers, each of which corresponds to a video event and at least two video time stamps, so as to realize video splitting based on the video event. For example, the video information is obtained by performing image recognition on the image frames in the original video. The video information contains a set S of a group of video content identifiers S = [#001, #002, #003], wherein #001, #002, and #003 are video content identifiers representing a piece of video content, respectively. The video information also includes video time stamps corresponding to #001, #002, and #003, respectively. For example, the video time stamp corresponding to #001 is [00:00:01-00:01:03]; the video time stamp corresponding to #002 is [00:01:04-00:03:14]; and the video time stamp corresponding to #003 is [00:03:15-00:04:00]. According to the video time stamps corresponding to the video content identifiers, the original video can be split, and the video clip [00:00:01-00:01:03], the video clip [00:01:04-00:03:14], and the video clip [00:03:15-00:04:00] are obtained.
[0040] Further, as shown in Figure 3 the specific implementation of step S101 includes:
[0041] Step S1011: Acquire a target video frame and a corresponding video subtitle in the original video.
[0042] Step S1012: Determine at least two video events according to the target video frame and the corresponding video subtitle.
[0043] Step S1013: Split the original video based on the video events to obtain video clips corresponding to the video events.
[0044] Exemplarily, after obtaining the original video, by parsing the original video, the image video frames in the original video and the corresponding video subtitles can be obtained. The video subtitles can be obtained in various ways. Based on the specific implementation of the video subtitles in the original video, such as external or embedded, the video subtitles can be obtained by reading the subtitle file or by performing text recognition on the subtitle area in the video frame. Here, no specific limitation is made.
[0045] Further, after obtaining the image video frames, the video frames with more content information can be determined as target video frames, so as to better express the video content. For example, the video frames containing more character pictures and object pictures are determined as target video frames; or the image definition is high, and the relatively adjacent video frames have great changes (such as key frames), which are determined as target video frames. Subsequently, based on the target video frames obtained by screening, the video segmentation is performed, which can improve the accuracy of the video segmentation, so that the video segments obtained by cutting have relatively independent and complete video content, thereby improving the content accuracy and content integrity of the commentary voice generated based on the video segments.
[0046] Further, the video subtitles and the target video frames have a corresponding relationship, that is, the video subtitles can be used to indirectly express the picture content in the target video frames, and there is a certain relevance between them in the semantic dimension. Therefore, based on the target video frames, the corresponding video subtitles are combined to determine the corresponding video events, improve the accuracy of the division of the video events, and then complete the cutting of the original video based on the video events, which can further improve the accuracy of the picture content expressed by the obtained video segments. In a possible implementation, according to the target video frames and the corresponding video subtitles, a dialogue scene in the original video can be determined as a video event, and then the original video is cut based on the video event to obtain the corresponding video segments.
[0047] Step S102: generating commentary voice corresponding to the video segments, the commentary voice being used to describe the picture content of the video segments.
[0048] Exemplarily, after the original video is split to obtain at least two video clips, image recognition is respectively performed on each video clip to determine the picture content expressed by each video clip, for example, the picture content expressed by the video clip #1 is "two athletes are running in a race". The picture content expressed by the video clip #2 is "an athlete is receiving an award". Since the visually impaired population cannot watch the video pictures and know the video content, the target video needs to convert the picture content that the visually impaired people cannot watch into a voice, that is, a commentary voice, so as to realize the expression of the picture content in the original video. The description manner of the commentary voice to the picture content can be various, for example, first describing the environment in the picture, then describing the characters in the picture, and then describing the clothing and expression of the characters and the like information; or first describing the clothing of the characters in the picture, then describing the actions of the characters, and finally describing the environment in the picture. The specific implementation manner can be based on the content of the picture in the video clip and the description rule pre-set by the user to execute. In a possible implementation manner, the video clip can be processed based on a pre-trained natural language model, and then a sentence used to describe the picture content of the video clip, that is, a commentary voice, is output. The specific training and use process of the natural language model is not described herein.
[0049] In a possible implementation manner, as shown in Figure 4 the specific implementation manner of step S102 includes:
[0050] Step S1021: content recognition is performed on the video clip to obtain content information representing the picture content of the video clip.
[0051] Step S1022: a corresponding commentary text is generated according to the content information.
[0052] Step S1023: a commentary voice is generated according to the commentary text.
[0053] Exemplarily, the video clip is composed of multiple video frames. First, image recognition is performed on each video frame in each video clip to obtain content information representing the image content. The content information can be a feature matrix representing the image content, or a feature array or a feature identifier. Then, the content information is mapped to a natural language, that is, a commentary text, based on a natural language model. Finally, the commentary text is converted into a commentary voice according to a pre-set voice component. The specific implementation manner is not described herein.
[0054] In the process, the conversion process of "image"-"text"-"voice" is sequentially passed. Among them, the conversion from "image" to "text" is based on the image content, that is, the picture content of the video segment, and is a semantic recognition process. In the semantic recognition process, a certain amount of context information is required, that is, when the content of the video segment is recognized (step S1021), a plurality of video frames in the video segment are required to provide context information for accurate content recognition. Therefore, in order to realize the content recognition of the video segment, it is required that each video segment can provide sufficient context information to meet the requirement of the precondition, that is, the reasonable division of the video segment; if the video segment is not reasonably divided, it may not be able to provide sufficient context information, resulting in inaccurate content recognition, and thus unable to generate accurate commentary voice; in the embodiment, the original video is divided based on the video content in the original video, so that the video segment is divided based on the video event (rather than fixed time length), so that the generated video segment can provide sufficient context information in the process of generating commentary voice, so that the generated commentary text can accurately express the picture content, thereby improving the accuracy and rationality of the generated commentary voice.
[0055] Step S103: generating a target video according to the video segment and the corresponding commentary voice.
[0056] An example action is to synthesize the video segment and the corresponding commentary voice after generating the commentary voice corresponding to the video segment, to generate a video containing a commentary voice track, that is, an accessible video (target video). Figure 5 A schematic diagram for generating a target video provided by an embodiment of the present disclosure is shown in Figure 5 As shown, after content analysis of the original video, the original video is divided into video segments Video_1, Video_2, and Video_3 based on video events. Then, the picture of the video segment Video_1, Video_2, and Video_3 is analyzed respectively to generate the commentary voice Lec_1, Lec_2, and Lec_3 corresponding to the video segment Video_1, Video_2, and Video_3 respectively; then, the commentary voice Lec_1 is inserted into Video_1 to obtain re_Video_1; the commentary voice Lec_2 is inserted into Video_2 to obtain re_Video_2; the commentary voice Lec_3 is inserted into Video_3 to obtain re_Video_3. Finally, re_Video_1, re_Video_2, and re_Video_3 are combined in order to generate a target video.
[0057] In the embodiment, the original video is acquired, and the original video is cut based on the video content in the original video to obtain at least two video clips; the commentary voice corresponding to the video clip is generated, and the commentary voice is used to describe the picture content of the video clip; and the target video is generated according to the video clip and the corresponding commentary voice. Since the original video is obtained, the original video is first cut into a plurality of video clips according to the video content, the corresponding commentary voice is generated based on the video clip, and then the video clip and the commentary voice are synthesized to obtain the target video, which realizes the automatic generation from the original video to the target video, guarantees the rationality and accuracy of the commentary voice in the target video, improves the video generation efficiency, and reduces the production cost.
[0058] Reference Figure 6 , Figure 6 Flowchart of a video generation method provided by the embodiment of the disclosure Figure 2 The embodiment further refines the step S102 and adds a step of determining the playing time period of the commentary voice on the basis of the embodiment shown in Figure 2 The video generation method includes the following steps:
[0059] Step S201: An original video is acquired, and the original video is cut based on the video content in the original video to obtain at least two video clips.
[0060] Step S202: The content of the video clip is identified to obtain content information representing the picture content of the video clip.
[0061] Exemplarily, the content of the video clip is identified to obtain a feature matrix representing the picture content, that is, the content information. Then, the feature extraction is performed on the content information to obtain the semantic feature of the current picture content, that is, the second semantic feature. Exemplarily, the second semantic feature can be a low-dimensional feature of the feature matrix, for example, a feature identifier representing the type of the picture content. More specifically, the second semantic feature includes a feature identifier #002, and the picture content represented by the feature identifier #002 is “two people arguing”.
[0062] Step S203: The video voice of the video clip is acquired, and the first semantic feature of the video voice is extracted.
[0063] Step S204: The second semantic feature of the content information is acquired, and the feature similarity between the second semantic feature and the first semantic feature is compared.
[0064] Further, video audio is acquired from video clips, for example, by capturing the audio track data of the video clip. Video audio includes dialogue between characters in the scene, as well as surrounding sounds. Then, semantic recognition is performed on the video audio to obtain information representing the content features of the video audio, i.e., the first semantic feature. The first semantic feature and the second semantic feature share the same data dimensions, making feature comparison between them easier. Specifically, for example, both the first and second semantic features are feature identifiers representing the content of the scene; by comparing the feature identifiers corresponding to the first and second semantic features respectively, their feature similarity is obtained. Alternatively, for example, both the first and second semantic features are feature arrays, feature matrices, or feature structures representing the content of the scene. By comparing the feature values, feature matrices, or feature structures corresponding to the first and second semantic features, their feature similarity is obtained. Furthermore, the feature similarity can be a normalized value; the higher the feature similarity, the more similar the two are, and vice versa; when the feature similarity is 1, it means the two are identical.
[0065] Figure 7 This is a schematic diagram illustrating a process for obtaining feature similarity according to an embodiment of the present disclosure. The following is in conjunction with... Figure 7 The above process will be further explained. For example... Figure 7 As shown, for video segment A, image content recognition is performed using multiple video frames in video segment A to obtain the content information corresponding to video segment A. Then, feature extraction is performed on the content information to obtain the second semantic feature. Of course, the above step of obtaining the second semantic feature can also be completed in one step, i.e., feature extraction is performed on multiple video frames in video segment A to obtain the second semantic feature. No specific limitation is made here; it can be set as needed. This process is equivalent to the step of extracting image features from video segment A. Next, the video audio of video segment A is acquired, and feature extraction is performed on the video audio to obtain the corresponding first semantic feature. The first semantic feature represents the content in the video audio; therefore, this process is equivalent to the step of extracting audio features from video segment A. Further, after obtaining the first semantic feature (audio feature) and the second semantic feature (image feature), the two are compared to obtain the feature similarity. Then, based on this feature similarity, it is determined whether narration audio needs to be generated.
[0066] Step S205: If the feature similarity is less than the similarity threshold, then generate explanatory text based on the second semantic features of the content information.
[0067] Exemplarily, after obtaining the feature similarity, it is judged whether it is necessary to generate the explanatory voice through the feature similarity. Specifically, in the process of generating the explanatory voice corresponding to the original video, not all segments of the original video need to be "commented" on. In the case that the video voice in the original video can clearly express the picture content in the video, the visually impaired group can understand the picture content in the video through the video voice in the original video. In this case, there is no need to generate the explanatory voice. Only when the video voice in the video segment cannot express the picture content, the explanatory voice needs to be generated. The following will be described in detail with an exemplary application scenario:
[0068] For example, when the action of two people in the video picture is "arguing", the video voice in the video segment at this time can clearly express that the characters in the video are arguing at this picture. At this time, there is no need to generate the explanatory voice to inform the visually impaired user that the people in the current picture are arguing. When the action of two people in the video picture is "holding hands", the video voice in the video segment at this time is pure music (no dialogue), so it cannot express the "holding hands" action of the characters through the sound. At this time, the explanatory voice needs to be generated to inform the visually impaired user that the characters in the current picture are performing the "holding hands" action.
[0069] Based on the above introduction, when the feature similarity of the first semantic feature and the second semantic feature is greater than the similarity threshold, it indicates that the semantics expressed by the first semantic feature and the second semantic feature are almost the same. In this case, the video voice (the first semantic feature) in the video segment can express the picture content (the second semantic feature) of the video segment, so there is no need to additionally generate the explanatory voice to supplement the explanation of the picture content of the video segment. In this case, the visually impaired user can understand the picture content in the current video according to the video voice in the video segment. When the feature similarity of the first semantic feature and the second semantic feature is less than the similarity threshold, it indicates that the semantics expressed by the first semantic feature and the second semantic feature have a large difference, i.e. the video voice (the first semantic feature) in the video segment cannot express the picture content (the second semantic feature) of the video segment. At this time, the corresponding explanatory voice needs to be generated according to the picture content of the video segment, so that the visually impaired user can "understand" the picture content through the explanatory voice, i.e. the explanatory text is generated according to the second semantic feature.
[0070] Exemplarily, the content information at least includes first information and second information, wherein the first information represents a target character in the video segment, and the second information represents a target action corresponding to the target character; as shown in Figure 8 The specific implementation steps of generating the corresponding explanatory text in step S205 include:
[0071] Step S2051: generating at least one group of sub-texts according to the first information and the second information, the sub-texts being used to represent the target person and the corresponding target action in the video clip at the corresponding playing time;
[0072] Step S2052: generating the commentary text according to the sub-texts and the corresponding playing time.
[0073] Exemplarily, the second semantic feature of the content information includes the first information and the second information, wherein the first information represents the target person in the video clip, and the second information represents the target action corresponding to the target person. According to the sending time of the target action corresponding to each target person, a sub-text is generated, and each sub-text corresponds to at least one target action of a target person. For example, sub-text A corresponds to time p1, and the content of sub-text A includes "person A walked out of the door"; sub-text B corresponds to time p2, and the content of sub-text B includes "person B walked into the door". The specific content of the above-mentioned sub-text A and sub-text B, and the corresponding playing time of the text A and the sub-text B, i.e. determined by the first information and the second information in the content information, the first information and the second information contain the above-mentioned content. Then, the playing time corresponding to each group of sub-texts is merged, and the commentary text is obtained.
[0074] Step S206: determining the target speed of the commentary voice according to the content information of the video clip.
[0075] Step S207: generating the commentary voice based on the target speed and the commentary text.
[0076] Exemplarily, further, the target speed of the commentary voice is the speed when the commentary text is converted into voice broadcast. After determining the content information of the video clip, the picture content represented by the content information needs to be converted into commentary voice for playing, and the commentary voice needs to be played within a certain time to match the content in the video image, so as to avoid the problem of lagging or leading of the commentary voice relative to the video image. Therefore, before generating the commentary voice, the target speed corresponding to the commentary voice is first determined according to the content information of the video clip.
[0077] The target speed can be a unit of time representing output of each character in the commentary voice, or a specific numerical value, for example, the target speed = 5, representing that the commentary voice outputs 5 Chinese characters per second; and the target speed = 8, representing that the commentary voice outputs 8 Chinese characters per second.
[0078] Further, the target speech speed can be determined based on an information amount in the content information, specifically, the information amount is, for example, a text variable of the commentary text corresponding to the content information, the more the information amount in the content information, the faster the target speech speed, and vice versa. Thus, the content information amount of the video segment is matched with the target speech speed, when the content exhibited by the video segment is rich, more commentary text is needed to describe, and therefore a faster target speech speed is matched; and when the content exhibited by the video segment is simple, less commentary text is needed to describe, and therefore a slower target speech speed is matched.
[0079] In this embodiment, the target speech speed of the commentary speech is explained according to the content information of the video segment, for example, the text variable of the commentary text corresponding to the content information, and then the commentary speech is generated based on the target speech speed and the commentary text, so that the playback speed of the commentary speech is matched with the information amount of the content information, the problem of the commentary speech lagging behind or leading the video picture is improved, and the content reporting effect of the commentary speech is improved.
[0080] Step S208: determining a second playback time period corresponding to the commentary speech according to the first playback time period of the video speech.
[0081] Step S209: inserting the commentary speech into a target position of the original video according to the second playback time period, to generate a target video.
[0082] Exemplarily, after the commentary speech is generated, since the video segment itself can have a video speech, the generated commentary speech and the video speech need to be distinguished in time to avoid the commentary speech and the video speech being played at the same time, causing sound aliasing and affecting normal listening of the user. In this embodiment, first, a second playback time period that does not overlap with the first playback time period corresponding to the video speech in the video segment is selected, and then the commentary speech is inserted into a playback position corresponding to the second playback time period, i.e., a target position, in the video segment, to obtain a target video. Through the above steps, the commentary speech in the obtained target video does not overlap with the video speech itself in time, the commentary speech and the video speech are independently played, mutual influence is avoided, and the playback clarity of the commentary speech is improved.
[0083] In this embodiment, step S201 is consistent with step S101 in the above embodiment, and the detailed description is referred to the description of step S201, which is not repeated here.
[0084] Corresponding to the video generation method of the above embodiment, Figure 9 A structural block diagram of a video generation apparatus provided by the embodiments of the present disclosure is shown. For ease of illustration, only parts related to the embodiments of the present disclosure are shown.
[0085] Reference Figure 9The video generation apparatus 3 comprises:
[0086] The acquisition module 31 is configured to acquire an original video, and split the original video based on video content in the original video to obtain at least two video clips.
[0087] The explanation module 32 is configured to generate an explanation voice corresponding to the video clip, the explanation voice being used to describe picture content of the video clip.
[0088] The generation module 33 is configured to generate a target video according to the video clip and the corresponding explanation voice.
[0089] In an embodiment of the present disclosure, the explanation module 32 is specifically configured to: perform content recognition on the video clip to obtain content information representing picture content of the video clip; generate the corresponding explanation text according to the content information; and generate the explanation voice according to the explanation text.
[0090] In an embodiment of the present disclosure, when generating the explanation voice according to the explanation text, the explanation module 32 is specifically configured to: determine a target speech speed of the explanation voice according to the content information of the video clip; and generate the explanation voice based on the target speech speed.
[0091] In an embodiment of the present disclosure, the explanation module 32 is further configured to: acquire a video voice of the video clip, and extract a first semantic feature of the video voice; acquire a second semantic feature of the content information, and compare a feature similarity between the second semantic feature and the first semantic feature; and when generating the corresponding explanation text according to the content information, the explanation module 32 is specifically configured to: if the feature similarity is less than a similarity threshold, generate the explanation text according to the second semantic feature.
[0092] In an embodiment of the present disclosure, the content information at least comprises first information and second information, wherein the first information represents a target character in the video clip, and the second information represents a target action corresponding to the target character; and when generating the corresponding explanation text according to the content information, the explanation module 32 is specifically configured to: generate at least one group of subtexts according to the first information and the second information, the subtexts being used to represent the target character and the target action in the video clip at corresponding playing time; and generate the explanation text according to the subtexts and the corresponding playing time.
[0093] In an embodiment of the present disclosure, the generation module 33 is further configured to: acquire a video voice in the video clip; determine a second playing time period corresponding to the explanation voice according to a first playing time period of the video voice; and when generating the target video according to the video clip and the corresponding explanation voice, the generation module 33 is specifically configured to: insert the explanation voice into a target position of the original video according to the second playing time period to generate the target video.
[0094] In one embodiment of the present disclosure, the acquisition module 31, when segmenting the original video based on the video content in the original video to obtain at least two video clips, is specifically configured to: acquire target video frames and corresponding video subtitles in the original video; determine at least two video events according to the target video frames and the corresponding video subtitles; and segment the original video based on the video events to obtain video clips corresponding to the video events.
[0095] The acquisition module 31, the explanation module 32, and the generation module 33 are connected in sequence. The video generation device 3 provided in this embodiment can execute the technical solutions of the method embodiments described above, and has similar principles and technical effects, which will not be described here again.
[0096] Figure 10 A structural schematic diagram of an electronic device provided in an embodiment of the present disclosure is shown in FIG. 4, which includes: Figure 10
[0097] A processor 41 and a memory 42 connected with the processor 41 in communication;
[0098] The memory 42 stores computer execution instructions.
[0099] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the video generation method in the embodiment shown in FIG. 3. Figures 2-8
[0100] Optionally, the processor 41 and the memory 42 are connected through a bus 43.
[0101] The related descriptions can be understood by referring to the related descriptions of the steps in the corresponding embodiments, which will not be described here in detail. Figures 2-8
[0102] An embodiment of the present disclosure provides a computer readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, the computer execution instructions are configured to implement the video generation method provided in any one of the embodiments of the present disclosure. Figures 2-8
[0103] Reference can be made to the related descriptions of the steps in the corresponding embodiments. Figure 11 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0104] like Figure 11 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0105] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0106] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0107] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the above.
[0108] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled in the electronic device.
[0109] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods illustrated by the embodiments described above.
[0110] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0111] The flow diagrams and the block diagrams in the drawings are meant as methodological and functional description of implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0112] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0113] The functions described in the above description above can be performed at least in part by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0114] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0115] In a first aspect, according to one or more embodiments of the present disclosure, a video generation method is provided, comprising:
[0116] obtaining an original video, and based on video content in the original video, segmenting the original video to obtain at least two video clips; generating commentary voice corresponding to the video clips, the commentary voice being used to describe picture content of the video clips; and generating a target video according to the video clips and the corresponding commentary voice.
[0117] According to one or more embodiments of the present disclosure, the generating commentary voice corresponding to the video clips comprises: performing content recognition on the video clips to obtain content information representing picture content of the video clips; generating corresponding commentary text according to the content information; and generating the commentary voice according to the commentary text.
[0118] According to one or more embodiments of the present disclosure, the generating the commentary voice according to the commentary text comprises: determining a target speech speed of the commentary voice according to the content information of the video clips; and generating the commentary voice based on the target speech speed.
[0119] According to one or more embodiments of the present disclosure, the method further comprises: obtaining video voice of the video clips, and extracting first semantic features of the video voice; obtaining second semantic features of the content information, and comparing feature similarity between the second semantic features and the first semantic features; and generating corresponding commentary text according to the content information, comprising: if the feature similarity is less than a similarity threshold, generating the commentary text according to the second semantic features.
[0120] According to one or more embodiments of the present disclosure, the content information includes at least first information and second information, wherein the first information represents a target person in the video segment, and the second information represents a target action corresponding to the target person; generating a corresponding commentary text according to the content information includes: generating at least one group of subtexts according to the first information and the second information, the subtexts being used to represent the target person and the target action in the video segment at a corresponding playing time; and generating the commentary text according to each subtext and the corresponding playing time.
[0121] According to one or more embodiments of the present disclosure, the method further includes: obtaining a video voice in the video segment; determining a second playing time period corresponding to the commentary voice according to a first playing time period of the video voice; and generating a target video according to the video segment and the corresponding commentary voice includes: inserting the commentary voice into a target position of the original video according to the second playing time period to generate the target video.
[0122] According to one or more embodiments of the present disclosure, the cutting of the original video based on the video content in the original video to obtain at least two video segments includes: obtaining a target video frame and a corresponding video subtitle in the original video; determining at least two video events according to the target video frame and the corresponding video subtitle; and cutting the original video based on the video events to obtain a video segment corresponding to the video events.
[0123] In a second aspect, according to one or more embodiments of the present disclosure, a video generation apparatus is provided, including:
[0124] The obtaining module is configured to obtain an original video, and cut the original video based on video content in the original video to obtain at least two video segments.
[0125] The commentary module is configured to generate a commentary voice corresponding to the video segment, the commentary voice being used to describe picture content of the video segment.
[0126] The generating module is configured to generate a target video according to the video segment and the corresponding commentary voice.
[0127] According to one or more embodiments of the present disclosure, the commentary module is specifically configured to: perform content recognition on the video segment to obtain content information representing picture content of the video segment; generate a corresponding commentary text according to the content information; and generate the commentary voice according to the commentary text.
[0128] According to one or more embodiments of the present disclosure, the explanation module is specifically configured to: determine a target speech speed of the explanation voice according to the content information of the video segment when generating the explanation voice according to the explanation text; and generate the explanation voice based on the target speech speed.
[0129] According to one or more embodiments of the present disclosure, the explanation module is further configured to: acquire a video voice of the video segment, and extract a first semantic feature of the video voice; acquire a second semantic feature of the content information, and compare a feature similarity between the second semantic feature and the first semantic feature; and the explanation module is specifically configured to: generate the explanation text according to the second semantic feature if the feature similarity is less than a similarity threshold when generating the corresponding explanation text according to the content information.
[0130] According to one or more embodiments of the present disclosure, the content information at least includes first information and second information, wherein the first information represents a target character in the video segment, and the second information represents a target action corresponding to the target character; and the explanation module is specifically configured to: generate at least one group of subtexts according to the first information and the second information when generating the corresponding explanation text according to the content information, wherein each subtext represents the target character and the target action in the video segment at a corresponding playing time; and generate the explanation text according to each subtext and the corresponding playing time.
[0131] According to one or more embodiments of the present disclosure, the generation module is further configured to: acquire a video voice in the video segment; and determine a second playing time period of the explanation voice according to a first playing time period of the video voice; and the generation module is specifically configured to: insert the explanation voice into a target position of the original video according to the second playing time period to generate the target video when generating the target video according to the video segment and the corresponding explanation voice.
[0132] According to one or more embodiments of the present disclosure, the acquisition module is specifically configured to: acquire a target video frame and a corresponding video subtitle in the original video when cutting the original video based on video content in the original video to obtain at least two video segments; determine at least two video events according to the target video frame and the corresponding video subtitle; and cut the original video based on the video events to obtain a video segment corresponding to the video events.
[0133] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, which includes: a processor, and a memory connected to the processor in communication;
[0134] The memory stores computer execution instructions.
[0135] The processor executes the computer-executed instructions stored in the memory to implement the video generation method according to the first aspect and possible designs thereof.
[0136] In a fourth aspect, a computer-readable storage medium is provided, in which computer-executed instructions are stored, and when a processor executes the computer-executed instructions, the video generation method according to the first aspect and possible designs thereof is implemented.
[0137] In a fifth aspect, a computer program product is provided, which includes a computer program, and when a processor executes the computer program, the video generation method according to the first aspect and possible designs thereof is implemented.
[0138] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed range of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0139] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0140] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method of video generation, the method comprising: The method comprises: obtaining an original video, and cutting the original video into at least two video clips based on video content in the original video; generating commentary voice corresponding to the video clips, the commentary voice being used to describe picture content of the video clips; generating a target video according to the video clips and the commentary voice; the generating of the commentary voice corresponding to the video clips comprises: performing content recognition on the video clips to obtain content information representing picture content of the video clips; obtaining video voice of the video clips, and extracting first semantic features of the video voice; obtaining second semantic features of the content information, and comparing feature similarity between the second semantic features and the first semantic features; if the feature similarity is less than a similarity threshold, generating commentary text according to the second semantic features; generating the commentary voice according to the commentary text.
2. The method of claim 1, wherein, the generating of the commentary voice according to the commentary text comprises: determining a target speech speed of the commentary voice according to the content information of the video clips; and generating the commentary voice based on the target speech speed.
3. The method of claim 1, wherein, the content information comprises at least first information and second information, wherein the first information represents a target character in the video clips, and the second information represents a target action corresponding to the target character; the generating of the commentary text according to the content information comprises: generating at least one group of sub-texts according to the first information and the second information, the sub-texts being used to represent the target character and the target action in the video clips at corresponding playing time; and generating the commentary text according to the sub-texts and the corresponding playing time.
4. The method of claim 1, wherein, The method further comprises: obtaining video voice in the video clips; determining a second playing time period corresponding to the commentary voice according to a first playing time period of the video voice; the generating of the target video according to the video clips and the commentary voice comprises: inserting the commentary voice into a target position of the original video according to the second playing time period to generate the target video.
5. The method of claim 1, wherein, the cutting of the original video into at least two video clips based on video content in the original video comprises: obtaining target video frames and corresponding video subtitles in the original video; determining at least two video events according to the target video frames and the corresponding video subtitles; cutting the original video based on the video events to obtain video clips corresponding to the video events.
6. A video generating apparatus characterized by comprising: The method comprises: an obtaining module, configured to obtain an original video, and cut the original video into at least two video clips based on video content in the original video; a commentary module, configured to generate commentary voice corresponding to the video clips, the commentary voice being used to describe picture content of the video clips; a generating module, configured to generate a target video according to the video clips and the commentary voice; the commentary module, when generating the commentary voice corresponding to the video clips, is specifically configured to: perform content recognition on the video clips to obtain content information representing picture content of the video clips; The content recognition is performed on the video clip to obtain content information representing picture content of the video clip; video speech of the video clip is acquired, and a first semantic feature of the video speech is extracted; a second semantic feature of the content information is acquired, and a feature similarity between the second semantic feature and the first semantic feature is compared; If the feature similarity is less than a similarity threshold, the commentary text is generated according to the second semantic feature; and the commentary speech is generated according to the commentary text.
7. An electronic device, comprising: Comprise: a processor, and a memory connected to the processor in communication; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method of any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the video generation method of any one of claims 1 to 5 is implemented.
9. A computer program product, characterised in that, The computer program is executed by the processor to implement the video generation method of any one of claims 1 to 5.
Citation Information
Patent Citations
Voice information playing method and device, computer equipment and storage medium
CN110519636A
Video editing method, device, electronic equipment and storage medium
CN113613065A