Video script generation method and device
By comprehensively considering the human body characteristics and voiceprint characteristics of the speaker in the video, combining the line text information, matching the target role from the character library and generating a video script, the problem of lack of character introduction or poor recognition of scripts in the existing technology is solved, and a more accurate reflection of video content is achieved.
Patent Information
- Application Number
- CN202510496117.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-01
AI Technical Summary
When generating video scripts, the prior art fails to effectively consider the human body characteristics and voiceprint characteristics of the speaker, resulting in the scripts lacking character introduction content or poor recognition effect.
By performing speech recognition of voice clips in the target video, the human body characteristics and voiceprint characteristics of the speaker are extracted, the target feature information is matched from the character library, and the video script is generated based on the line text information.
The generated video script can accurately reflect the video content, improve the understanding of the video, and fully reflect the character lines and background descriptions.
Smart Images

Figure CN120238703A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of AI (Artificial Intelligence), specifically to technical fields such as artificial intelligence, big data, and deep learning, and particularly relates to a method and apparatus for generating a video script. Background Art
[0002] With the continuous explosion of short videos produced, in actual working scenarios, video scripts are of great auxiliary significance for further analysis and understanding of the produced videos, as well as for generating exciting highlight content subsequently. The content of videos is complex and diverse. Generally speaking, the information describing the characters appearing in the video is the core content of the script.
[0003] Therefore, how to generate a video script becomes particularly important. Summary of the Invention
[0004] The present disclosure provides a method and apparatus for generating a video script.
[0005] According to one aspect of the present disclosure, there is provided a method for generating a video script, including: obtaining a target video; performing speech recognition on a speech segment in the target video to obtain corresponding line text information; extracting the human body features of the speaker from the image frame associated with the speech segment in the target video, and extracting the voiceprint features of the speaker from the speech segment; determining, from a character library corresponding to the target video, target feature information that matches at least one of the human body features of the speaker and the voiceprint features of the speaker, and determining a target character corresponding to the target feature information in the character library; generating a video script of the target video according to the target character and the line text information.
[0006] According to another aspect of the present disclosure, there is provided a video script generating apparatus, including: a first obtaining module, configured to obtain a target video; a first determining module, configured to perform speech recognition on a speech segment in the target video to obtain corresponding line text information; an extracting module, configured to extract the human body features of the speaker from the image frame associated with the speech segment in the target video, and extract the voiceprint features of the speaker from the speech segment; a second determining module, configured to determine, from a character library corresponding to the target video, target feature information that matches at least one of the human body features of the speaker and the voiceprint features of the speaker, and determine a target character corresponding to the target feature information in the character library; a generating module, configured to generate a video script of the target video according to the target character and the line text information.
[0007] According to still another aspect of the present disclosure, there is provided an electronic device, including:
[0008] At least one processor; and
[0009] A memory communicatively connected to the at least one processor; wherein
[0010] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the video script generation method provided in one aspect of the present disclosure above.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the video script generation method provided in one aspect of the present disclosure above.
[0012] According to still another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the video script generation method provided in one aspect of the present disclosure above.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0015] Figure 1 is a schematic flowchart of the video script generation method provided in Embodiment 1 of the present disclosure;
[0016] Figure 2 is a schematic flowchart of the video script generation method provided in Embodiment 2 of the present disclosure;
[0017] Figure 3 is a schematic flowchart of the video script generation method provided in Embodiment 3 of the present disclosure;
[0018] Figure 4 is a schematic flowchart of the video script generation method provided in Embodiment 4 of the present disclosure;
[0019] Figure 5 is a schematic flowchart of the video script generation method provided in Embodiment 5 of the present disclosure;
[0020] Figure 6 is a schematic structural diagram of the video script generation device provided in Embodiment 6 of the present disclosure;
[0021] Figure 7 shows a schematic block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed Implementation Modes
[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0023] Generally speaking, the content of videos is complex and diverse. There may be multiple characters in a video, and different characters can appear at different times in the video. Even the same character can appear in different segments of the video. The segments where the character appears are important content that makes up the video, and the line text expressed by the character is also very important information for understanding the video. In order to make the generated script truly reflect the core content of the video, how to generate a video script is very important.
[0024] In the related art, in order to generate a video script, the following several methods are usually adopted:
[0025] In the first method, when generating a video script, during the process of identifying the speaker who describes the line text through voice in the video, the role played by the speaker and the introduction information of the role are not taken into account.
[0026] The second method is to identify the characters in the video only by human body features alone, or only by voiceprint features to identify the characters in the video.
[0027] It should be noted that when generating a video script by the above methods, there are at least the following problems:
[0028] For the first method, ignoring the information of the role played by the speaker makes the generated script lack the content of character introduction, which will affect the generated script's inability to summarize all the core content of the video;
[0029] For the second method, only identifying the speaker in the video by human body features alone, or only by voiceprint features, there can be multiple characters in the video, and for a character, both human body features and voiceprint features are very important features. Therefore, if the human body features and voiceprint features are not considered comprehensively, it will affect the recognition effect of the characters.
[0030] For the convenience of understanding, first, relevant concepts that may be involved in the embodiments of the present application are briefly described:
[0031] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.
[0032] Figure 1 It is a schematic flowchart of the video script generation method provided by Embodiment 1 of the present disclosure. As Figure 1 shown, the method includes:
[0033] In this embodiment of the present disclosure, the video script generation method is configured in a video script generation device as an example. The video script generation device can be applied to any electronic device so that the electronic device can perform the video script generation function.
[0034] Among them, the electronic device can be any device with computing power, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens. In this embodiment, the server is taken as an example to illustrate the process of the video script generation method.
[0035] Step 101, obtain the target video.
[0036] Among them, it can be understood that the target video includes voice segments and silent segments.
[0037] It should be noted that in the actual application scenario, the target video is usually not a continuous voice video. In order to make the content presented by the produced target video logical and interesting, the target video can have segmented voice segments, and there can also be silent segments between different voice segments. Among them, the characters with voice output line texts in the voice segments are the core information of the content presented by the target video. Therefore, focusing on the voice segments of the target video is crucial for the generation of the video script.
[0038] Step 102, perform speech recognition on the voice segments in the target video to obtain the corresponding line text information.
[0039] Among them, it should be noted that different voice segments can have different numbers of characters and characters with different voice output line texts. Therefore, the corresponding line text information obtained from different voice segments is also different.
[0040] Step 103, extract the human features of the speaker from the image frames associated with the voice segments in the target video, and extract the voiceprint features of the speaker from the voice segments.
[0041] Among them, it can be understood that the speaker is the person who indicates the speech output line text in the speech segment.
[0042] It should be understood that when extracting image features from the image frames associated with the speech segment in the target video, multiple human body features can be extracted. However, not all the characters appearing in the target video have speech line text. Therefore, not all the extracted human body features belong to the speaker. Although voiceprint extraction is performed on the speech segment to determine the extracted voiceprint features from the speaker in the speech segment, there can be multiple speakers in the speech segment. Thus, further analysis of the extracted voiceprint features is required.
[0043] For example, when extracting image features from the image frames associated with the speech segment in the target video, multiple human body features are extracted. Clustering is performed on the multiple human body features. Suppose three human body features are determined, namely human body feature A, human body feature B, and human body feature C. Voiceprint extraction is performed on the speech segment. Suppose two voiceprint features are extracted, namely voiceprint feature A and voiceprint feature B. Among them, there is a corresponding relationship between human body feature A and voiceprint feature A, which come from the same person. There is a corresponding relationship between human body feature B and voiceprint feature B, which come from the same person. Then there are speaker A and speaker B in this speech segment. The human body feature of speaker A is human body feature A and the voiceprint feature of speaker A is voiceprint feature A; the human body feature of speaker B is human body feature B and the voiceprint feature of speaker B is voiceprint feature B.
[0044] Step 104: Determine, from the character library corresponding to the target video, the target feature information that matches at least one of the human body feature of the speaker and the voiceprint feature of the speaker, and determine the target character corresponding to the target feature information in the character library.
[0045] Among them, it should be noted that the character library includes the feature information corresponding to each character.
[0046] For example, the character library includes character A, character B, and character C. Suppose the human body feature A of speaker A and the voiceprint feature A of speaker A are determined, and there is a mutual match between human body feature A and the feature information A corresponding to character A. Then the feature information A can be used as the target feature information, and the target character (character A) corresponding to the target feature information (feature information A) in the character library can be determined.
[0047] Step 105: Generate the video script of the target video according to the target character and the speech line text information.
[0048] It should be understood that there can be multiple speech segments in the target video. The determined target character and line text information come from one speech segment in the target video. The target character and line text information determined from different speech segments are different, and the corresponding relationship between the target character and line text information in the same speech segment is relatively diverse. Therefore, in order to quickly generate the video script of the target video, the target video can be analyzed segment by segment based on the speech segments to determine the corresponding relationship between each line of the line text information of the target character and line text information in the same speech segment, and then determine the video script of any speech segment, so as to generate the video script of the target video.
[0049] In summary, the video script generation method proposed in this disclosure obtains the target video; performs speech recognition on the speech segments in the target video to obtain the corresponding line text information; extracts the human body features of the speaker from the image frames associated with the speech segments in the target video, and extracts the voiceprint features of the speaker from the speech segments; determines, from the character library corresponding to the target video, target feature information that matches at least one of the human body features of the speaker and the voiceprint features of the speaker, and determines the target character corresponding to the target feature information in the character library; generates the video script of the target video according to the target character and the line text information. Thus, by extracting image features from each frame of the target video, not only the human body features of the speaker are considered, but also the voiceprint features of the speaker are considered. Through the character library, the human body features and voiceprint features are comprehensively used to identify the target characters appearing in the video, so as to accurately identify the characters appearing in the video, and in the process of generating the script, from the perspective of the target characters in the video, focus on the line text information described by the target characters, so that the generated video script can truly reflect the content of the video, which helps to deepen the understanding of the video, and thus facilitates various applications based on the generated video script.
[0050] It should be noted that in the technical solution of this disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processes are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0051] To illustrate how to extract the human body features of the speaker from the image frames associated with the speech segments in the target video and extract the voiceprint features of the speaker from the speech segments in this disclosure, this disclosure also proposes a video script generation method.
[0052] Figure 2 It is the flowchart of the video script generation method provided in the second embodiment of this disclosure.
[0053] As Figure 2As shown in the figure, the video script generation method may include the following steps 201 to 207:
[0054] Step 201, obtain a target video.
[0055] Step 202, perform speech recognition on the speech segment in the target video to obtain the corresponding line text information.
[0056] For the explanations of steps 201 to 202, reference may be made to the relevant descriptions in the embodiments of the present disclosure, and details are not described herein again.
[0057] Step 203, for any speech segment in the target video, extract the voiceprint of the speech segment to determine the voiceprint feature of the speaker.
[0058] It should be noted that the voiceprint feature extracted from the speech segment and the speaker have a one-to-one correspondence relationship, that is, the number of extracted voiceprint features is the same as the number of speakers.
[0059] For example, when extracting the voiceprint of a speech segment and two voiceprint features are extracted, namely voiceprint feature A and voiceprint feature B, the speaker corresponding to voiceprint feature A (speaker A) is different from the speaker corresponding to voiceprint feature B (speaker B), and there is a one-to-one correspondence relationship between voiceprint feature A and speaker A; there is a one-to-one correspondence relationship between voiceprint feature B and speaker B.
[0060] Step 204, determine a first image frame from multiple image frames associated with the speech segment; wherein, the human body movement in the first image frame meets the set requirements of the speaker.
[0061] It should be noted that the set requirements are determined according to actual needs. For example, the set requirements can be set as the human body movement with an open mouth as the facial movement of the human body.
[0062] In order to accurately determine the first image frame from multiple image frames associated with the speech segment, as a possible implementation, according to the human body region in any image frame of the multiple image frames, determine the first image frame whose human body movement meets the set requirements through the human body movement recognized within the human body region.
[0063] As an example, the image frames synchronously displayed with the speech segment are used as the multiple image frames associated with the speech segment; perform target recognition on any image frame of the multiple image frames to obtain the human body region within the corresponding image frame and perform action recognition on the human body region to obtain the human body movement; determine the first image frame that meets the set requirements from the multiple image frames according to the human body movement recognized within each image frame.
[0064] Step 205, extract the image features of the first image frame to obtain the human body features of the speaker.
[0065] It should be understood that there may be multiple human bodies within the human body region in the first image frame, and not all of the multiple human bodies are from the speaker. Therefore, it is also necessary to further screen the human body region in the first image frame to determine the human body region where the speaker is located.
[0066] As a possible implementation manner, in order to accurately obtain the human body characteristics of the speaker, the human body characteristics of the speaker are determined according to the human body region in the first image frame where the human body movements meet the set requirements.
[0067] As an example, image feature extraction is performed on the target human body region in the first image frame to obtain the human body characteristics of the speaker, where the target human body region is the human body region where the human body movements meet the set requirements.
[0068] In order to perform action recognition on the human body region and completely obtain the human body movements within the human body region, as a possible implementation manner, the human body movements are determined according to the positional relationship between multiple human body key points recognized within the human body region.
[0069] As an example, human body key point recognition is performed on the human body region; the human body movements are determined according to the positional relationship between the multiple human body key points obtained by the recognition.
[0070] Step 206: Determine, from the role library corresponding to the target video, target feature information that matches at least one of the human body characteristics of the speaker and the voiceprint characteristics of the speaker, and determine the target role corresponding to the target feature information in the role library.
[0071] Step 207: Generate a video script for the target video according to the target role and the line text information.
[0072] For the explanation of steps 206 to 207, reference can be made to the relevant descriptions in the embodiments of the present disclosure, which will not be elaborated here.
[0073] In the video script generation method according to the embodiments of the present disclosure, for any voice segment in the target video, voiceprint extraction is performed on the voice segment to determine the voiceprint characteristics of the speaker, and the first image frame is determined from multiple image frames associated with the voice segment; where the human body movements in the first image frame meet the set requirements of the speaker, and image feature extraction is performed on the first image frame to obtain the human body characteristics of the speaker. Thus, the voiceprint characteristics of the speaker are extracted first, and then the first image frame where the speaker is located is gradually determined from multiple image frames. Furthermore, according to the first image frame where the speaker is located, the human body characteristics of the speaker are obtained, and multiple image frames are gradually analyzed, so that not only can the speaker be accurately determined, but also the human body characteristics of the speaker can be quickly extracted.
[0074] To illustrate how to determine, in the character library corresponding to the target video, target feature information that matches at least one of the human features of the speaker and the voiceprint features of the speaker, and to determine the target character corresponding to the target feature information in the character library, the present disclosure also proposes a video script generation method.
[0075] Figure 3 It is a flowchart of the video script generation method provided in Embodiment 3 of the present disclosure.
[0076] As Figure 3 shown, the video script generation method may include the following steps 301 to 307:
[0077] Step 301, obtain the target video.
[0078] Step 302, perform speech recognition on the voice segment in the target video to obtain the corresponding line text information.
[0079] Step 303, extract the human features of the speaker from the image frames associated with the voice segment in the target video, and extract the voiceprint features of the speaker from the voice segment.
[0080] For the explanatory descriptions of Steps 301 to 303, reference can be made to the relevant descriptions in the embodiments of the present disclosure, and details will not be elaborated here.
[0081] Step 304, obtain the feature information to be matched among the human features of the speaker and the voiceprint features of the speaker.
[0082] It should be noted that the specific information of the feature information to be matched can be determined according to the actual situation, and no specific limitation is made in this embodiment. For example, the specific information of the feature information to be matched can be divided into three cases, namely, the feature information to be matched includes the human features of the speaker and the voiceprint features of the speaker, the feature information to be matched only includes the human features of the speaker, and the feature information to be matched only includes the voiceprint features of the speaker.
[0083] Step 305, determine the similarity between each feature information in the set of feature information included in the character library and the feature information to be matched.
[0084] It should be noted that the set of feature information includes the feature information corresponding to each character in the character library.
[0085] Step 306, according to the similarity, determine the target feature information whose similarity exceeds the similarity threshold from the set of feature information, and determine the target character corresponding to the target feature information in the character library.
[0086] Among them, it should be noted that the specific value of the similarity threshold can be determined according to the actual situation, and this embodiment does not make specific limitations.
[0087] For example, in the feature information set in the character library, there are feature information A and feature information B. Feature information A has a corresponding relationship with character A, and feature information B has a corresponding relationship with character B. Assuming that the similarity between the feature information to be matched and feature information A exceeds the similarity threshold, then feature information A can be used as the target feature information, and the target character (character A) corresponding to the target feature information (feature information A) in the character library can be determined.
[0088] Step 307, generate a video script for the target video according to the target character and the line text information.
[0089] For the explanation of step 307, reference can be made to the relevant description in the embodiments of the present disclosure, and details are not described herein.
[0090] In the video script generation method of the embodiments of the present disclosure, the feature information to be matched among the human body features of the speaker and the voiceprint features of the speaker is obtained. In the feature information set included in the character library, the similarity between each feature information in the feature information set and the feature information to be matched is determined. According to the similarity, the target feature information whose similarity exceeds the similarity threshold is determined from the feature information set, and the target character corresponding to the target feature information in the character library is determined. Thus, the human body features of the speaker and the voiceprint features of the speaker are comprehensively considered, so that the target feature information matching the feature information to be matched is determined from the feature information set through the similarity between each feature information and the feature information to be matched. Furthermore, the accuracy of obtaining the target character is improved, and through the similarity, it is convenient to quickly locate the target feature information from multiple feature information.
[0091] To illustrate how to generate a video script for the target video according to the target character and the line text information in the embodiments of the present disclosure, the present disclosure also proposes a video script generation method.
[0092] Figure 4 It is a flowchart of the video script generation method provided in Embodiment 4 of the present disclosure.
[0093] As Figure 4 shown, the video script generation method may include the following steps 401 to 406:
[0094] Step 401, obtain the target video.
[0095] Step 402, perform speech recognition on the voice segment in the target video to obtain the corresponding line text information.
[0096] Step 403: Extract the human body features of the speaker from the image frames associated with the voice segments in the target video, and extract the voiceprint features of the speaker from the voice segments.
[0097] Step 404: Determine, from the character library corresponding to the target video, the target feature information that matches at least one of the human body features of the speaker and the voiceprint features of the speaker, and determine the target character corresponding to the target feature information in the character library.
[0098] For the explanations of Steps 401 to 404, reference can be made to the relevant descriptions in the embodiments of the present disclosure, and details are not elaborated herein.
[0099] Step 405: Determine the target character and the line text information that come from the same voice segment according to the voice segment where the target character is located and the voice segment where the line text information is located.
[0100] It should be noted that the target video segment also includes silent segments. For the video script, the scene-by-scene script is also an indispensable content. Therefore, in order to make the generated video script more abundant and be able to include all the important contents in the target video, in addition to focusing on the voice segments, the background description information of the silent segments can also be extracted to form a part of the video script.
[0101] As a possible implementation manner to obtain a complete video script, when there is a background switch in the silent segment relative to the adjacent voice segment, the background description information of the silent segment is inserted into the video script.
[0102] As an example, perform background recognition on the adjacent voice segments between the silent segments; when there is a background switch in the silent segment relative to the adjacent voice segment, perform semantic recognition on the background of the silent segment to obtain the background description information of the silent segment; insert the background description information into the video script according to the appearance order of the silent segment in the target video.
[0103] Step 406: Arrange the target characters and the line text information of each voice segment in sequence according to the appearance order of each voice segment in the target video to obtain the video script.
[0104] It should be noted that the line text information described by the target character is the core content of the video script. However, the introduction text of the target character itself also helps to deepen the understanding of the target character that appears in the target video. Therefore, during the process of generating the video script, the introduction text information of the target character can be associated with the line text information so that it can be intuitively understood which target character and which line text correspond in the subsequent generated video script.
[0105] To obtain a more perfect video script, as a possible implementation, through a character library, the introduction text information corresponding to the target character is supplemented into the video script.
[0106] As an example, according to the target character, query the character library to obtain the introduction text information corresponding to the target character; wherein, the character library includes the introduction text information corresponding to each character; according to the position of the line text information of the target character in the video script, the introduction text information is supplemented into the video script.
[0107] The video script generation method of the present disclosure embodiment determines the target character and the line text information from the same voice segment according to the voice segment where the target character is located and the voice segment where the line text information is located, and arranges the target character and the line text information of each voice segment in sequence according to the appearance order of each voice segment in the target video to obtain a video script. Thus, the target character and the line text information of each voice segment are arranged to form a video script according to the appearance order of each voice segment in the target video, focusing on the core content of the video script, so that the generated video script is more complete and can deeply reflect the content of the target video.
[0108] To illustrate how the character library is generated in the present disclosure embodiment, the present disclosure also proposes a video script generation method.
[0109] Figure 5 It is a flowchart of the video script generation method provided in the fifth embodiment of the present disclosure.
[0110] As Figure 5 shown, the video script generation method may include the following steps 501 to 505:
[0111] Step 501, perform human body recognition on each frame of video image in the target video to determine multiple human body features corresponding to the target video.
[0112] It should be understood that since the same character can appear in multiple video images, among the multiple human body features obtained by performing human body recognition on each frame of video image in the target video, there may be approximate human body features, and these approximate human body features may come from the same character. For determining the feature information corresponding to the character, too many approximate human body features are redundant and will also slow down the progress of generating the character library.
[0113] Step 502, perform voiceprint extraction on the target video to determine multiple voiceprint features corresponding to the target video.
[0114] It should be understood that since there can be multiple speaking characters in the target video, and the speaking characters can appear in different segments of the target video, when extracting voiceprint features from the target video, there may be approximate voiceprint features among the multiple obtained voiceprint features, and these approximate voiceprint features can come from the same speaking character. For determining the feature information corresponding to the character, too many approximate voices are redundant and will also affect the progress of generating the character library.
[0115] Step 503: Cluster the multiple voiceprint features corresponding to the target video to obtain at least one voiceprint feature category, and cluster the multiple human body features corresponding to the target video to obtain at least one human body feature category.
[0116] It should be noted that the voiceprint feature category is used to indicate the character corresponding to the voiceprint feature, and the human body feature category is used to indicate the character corresponding to the human body feature.
[0117] For example, it is determined that the multiple voiceprint features corresponding to the target video are voiceprint feature a, voiceprint feature b, and voiceprint feature c respectively. Among them, voiceprint feature a is approximate to voiceprint feature b. Then, cluster voiceprint feature a, voiceprint feature b, and voiceprint feature c to obtain the voiceprint feature category AB formed by clustering voiceprint feature a and voiceprint feature b, and the voiceprint feature category C formed by clustering voiceprint feature c. Among them, the voiceprint feature category AB is used to indicate the character AB corresponding to voiceprint feature a and the voiceprint feature; the voiceprint feature category C is used to indicate the character C corresponding to voiceprint feature c.
[0118] For example, it is determined that the multiple human body features corresponding to the target video are human body feature a, human body feature b, human body feature c, and human body feature d respectively. Among them, human body feature c and human body feature d are approximate. Then, cluster human body feature a, human body feature b, human body feature c, and human body feature d to obtain the human body feature category A formed by clustering voiceprint feature a; the human body feature category B formed by clustering voiceprint feature b; the human body feature category CD formed by clustering human body feature c and human body feature d.
[0119] Step 504: Determine the feature information corresponding to each character according to the voiceprint features belonging to the same voiceprint feature category and the human body features belonging to the same human body feature category.
[0120] To accurately obtain the feature information corresponding to each character, as a possible implementation, determine the feature information of the corresponding character according to the voiceprint feature category and the human body feature category that belong to the same character among each voiceprint feature and each human body feature.
[0121] As an example, determine the speech segments from which the voiceprint features within each voiceprint feature category are derived, and determine the video images from which the human body features within each human body feature category are derived; based on the co-occurrence relationship between the speech segments and the video images, determine the voiceprint feature categories and human body feature categories that belong to the same role; based on the voiceprint feature categories and human body feature categories that belong to the same role, determine the feature information corresponding to the role.
[0122] In this embodiment, the voiceprint features at the clustering center within the voiceprint feature categories that belong to the same role, and the human body features at the clustering center within the human body feature categories, are determined as the feature information corresponding to the role.
[0123] It should be noted that the specific method for determining whether a voiceprint feature is at the clustering center within a voiceprint feature category, and the specific method for determining whether a human body feature is at the clustering center within a human body feature category, can both be determined according to existing related technologies, and will not be elaborated in this embodiment.
[0124] Step 505, determine the role library according to the feature information corresponding to the role.
[0125] It should be understood that there may also be introduction text information of the role in the target video. Generally speaking, the introduction text information usually appears in the image where the role first appears in the target video. For example, in the image where role A first appears in the target video, the name of role A can be marked within the image of role A, so the name of role A can be used as the introduction text information of role A. Therefore, in order to make the data saved for each role in the role library more abundant, the introduction text information of the role can also be stored in the role library.
[0126] In this embodiment, for any role, determine the first appearance image ranked first from multiple frames of images associated with the role; perform text recognition on the first appearance image to determine the introduction text information corresponding to the role; store the introduction text information in the role library.
[0127] The video script generation method of the embodiment of the present disclosure performs human body recognition on each frame of video image in the target video to determine multiple human body features corresponding to the target video, performs voiceprint extraction on the target video to determine multiple voiceprint features corresponding to the target video, clusters the multiple voiceprint features corresponding to the target video to obtain at least one voiceprint feature category, and clusters the multiple human body features corresponding to the target video to obtain at least one human body feature category, determines feature information corresponding to each role according to voiceprint features belonging to the same voiceprint feature category and human body features belonging to the same human body feature category, and determines a role library according to the feature information corresponding to the role, thereby, according to the voiceprint feature category and human body feature category extracted by clustering from the target video, determines voiceprint features belonging to the same voiceprint feature category and human body features belonging to the same human body feature category, and then generates a role library. Through the pre-generated role library, there is no need to repeatedly match the feature information to be matched extracted from the speaker's human body features and the speaker's voiceprint features with the multiple human body features corresponding to the target video and the multiple voiceprint features corresponding to the target video to determine the target role, so that the target role can be determined quickly and accurately in the future.
[0128] With the above Figures 1 to 5 Corresponding to the video script generation method provided in the embodiment, the present disclosure also provides a video script generation device. Since the video script generation device provided in the embodiment of the present disclosure is similar to the above-mentioned Figures 1 to 5 The video script generation method provided in the embodiment corresponds to the embodiment, so the implementation method of the video script generation method is also applicable to the video script generation device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0129] Figure 6 This is a structural diagram of the video script generation device provided in Example 6 of the present disclosure.
[0130] like Figure 6 As shown, the video script generating device 600 may include: a first acquisition module 601, a first determination module 602, an extraction module 603, a second determination module 604 and a generation module 605.
[0131] A first acquisition module 601 is used to acquire a target video;
[0132] The first determination module 602 is used to perform speech recognition on the speech segment in the target video to obtain corresponding line text information;
[0133] An extraction module 603 is used to extract the speaker's body features from the image frames associated with the speech segment in the target video, and to extract the speaker's voiceprint features from the speech segment;
[0134] A second determination module 604, configured to determine, from a character library corresponding to the target video, target feature information that matches at least one of the human body features and voiceprint features of the speaker, and determine a target character corresponding to the target feature information in the character library;
[0135] A generation module 605, configured to generate a video script for the target video according to the target character and the line text information.
[0136] In a possible implementation manner of the embodiments of the present disclosure, the extraction module 603 is specifically configured to:
[0137] For any voice segment in the target video, perform voiceprint extraction on the voice segment to determine the voiceprint feature of the speaker;
[0138] Determine a first image frame from multiple image frames associated with the voice segment; wherein, the human body action in the first image frame meets the set requirements of the speaker;
[0139] Perform image feature extraction on the first image frame to obtain the human body features of the speaker.
[0140] In a possible implementation manner of the embodiments of the present disclosure, the extraction module 603 is specifically configured to:
[0141] Use the image frames synchronously displayed with the voice segment as the multiple image frames associated with the voice segment;
[0142] Perform target recognition on any image frame in the multiple image frames to obtain the human body region within the corresponding image frame, and perform action recognition on the human body region to obtain the human body action;
[0143] Determine a first image frame that meets the set requirements from the multiple image frames according to the human body actions recognized in each image frame.
[0144] In a possible implementation manner of the embodiments of the present disclosure, the extraction module 603 is specifically configured to:
[0145] Perform image feature extraction on the target human body region within the first image frame to obtain the human body features of the speaker, where the target human body region is the human body region whose human body action meets the set requirements.
[0146] In a possible implementation manner of the embodiments of the present disclosure, the extraction module 603 is further specifically configured to:
[0147] Perform human body key point recognition on the human body region;
[0148] Determine the human body action according to the positional relationship between multiple human body key points obtained by recognition.
[0149] In a possible implementation manner of the embodiment of the present disclosure, the second determination module 604 is specifically configured to:
[0150] Obtain the feature information to be matched among the human body features of the speaker and the voiceprint features of the speaker;
[0151] In the set of feature information included in the role library, determine the similarity between each piece of feature information in the set of feature information and the feature information to be matched; wherein, the set of feature information includes the feature information corresponding to each role in the role library;
[0152] According to the similarity, determine the target feature information whose similarity exceeds the similarity threshold from the set of feature information, and determine the target role corresponding to the target feature information in the role library.
[0153] In a possible implementation manner of the embodiment of the present disclosure, the generation module 605 is specifically configured to:
[0154] According to the voice segment where the target role is located and the voice segment where the line text information is located, determine the target role and the line text information that come from the same voice segment;
[0155] According to the appearance order of each voice segment in the target video, arrange the target role and the line text information of each voice segment in sequence to obtain the video script.
[0156] In a possible implementation manner of the embodiment of the present disclosure, the generation module 605 is further specifically configured to:
[0157] Perform background recognition on adjacent voice segments between silent segments;
[0158] In the case where there is a background switch of the silent segment relative to the adjacent voice segment, perform semantic recognition on the background of the silent segment to obtain the background description information of the silent segment;
[0159] According to the appearance order of the silent segment in the target video, insert the background description information into the video script.
[0160] In a possible implementation manner of the embodiment of the present disclosure, the generation module 605 is further specifically configured to:
[0161] According to the target role, query the role library to obtain the introduction text information corresponding to the target role; wherein, the role library includes the introduction text information corresponding to each role;
[0162] According to the position of the line text information of the target role in the video script, supplement the introduction text information to the video script.
[0163] In a possible implementation manner of the embodiment of the present disclosure, the device further includes a role library module, specifically configured to:
[0164] Perform human body recognition on each frame of video image in the target video to determine multiple human body characteristics corresponding to the target video;
[0165] Extract voiceprints from the target video to determine multiple voiceprint characteristics corresponding to the target video;
[0166] Cluster the multiple voiceprint characteristics corresponding to the target video to obtain at least one voiceprint characteristic category, and cluster the multiple human body characteristics corresponding to the target video to obtain at least one human body characteristic category;
[0167] Determine the characteristic information corresponding to each role according to the voiceprint characteristics belonging to the same voiceprint characteristic category and according to the human body characteristics belonging to the same human body characteristic category;
[0168] Determine the role library according to the characteristic information corresponding to the role.
[0169] In a possible implementation manner of the embodiment of the present disclosure, the role library module is further specifically configured to:
[0170] Determine the voice segments from which the voiceprint characteristics within each voiceprint characteristic category are derived, and determine the video images from which the human body characteristics within each human body characteristic category are derived;
[0171] Determine the voiceprint characteristic category and the human body characteristic category belonging to the same role according to the co-occurrence relationship between the voice segments and the video images;
[0172] Determine the characteristic information of the corresponding role based on the voiceprint characteristic category and the human body characteristic category belonging to the same role.
[0173] In a possible implementation manner of the embodiment of the present disclosure, the role library module is further specifically configured to:
[0174] Determine the voiceprint characteristics at the clustering center in the voiceprint characteristic category belonging to the same role and the human body characteristics at the clustering center in the human body characteristic category as the characteristic information of the corresponding role.
[0175] In a possible implementation manner of the embodiment of the present disclosure, the role library module is further specifically configured to:
[0176] For any role, determine the first-occurring image ranked first from multiple frames of images associated with the role;
[0177] Perform text recognition on the first-occurring image to determine the introduction text information corresponding to the role;
[0178] Store the introduction text information in the role library.
[0179] In summary, the video script generation device proposed by the present disclosure obtains a target video; performs speech recognition on the speech segments in the target video to obtain corresponding line text information; extracts the human features of the speaker from the image frames associated with the speech segments in the target video, and extracts the voiceprint features of the speaker from the speech segments; determines, from the character library corresponding to the target video, target feature information that matches at least one of the human features of the speaker and the voiceprint features of the speaker, and determines the target character corresponding to the target feature information in the character library; generates a video script for the target video according to the target character and the line text information. Thus, by extracting image features from each frame image in the target video, not only the human features of the speaker are considered, but also the voiceprint features of the speaker are considered. Through the character library, the human features and voiceprint features are comprehensively used to identify the target characters appearing in the video, so as to accurately identify the characters appearing in the video. And in the process of generating the script, from the perspective of the target characters in the video, the line text information described by the target characters is focused on, so that the generated video script can truly reflect the content of the video, which helps to deepen the understanding of the video, and thus facilitates various applications based on the generated video script in the future.
[0180] To implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the video script generation method proposed in any of the above embodiments of the present disclosure.
[0181] To implement the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the video script generation method proposed in any of the above embodiments of the present disclosure.
[0182] To implement the above embodiments, the present disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the video script generation method proposed in any of the above embodiments of the present disclosure.
[0183] Figure 7FIG. 0 is a schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0184] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 702 or a computer program loaded from a storage unit 707 into a RAM (Random Access Memory) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An I / O (Input / Output) interface 705 is also connected to the bus 704.
[0185] Multiple components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0186] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the video script generation method. For example, in some embodiments, the video script generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the video script generation method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the video script generation method by any other suitable means (e.g., by means of firmware).
[0187] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, FPGA (Field Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), ASSP (Application Specific Standard Product), SOC (System On Chip), CPLD (Complex Programmable Logic Device), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0188] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0189] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0190] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0191] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0192] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0193] It should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0194] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0195] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating a video script, the method comprising: Get the target video; Performing speech recognition on the speech clips in the target video to obtain corresponding dialogue text information; Extracting body features of a speaker from image frames associated with the speech segment in the target video, and extracting voiceprint features of the speaker from the speech segment; Determine, from the character library corresponding to the target video, target feature information that matches at least one of the body feature of the speaker and the voiceprint feature of the speaker, and determine the target character corresponding to the target feature information in the character library; A video script of the target video is generated according to the target role and the dialogue text information.
2. The method according to claim 1, wherein: The step of extracting the speaker's body features from the image frames associated with the speech segment in the target video, and extracting the speaker's voiceprint features from the speech segment, includes: For any voice segment in the target video, extract the voiceprint of the voice segment to determine the voiceprint characteristics of the speaker; Determining a first image frame from a plurality of image frames associated with the voice segment; wherein a human body movement in the first image frame meets the setting requirements of the speaker; Image feature extraction is performed on the first image frame to obtain body features of the speaker.
3. The method according to claim 2, wherein: The step of determining a first image frame from a plurality of image frames associated with the voice segment comprises: Using the image frames synchronously displayed with the speech segment as multiple image frames associated with the speech segment; Performing target recognition on any image frame of the multiple image frames to obtain a human body region in the corresponding image frame, and performing action recognition on the human body region to obtain a human body action; According to the human body motion recognized in each image frame, a first image frame meeting the set requirement is determined from the multiple image frames.
4. The method according to claim 3, wherein: The extracting image features from the first image frame to obtain the body features of the speaker includes: Image features are extracted from a target human body region in the first image frame to obtain human body features of the speaker, wherein the target human body region is a human body region whose human body movements meet the set requirements.
5. The method according to claim 2, wherein: The step of performing action recognition on the human body region to obtain a human body action includes: Performing human body key point recognition on the human body region; The human body action is determined according to the positional relationship between the multiple human body key points obtained by identification.
6. The method according to claim 1, wherein: The step of determining target feature information that matches at least one of a human feature of the speaker and a voiceprint feature of the speaker from a character library corresponding to the target video, and determining a target character corresponding to the target feature information in the character library includes: Acquire feature information to be matched from the speaker's body features and the speaker's voiceprint features; In the feature information set contained in the role library, determine the similarity between each feature information in the feature information set and the feature information to be matched; wherein the feature information set includes the feature information corresponding to each role in the role library; According to the similarity, target feature information whose similarity exceeds a similarity threshold is determined from the feature information set, and a target role corresponding to the target feature information in the role library is determined.
7. The method according to claim 1, wherein: The step of generating a video script of the target video according to the target role and the line text information includes: According to the voice segment where the target character is located and the voice segment where the line text information is located, determining the target character and the line text information from the same voice segment; According to the order in which each voice segment appears in the target video, the target role and line text information of each voice segment are arranged in sequence to obtain the video script.
8. The method according to claim 7, wherein: The target video segment also includes a silent segment, and the method further includes: Performing background recognition on the silent segment and the adjacent speech segments between the silent segments; When there is a background switch between the silent segment and the adjacent speech segment, semantic recognition is performed on the background of the silent segment to obtain background description information of the silent segment; The background description information is inserted into the video script according to the order in which the silent segments appear in the target video.
9. The method according to claim 8, wherein: The method further comprises: According to the target role, query the role library to obtain the introduction text information corresponding to the target role; wherein the role library includes the introduction text information corresponding to each role; The introduction text information is supplemented into the video script according to the position of the target character's dialogue text information in the video script.
10. The method according to claim 1, wherein: The role library is generated in the following manner, including: Performing human body recognition on each frame of video image in the target video to determine a plurality of human body features corresponding to the target video; Performing voiceprint extraction on the target video to determine a plurality of voiceprint features corresponding to the target video; Clustering a plurality of voiceprint features corresponding to the target video to obtain at least one voiceprint feature category, and clustering a plurality of human body features corresponding to the target video to obtain at least one human body feature category; Determine feature information corresponding to each character based on voiceprint features belonging to the same voiceprint feature category and based on human features belonging to the same human feature category; The role library is determined according to the feature information corresponding to the role.
11. The method according to claim 10, wherein: The determining of feature information corresponding to each role according to voiceprint features belonging to the same voiceprint feature category and human features belonging to the same human feature category includes: Determine the voice segment from which the voiceprint features in each of the voiceprint feature categories come, and determine the video image from which the human body features in each of the human body feature categories come; Determining, according to the co-occurrence relationship between the voice segment and the video image, a voiceprint feature category and a human feature category belonging to the same character; Based on the voiceprint feature category and the body feature category belonging to the same role, the feature information of the corresponding role is determined.
12. The method according to claim 11, wherein: The determining of the feature information of the corresponding role based on the voiceprint feature category and the human body feature category belonging to the same role includes: The voiceprint features at the cluster center in the voiceprint feature category belonging to the same character and the human body features at the cluster center in the human body feature category are determined as feature information of the corresponding character.
13. The method according to claim 10, wherein: The method further comprises: For any character, determine the first appearing image ranked first from multiple frames of images associated with the character; Performing text recognition on the first-appearance image to determine introduction text information corresponding to the character; The introduction text information is stored in the role library.
14. A video script generation device, wherein: The device comprises: A first acquisition module, used to acquire a target video; A first determination module is used to perform speech recognition on the speech segment in the target video to obtain corresponding line text information; An extraction module, configured to extract the speaker's body features from the image frames associated with the speech segment in the target video, and to extract the speaker's voiceprint features from the speech segment; A second determination module is used to determine target feature information that matches at least one of the human features of the speaker and the voiceprint features of the speaker from a character library corresponding to the target video, and determine a target character corresponding to the target feature information in the character library; A generation module is used to generate a video script of the target video based on the target role and the dialogue text information.
15. An electronic device, characterized in that: including a processor and a memory; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method according to any one of claims 1 to 13.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 13.
Citation Information
Cited By
Video generation method and device, terminal and computer readable storage medium
CN121000947A