Video processing method, apparatus, medium, and computer program
The video processing method improves efficiency by segmenting video generation into template and variable text segments, reducing processing time and enhancing continuity, addressing the inefficiencies of existing technologies.
Patent Information
- Application Number
- JP2023554305
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-24
- Filing Date
- 2022-08-30
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2042-08-30
AI Technical Summary
Existing video processing technologies face inefficiencies when generating complete videos for complete texts, leading to low processing efficiency due to time-consuming processes.
A video processing method that involves obtaining a first video segment corresponding to a template text, generating a second video segment for variable text, and combining these segments to produce a video, thereby reducing processing time and improving efficiency.
This approach enhances video processing efficiency by reducing the length of generated videos and corresponding time costs, while also improving continuity at splicing positions through audio pauses in the first video segment.
Smart Images

Figure 0007697027000001 
Figure 0007697027000002 
Figure 0007697027000003
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and particularly to a video processing method, apparatus, medium, and program product.
[0002] This application claims the priority of a Chinese patent application filed with the Chinese Patent Office on September 24, 2021, with an application number of 202111124169.4 and an application title of "Video Processing Method, Apparatus, and Medium", the entire content of which is incorporated herein by reference.
Background Art
[0003] With the development of communication technologies, virtual objects can be widely applied in application scenarios such as announcement scenes, education scenes, medical scenes, and customer service scenes. In these application scenarios, virtual objects usually need to represent text, and correspondingly, videos corresponding to the virtual objects can be generated and played. The video can represent the process in which the virtual object represents text. The video generation process usually includes an audio generation process and an image sequence generation process. Among them, the audio generation process usually uses speech synthesis technology. The image sequence generation process usually uses image processing technology.
[0004] The inventor has found that in the process of implementing the embodiments of this application, when related technologies generate corresponding complete videos for complete texts, it usually takes a lot of time costs, resulting in relatively low video processing efficiency.
Summary of the Invention
Problems to be Solved by the Invention
[0005] How to improve the processing efficiency of video is a technical problem that those skilled in the art need to solve. In view of the above problems, the embodiments of the present application propose a video processing method, apparatus, medium, and program product that solve the above problems or at least partially solve the above problems.
Means for Solving the Problem
[0006] To solve the above problems, the present application discloses a video processing method, which is executed in an electronic device, and the method includes: Obtaining a first video segment, where the first video segment corresponds to a template text in a first text of a video to be generated, and the first video segment includes a video sub-segment where the audio pauses, and the position of the video sub-segment corresponds to a boundary position between the template text and variable text to be processed in the first text; Generating a second video segment corresponding to the variable text to be processed; Obtaining a video corresponding to the first text by combining the first video segment and the second video segment.
[0007] In another aspect, the present application discloses a video processing apparatus, A providing module used to obtain a first video segment, where the first video segment corresponds to a template text in a first text of a video to be generated, and the first video segment includes a video sub-segment where the audio pauses, and the position of the video sub-segment corresponds to a boundary position between the template text and variable text to be processed in the first text; A generating module used to generate a second video segment corresponding to the variable text to be processed; A combining module used to obtain a video corresponding to the first text by combining the first video segment and the second video segment.
[0008] In a further aspect, the present application discloses an apparatus for video processing, including a memory and one or more programs, wherein one or more programs are stored in the memory, and when the programs are executed by one or more processors, the steps of the method are realized.
[0009] In another aspect, embodiments of the present application disclose one or more machine-readable media, in which commands are stored, and when executed by one or more processors, cause the apparatus to execute the one or more methods.
[0010] In another aspect, embodiments of the present application disclose a computer program product, the program product including computer commands, the computer commands being stored in a computer-readable storage medium, and when a processor executes the computer commands, the processor executes the video processing method of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
[0012] To make the above objects, features, and advantages of the present application clearer and easier to understand, the present application will be described in more detail below by combining the drawings and specific embodiments.
[0013] In an embodiment of the present application, the virtual object is a clear, natural, and virtual object close to an actual object obtained by technologies such as object modeling and motion capture, and can be given the ability to recognize, understand, or express the virtual object by artificial intelligence technologies such as voice recognition and natural language understanding. The virtual object specifically includes a virtual person, or a virtual animal, or a two-dimensional comic object, or a three-dimensional comic object, etc.
[0014] For example, in an announcement scene, the virtual object can, for example, announce the news or explain the game on behalf of media personnel. Also, for example, in a medical scene, the virtual object can, for example, provide medical guidance on behalf of medical staff.
[0015] In a specific implementation, the virtual object can represent text. The embodiment of the present application can generate text and a video corresponding to the virtual object. The video may specifically include an audio sequence corresponding to the text and an image frame sequence corresponding to the audio sequence.
[0016] In some application scenarios, the text of the video to be generated specifically includes template text and variable text. Among them, the template text is relatively fixed, and the variable text can usually change based on preset elements such as user input.
[0017] For example, the variable text may be determined based on user input. Taking the medical scenario as an example, based on the disease name included in the user input, the corresponding variable text can be determined. Optionally, the fields corresponding to the variable text specifically include a disease name field, a food type field, a number of food ingredients field, etc., and these fields can be determined based on the disease name included in the user input.
[0018] As can be understood, those skilled in the art can determine the variable text in the text according to the actual application needs, so the embodiments of the present application do not limit the specific determination method of the variable text.
[0019] In order to ensure that the video quality meets the requirements, in the related art, in the situation where the variable text changes, usually, for the complete text after the change, the corresponding complete video is generated. However, generating the corresponding complete video for the complete text after the change usually incurs more time costs and causes a decrease in the video processing efficiency.
[0020] Regarding the technical problem of how to improve the video processing efficiency, the embodiments of the present application provide a video processing solution, which specifically includes the steps of obtaining a first video segment corresponding to the template text in the first text of the video to be generated, and the first video segment includes a video sub-segment where the audio pauses, and the position of the video sub-segment corresponds to the boundary position between the template text and the variable text to be processed in the first text, and the first text includes the template text and the variable text to be processed; generating a second video segment corresponding to the variable text to be processed; and obtaining the video corresponding to the first text by combining the first video segment and the second video segment.
[0021] The embodiments of the present application combine a first video segment corresponding to a template text and a second video segment corresponding to variable text to be processed. Among them, the first video segment may be a pre-stored video segment, and a second video segment corresponding to the variable text to be processed in the video processing process can be generated. Since the length of the variable text to be processed is smaller than the length of the complete text, the embodiments of the present application can reduce the length of the generated video and the corresponding time cost, and thus can improve the processing efficiency of the video.
[0022] Furthermore, the first video segment of the embodiments of the present application includes a video sub-segment where the audio pauses. Here, the audio pausing means that the audio has stopped, for example, the virtual object is not speaking. The position of the video sub-segment corresponds to the boundary position between the template text and the variable text to be processed in the first text. The video sub-segment where the audio pauses in the above first video segment contributes to solving the problem of hopping or jitter at the splicing position, and thus can improve the continuity at the splicing position.
[0023] The video processing method provided by the embodiments of the present application can be applied to application scenarios corresponding to client terminals and server terminals. For example, FIG. 1A shows a schematic diagram of an application scenario according to an embodiment of the present application. The client terminal and the server terminal are in a wired or wireless network, and through the wired or wireless network, the client terminal performs data interaction with the server terminal.
[0024] The client terminal and the server terminal may be collectively referred to as electronic devices. The client terminal includes, for example, but is not limited to, smartphones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop portable computers, in-vehicle computers, desktop personal computers, set-top boxes, smart TVs, and wearable devices, etc. The server terminal is a device such as, for example, a server with independent hardware, a virtual server, or a server cluster.
[0025] The client terminal corresponds to the server terminal and refers to a program that provides local services to users. The client terminal in the embodiments of the present application can receive user input and provide a video corresponding to the user input. The video may be generated by the client terminal or the server terminal, and the embodiments of the present application do not limit the specific generation entity of the video.
[0026] In one embodiment of the present application, the client terminal can receive user input and upload the user input to the server terminal, so that the server terminal can generate a video corresponding to the user input. The server terminal determines the variable text to be processed based on the user input, generates a second video segment corresponding to the variable text to be processed, and combines the pre-stored first video segment and the second video segment to obtain a video corresponding to the template text and the variable text to be processed.
[0027] Method Embodiment 1
[0028] As shown in reference to FIG. 1B, a flowchart of the video processing method of the present application is shown, and specifically may include the following steps. The video processing method may be executed by, for example, an electronic device.
[0029] Step 101: Obtain a first video segment, where the first video segment corresponds to a template text in a first text of the video to be generated, and the first video segment includes a video sub-segment where the audio pauses. The position of the video sub-segment corresponds to the boundary position between the template text and the variable text to be processed in the first text.
[0030] Step 102: Generate a second video segment corresponding to the variable text to be processed.
[0031] Step 103: Obtain a video corresponding to the first text by combining the first video segment and the second video segment.
[0032] In one embodiment, in Step 101, a first video segment corresponding to the template text can be pre-generated and saved. The first video segment includes a video sub-segment where the audio pauses. Here, "where the audio pauses" means that the audio stops or temporarily does not output. The video sub-segment where the audio pauses may be regarded as a video sub-segment without audio. The position of the video sub-segment corresponds to the boundary position between the template text and the variable text to be processed in the first text, and the video sub-segment can improve the continuity at the combination position.
[0033] The structure of the text of the embodiment of the present application specifically includes a template text and a variable text. The boundary position can be used to distinguish between adjacent template text and variable text.
[0034] Regarding "<Diabetes>" and the issue of "<Fruits>", I am still conducting research. I think the dietary advice for this "<Diabetes>" may also be helpful to you, but since it includes recommendations and taboos for about <1800> types of ingredients, please click to check. Taking text A as an example, there are multiple boundary positions in text A. For example, there is a corresponding boundary position between the template text "Regarding" and the variable text "<Diabetes>", a corresponding boundary position between the variable text "<Diabetes>" and the template text "and", a corresponding boundary position between the template text "and" and the variable text "<Fruits>", a corresponding boundary position between the variable text "<Fruits>" and the template text "of", and so on.
[0035] In one embodiment, the process of determining the first video segment may include the step of generating a preset video based on the template text, the preset variable text, and the pose information at the corresponding boundary positions, and the step of cutting out the first video segment corresponding to the template text from the preset video.
[0036] Among them, the preset variable text may be any variable text, or the preset variable text may be any instance of the variable text.
[0037] The embodiments of the present application can generate a preset video based on the template text and the preset complete text corresponding to the preset variable text. Among them, the pose information at the boundary positions may be considered in the process of generating the preset video. The pose information indicates, for example, the audio pause for a predetermined time.
[0038] In actual application, the preset video may include the preset audio corresponding to the audio part and the preset image sequence corresponding to the image part.
[0039] In a specific implementation, the TTS (Text To Speech) technology can be used to convert a preset complete text into a preset voice. The preset voice can be represented in the form of a waveform.
[0040] Converting the preset complete text of the embodiment of the present application into a preset voice specifically includes a language analysis process and an acoustic system process. Among them, the language analysis process is used to generate corresponding language information based on the preset complete text and its corresponding pause information, and the acoustic system process mainly generates the corresponding preset voice based on the language information provided by the voice analysis process and realizes the pronunciation function.
[0041] In one embodiment, the processing of the language analysis process may specifically include judgment of text structure and language type, text standardization, conversion from text to phonemes, and prosody prediction. The language information may be the result of the voice analysis process.
[0042] Among them, the judgment of text structure and language type is used to judge the language type of the preset complete text, such as the language types of Chinese, English, Tibetan, Uyghur, etc., and based on the grammar rules of the corresponding language type, the preset complete text is segmented into phrases and the segmented phrases are transmitted to the subsequent processing module.
[0043] Text standardization is used to standardize the segmented phrases based on the set rules.
[0044] The conversion from text to phonemes is used to determine the phoneme features corresponding to the phrases.
[0045] When humans express language, they usually have a tone and emotions. Therefore, the purpose of speech synthesis is generally to imitate the voice of an actual person. Thus, prosody prediction can be used to determine where pauses are needed in a sentence, how long the pauses should be, which characters or words need to be pronounced heavily, and which words need to be pronounced lightly, etc., and further to achieve pitch changes and intonations of sounds.
[0046] The embodiments of the present application can first utilize prosody prediction technology to determine a prosody prediction result, and then update the prosody prediction result based on pause information.
[0047] Taking Text A as an example, the pause information may be preset pause information added between the template text "regarding" and the variable text "<diabetes>", and updating the prosody prediction result may specifically include adding preset pause information between the phoneme features "guan", "yu" of the template text "regarding" and the phoneme features "tang", "niao", "bing" of the variable text "<diabetes>". The updated prosody prediction result may be "guan", "yu", "N milliseconds pause", "tang", "niao", "bing", etc. Among them, N may be a natural number greater than 0, and the value of N can be determined by those skilled in the art according to actual application needs.
[0048] The acoustic system process can obtain a preset voice that meets the needs according to the speech synthesis parameters.
[0049] Optionally, the speech synthesis parameters may include timbre parameters. The timbre parameters may refer to the unique characteristics in which different sound frequencies appear on the waveform surface. Usually, different sound sources correspond to different timbres. Therefore, according to the timbre parameters, a speech sequence matching the timbre of the target sound source can be obtained. The target sound source may be specified by the user. For example, the target sound source may be a specified medical staff or the like. In actual applications, according to the audio of a preset length of the target sound source, the timbre parameters of the target sound source can be obtained.
[0050] The preset image sequence corresponding to the image part can be obtained based on the virtual object image. In other words, the embodiments of the present application can obtain the preset image sequence by endowing the virtual object image with state characteristics. The virtual object image may be specified by the user. For example, the virtual object image may be an image of a famous person (such as a host).
[0051] The above state characteristics facial expression characteristics, lip characteristics, and may include at least one of limb characteristics.
[0052] Facial expressions express emotions and feelings and may refer to the emotions and feelings shown on the face.
[0053] Facial expression characteristics usually refer to those of the entire face. Lip characteristics may particularly refer to those of the lips and are related to the text content, voice, and pronunciation method of the text. Therefore, the naturalness of the expression corresponding to the preset image sequence can be improved.
[0054] Body features can convey a person's thoughts through the coordinated activities of human body parts such as the head, eyes, neck, hands, elbows, arms, body, thighs, and feet, and can vividly convey emotions and feelings. Body features may include turning around, shrugging shoulders, and gestures, etc., which can improve the richness of expressions corresponding to the image sequence. For example, at least one arm naturally hangs down when speaking, and at least one arm is naturally placed on the abdomen when not speaking, etc.
[0055] In the process of generating the preset image part of the video in the embodiments of the present application, image parameters can be determined based on the preset complete text and pose information, and the image parameters can represent the state features of the virtual object, and a preset image sequence corresponding to the image part is generated based on the image parameters.
[0056] Among them, the image parameters may include pose image parameters, and the pose image parameters can represent the pose state features corresponding to the pose information. In other words, the pose image parameters indicate the state features of aspects such as the body shape and expression that appear on the virtual object when the virtual object stops speaking. Correspondingly, the preset image sequence may include an image sequence corresponding to the pose state features. For example, the pose state features may include a neutral expression, the closed state of the lips, and the state of the arms hanging down, etc.
[0057] After generating the preset voice and the preset image sequence, the preset voice and the preset image sequence can be fused to obtain the corresponding preset video.
[0058] After obtaining the preset video, a first video segment corresponding to the template text can be cut out from the above preset video. Specifically, the cutting of the first video segment can be performed based on the start position and the end position of the preset variable text in the preset video.
[0059] Taking Text A as an example, assuming that the start position of the preset variable text "<Diabetes>" in the text corresponds to the start position T1 in the preset video, and the end position of the preset variable text "<Diabetes>" corresponds to the end position T2 in the preset video, then the video segment before T1 can be cut out from the preset video as the first video segment corresponding to the template text "Regarding". It should be noted that in the process of generating the preset video, the pause information at the boundary position is utilized. Therefore, the first video segment before T1 has the pause information (that is, the first video segment includes the video sub-segment where the audio pauses), and thus, the continuity at the joining position can be improved in the subsequent joining process.
[0060] Taking Text A as an example, assuming that the start position of the preset variable text "<Fruit>" in the text corresponds to the start position T3 in the preset video, and the start position of the preset variable text "<Fruit>" corresponds to the end position T4 in the preset video, then the video segment between T2 and T3 can be cut out from the preset video as the first video segment corresponding to the template text "and".
[0061] Since the template text in the preset complete text is divided into multiple parts by the preset variable text, in actual applications, the first video segments corresponding to multiple template texts can be extracted from the preset video respectively.
[0062] As can be understood, the acquisition method of obtaining the first video segment by utilizing the pause information at the boundary position in the above process of generating the preset video is just a selectable embodiment. In fact, those skilled in the art may also use other acquisition methods according to the actual application needs.
[0063] In one embodiment, the video sub-segment in the first video segment is such that not only is the audio in a pause state, but also the virtual object in the image of the video sub-segment is in a non-speaking state.
[0064] In one embodiment, the video sub-segment is a sub-segment obtained after undergoing pause processing.
[0065] The pause processing for the video sub-segment is obtaining an audio signal sub-segment in a paused state of the audio by performing a weighting process on the audio signal sub-segment and the mute signal at a combination position corresponding to the boundary position in the first video segment; and obtaining the image sub-sequence in which the virtual object is in a non-speaking state by performing a weighting process on the image sub-sequence and the image sequence of the target state feature at the combination position of the first video segment, wherein the target state feature indicates a feature in which the virtual object is in a non-speaking state. Thus, the audio signal sub-segment in a paused state of the audio and the image sub-sequence in which the virtual object is in a non-speaking state can constitute the video sub-segment.
[0066] In one embodiment, one way to obtain the first video segment may include generating a first video based on a template text and a preset variable text, cutting out the first video segment corresponding to the template text from the first video, and performing pause processing on the first video segment at the boundary position.
[0067] Taking the pause processing of the audio part as an example, the pause processing of the audio part can be realized by performing a weighting process on the audio signal sub-segment and the mute signal at the boundary position of the video segment. Taking the pause processing of the image part as an example, the pause processing of the image part can be realized by performing a weighting process on the image sub-sequence at the boundary position of the video segment and the image sequence of the target state feature corresponding to the pause information.
[0068] After obtaining the first video segment, by saving the first video segment, in the situation where the variable text changes, the first video segment can be combined with the second video segment corresponding to the variable text after the change (hereinafter abbreviated as the variable text to be processed).
[0069] In step 102, the variable text to be processed can be obtained based on user input. As can be understood, the embodiments of the present application do not limit the specific determination method of the variable text to be processed.
[0070] The embodiments of the present application can provide the following technical solutions for generating the second video segment corresponding to the variable text to be processed.
[0071] Technical solution 1
[0072] In Technical Solution 1, the step of generating a second video segment corresponding to the variable text to be processed specifically includes: determining corresponding audio parameters and image parameters for a clause in the first text that contains the variable text to be processed, where the image parameters represent the state characteristics of the virtual object that is to appear in the video corresponding to the first text, and the audio parameters are used to represent the parameters corresponding to voice synthesis; extracting target audio parameters and target image parameters corresponding to the variable text to be processed from among the audio parameters and image parameters; and generating a second video segment corresponding to the variable text to be processed based on the target audio parameters and target image parameters.
[0073] Technical Solution 1 first determines corresponding audio parameters and image parameters in units of the clause where the variable text to be processed is located, and then extracts target audio parameters and target image parameters corresponding to the variable text to be processed from among the audio parameters and image parameters.
[0074] A clause is a grammatically independent unit, which is composed of one word or a group of words that are syntactically connected, and expresses a statement, question, command, wish, or exclamation.
[0075] In a situation where the variable text to be processed corresponds to words, a sentence usually includes template text and also variable text to be processed. Since the speech parameters and image parameters corresponding to the sentence have a certain continuity, the target speech parameters and target image parameters corresponding to the variable text extracted therefrom, and the speech parameters and image parameters corresponding to the template text in the sentence have a certain continuity. Based on this, the continuity between the second video segment corresponding to the variable text to be processed and the first video segment corresponding to the template text in the sentence can be improved, and further the continuity at the connection position can be improved.
[0076] In actual applications, the speech parameters can represent the parameters corresponding to speech synthesis. The speech parameters may include linguistic features and / or acoustic features.
[0077] The linguistic features may include phoneme features. A phoneme is the smallest speech unit divided based on the natural attributes of speech. Analyzing according to the pronunciation actions in a syllable, one action constitutes one phoneme. Phonemes may include vowels and consonants.
[0078] The acoustic features can represent the features of speech from the perspective of pronunciation.
[0079] The acoustic features are prosodic features (suprasegmental features / supra-linguistic features), specifically including time-related features, fundamental frequency-related features, energy-related features, etc., prosodic features, and voice quality features, and A relevance analysis feature based on spectrum, which is an embodiment of the relevance between the change in vocal tract shape and the articulatory movement. Currently, the spectrum-based relevance features may include, but are not limited to, spectrum-based relevance analysis features mainly including linear prediction cepstrum coefficients (LPCC), Mel frequency cepstrum coefficients (MFCC), etc.
[0080] For better understanding, the above voice parameters are merely examples, and the embodiments of the present application do not limit specific voice parameters.
[0081] In a specific implementation, based on the target voice parameters, voice synthesis is performed on the variable text to be processed, so that the variable text to be processed can be converted into the target voice.
[0082] The image parameters may be parameters corresponding to the generation of an image sequence. The image parameters can be used to determine the state features corresponding to the virtual object, or the image parameters may include the state features corresponding to the virtual object. For example, the image parameters may include lip features.
[0083] In a specific implementation, by endowing the virtual object image with the state features corresponding to the target image parameters, a target image sequence can be obtained. The target voice and the target image sequence can be fused to obtain a second video segment.
[0084] Technical solution 2
[0085] In Technical Solution 2, the step of generating a second video segment corresponding to the variable text to be processed specifically includes performing a smoothing process on the target image parameters corresponding to the variable text to be processed based on the preset image parameters at the boundary positions of the preset variable text, so as to improve the continuity at the boundary position between the target image parameters and the image parameters of the template text, and generating a second video segment corresponding to the variable text to be processed based on the target image parameters after the smoothing process.
[0086] Technical Solution 2 performs a smoothing process on the target image parameters corresponding to the variable text to be processed based on the preset image parameters at the boundary positions of the preset variable text. Since the preset image parameters at the boundary positions of the preset variable text and the image parameters at the boundary positions of the template text have a certain continuity, the above smoothing process can improve the continuity at the boundary position between the target image parameters after the smoothing process and the image parameters of the template text. Based on this, the continuity between the second video segment corresponding to the variable text to be processed and the first video segment corresponding to the template text in the sentence can be improved, and further the continuity at the joining position can be improved.
[0087] In a specific implementation, a window function such as a Hann window can be used to perform a smoothing process on the target image parameters corresponding to the variable text to be processed based on the preset image parameters. As can be understood, the embodiments of the present application do not limit the specific process of the smoothing process.
[0088] According to the above description, in the process of generating the image part of the preset video in the embodiments of the present application, image parameters can be determined based on the preset complete text and pose information. The embodiments of the present application can extract the preset image parameters at the boundary positions of the preset variable text from the image parameters and save the preset image parameters.
[0089] Taking text A as an example, assuming that the start position of the preset variable text "<diabetes>" corresponds to the start position T1 in the preset video and the start position of the preset variable text "<diabetes>" corresponds to the end position T2 in the preset video, the image parameters between T1 and T2 can be extracted as the preset image parameters at the boundary positions of the preset variable text "<diabetes>".
[0090] Technical solution 3
[0091] In technical solution 3, the image sequence corresponding to the video includes a background image sequence and a moving image sequence. In that case, the step of generating a second video segment corresponding to the variable text to be processed specifically includes the step of generating a target moving image sequence corresponding to the variable text to be processed, the step of determining a target background image sequence corresponding to the variable text to be processed based on the preset background image sequence, and the step of obtaining a second video segment corresponding to the variable text to be processed by fusing the target moving image sequence and the target background image sequence.
[0092] In actual applications, the image sequence corresponding to the video can be decomposed into two parts. The first part is a moving image sequence, which can be used to represent the parts that move when the virtual object is represented, and usually corresponds to preset parts such as the lips, eyes, and arm parts. The second part is a background image sequence, which can be used to represent the relatively static parts when the virtual object is represented, and usually corresponds to the parts excluding the preset parts.
[0093] In a specific implementation, the background image sequence may be obtained by presetting. For example, a preset background image sequence for a preset time can be preset, and circular arrangement (which may also be called circular appearance) can be performed on the preset background image sequence in the image sequence. Based on the target image parameters corresponding to the variable text to be processed, a moving image sequence can be generated.
[0094] In actual applications, an image sequence can be obtained by fusing the moving image sequence and the background image sequence. For example, an image sequence can be obtained by pasting the moving image sequence on top of the background image sequence.
[0095] Technical solution 3 can determine the target background image sequence corresponding to the variable text to be processed based on the preset background image sequence corresponding to the variable text, improve the matching degree between the target background image sequence and the preset background image sequence, and further improve the matching degree and continuity between the target background image sequence corresponding to the variable text to be processed and the background image sequence corresponding to the template text.
[0096] According to the above description, in the process of generating the preset image part of the video in the embodiments of the present application, the information of the preset background image sequence corresponding to the preset variable text can be recorded. For example, the information of the preset background image sequence may include the start frame identifier, end frame identifier, etc. in the preset video of the preset background image sequence. For example, the information of the preset background image sequence may include the start frame number 100, end frame number 125, etc.
[0097] In one embodiment, in order to improve the matching degree at the start position or end position between the target background image sequence and the preset background image sequence, the background images at the start and end positions of the target background image sequence match the background images at the start and end positions of the preset background image sequence.
[0098] The start position may refer to the start location, and the end position may refer to the end location. Specifically, the background image at the start position of the target background image sequence matches the background image at the start position of the preset background image sequence. Or the background image at the end position of the target background image sequence matches the background image at the end position of the preset background image sequence.
[0099] Since the preset background image sequence and the background image sequence corresponding to the template text match and are continuous at the boundary position, under the situation where the target background image sequence and the preset background image sequence match at the boundary position, the matching degree and continuity at the combined position between the target background image sequence and the background image sequence corresponding to the template text can also be improved.
[0100] In order to realize that the target background image sequence matches the preset background image sequence at the boundary position, the determination method used to determine the target background image sequence corresponding to the variable text to be processed may specifically include the following determination methods. Determination method 1: Under the condition that the number N1 of images corresponding to the preset background image sequence matches the number N2 of images corresponding to the target moving image sequence, determine the preset background image sequence as the target background image sequence, or Determination method 2: When the number N1 of images corresponding to the preset background image sequence is greater than the number N2 of images corresponding to the target moving image sequence, discard the first background image at the middle position from the preset background image sequence. In the situation of discarding at least two frames of the first background image, at least two frames of the first background image are discontinuously distributed in the preset background image sequence, or Determination method 3: When the number N1 of images corresponding to the preset background image sequence is smaller than the number N2 of images corresponding to the target moving image sequence, add a second background image based on the preset background image sequence.
[0101] Regarding determination method 1, when N1 and N2 are equal, determine the preset background image sequence as the target background image sequence, and it is possible to realize the matching at the boundary position between the target background image sequence and the preset background image sequence.
[0102] In actual applications, based on the audio time information corresponding to the variable text to be processed, the number N2 of images corresponding to the target moving image sequence can be determined. The audio time information may be determined based on the audio parameters corresponding to the variable text to be processed, or the audio time information may be determined based on the time of the audio segment corresponding to the variable text to be processed.
[0103] Regarding determination method 2, under the situation where N1 is larger than N2, the first background image at the middle position can be discarded from the preset background image sequence, and matching at the boundary position between the target background image sequence and the preset background image sequence can be realized.
[0104] The middle position may be different from the starting position or the ending position. And at least two frames of the discarded first background images are discontinuously distributed in the preset background image sequence. In this way, the problem of poor continuity of the background images caused by discarding continuous background images can be avoided to a certain extent.
[0105] In actual applications, the number of the first background images can match the difference value between N1 and N2. For example, the information of the preset background image sequence may include the starting frame number 100, the ending frame number 125, etc. Assuming that the value of N1 is 26 and the number N2 of the images corresponding to the target moving image sequence is 24, two frames of the first background images at the middle position and with discontinuous positions can be discarded from the preset background image sequence.
[0106] Regarding determination method 3, under the situation where N1 is smaller than N2, a second background image is added based on the preset background image sequence, and matching at the boundary position between the target background image sequence and the preset background image sequence can be realized.
[0107] In an alternative embodiment of the present application, the second background image may be from the preset background image sequence. In other words, the second background image to be added can be determined from the preset background image sequence.
[0108] In one embodiment, first, according to the forward order, a preset background image sequence is determined as the first part of the target background image sequence. Next, according to the reverse order, the preset background image sequence is determined as the second part of the target background image sequence. Subsequently, according to the forward order, the preset background image sequence can be determined as the third part of the target background image sequence, wherein the end frame of the third part matches the end frame of the preset background image sequence.
[0109] For example, the information of the preset background image sequence may include a start frame number 100, an end frame number 125, etc. Assuming that the value of N1 is 26 and the number N2 of images corresponding to the target moving image sequence is 30, the frame numbers corresponding to the first part of the target background image sequence may be 100→125, the frame numbers corresponding to the second part of the target background image sequence may be 125→124, and the frame numbers corresponding to the third part of the target background image sequence may be 124→125.
[0110] In another alternative embodiment of the present application, the second background image may be from a background image sequence other than the preset background image sequence. For example, the second background image can be determined from the background image sequence after the preset background image sequence.
[0111] In one embodiment, first, according to the forward order, a preset background image sequence is determined as the first part of the target background image sequence. Next, according to the forward order, the subsequent background image sequence of the preset background image sequence is determined as the second part of the target background image sequence. Subsequently, according to the reverse order, the subsequent background image sequence of the preset background image sequence and the end frame of the preset background image sequence can be determined as the third part of the target background image sequence, wherein the end frame of the third part matches the end frame of the preset background image sequence.
[0112] For example, the information of the preset background image sequence may include a start frame number 100, an end frame number 125, etc. Assuming that the value of N1 is 26 and the number N2 of images corresponding to the target video sequence is 30, the frame numbers corresponding to the first part of the target background image sequence may be 100→125, the frame numbers corresponding to the second part of the target background image sequence may be 126→127, and the frame numbers corresponding to the third part of the target background image sequence may be 127→125.
[0113] As can be understood, the above implementation of adding the second background image based on the preset background image sequence is merely an example. In fact, those skilled in the art can use other implementations according to actual application needs. Any implementation that can achieve matching at the boundary position between the target background image sequence and the preset background image sequence is included within the protection scope of the implementation of the embodiments of the present application.
[0114] For example, in other embodiments, a reverse target background image sequence may be further determined. The corresponding determination process may first include determining a preset background image sequence as the first part of the target background image sequence according to the reverse order, and then determining the preset background image sequence as the second part of the target background image sequence according to the forward order, and subsequently determining the preset background image sequence as the third part of the target background image sequence according to the reverse order, where the start frame of the third part matches the start frame of the preset background image sequence.
[0115] For example, the information of the preset background image sequence may include a start frame number 100, an end frame number 125, etc. Assuming that the value of N1 is 26 and the number N2 of images corresponding to the target video sequence is 30, the frame numbers corresponding to the first part of the target background image sequence may be 125→100, the frame numbers corresponding to the second part of the target background image sequence may be 100→101, and the frame numbers corresponding to the third part of the target background image sequence may be 101→100. The frame numbers of the target background image sequence obtained in such a situation may be 100→101→101→100→100→125.
[0116] As described above, the process of generating the second video segment corresponding to the variable text to be processed has been described in detail by means 1 to 3 for technical solutions. As can be understood, those skilled in the art can use any one or a combination of means 1 to 3 for technical solutions according to actual application needs, but the embodiments of the present application do not limit the specific process of generating the second video segment corresponding to the variable text to be processed.
[0117] In step 103, by combining the first video segment and the second video segment, a video corresponding to the first text can be obtained.
[0118] In an alternative embodiment of the present application, the first video segment may specifically include a first audio segment, and the second video segment may specifically include a second audio segment.
[0119] In that case, the step of combining the first video segment and the second video segment may specifically include a step of performing a smoothing process on the audio sub-segments at the respective combination positions of the first audio segment and the second audio segment, and a step of combining the smoothed first audio segment and the smoothed second audio segment.
[0120] In the embodiment of the present application, first, a smoothing process is performed on the audio sub-segments at the respective combination positions of the first audio segment and the second audio segment, and then the smoothed first audio segment and the smoothed second audio segment are combined. The smoothing process can improve the continuity between the smoothed first audio segment and the second audio segment, and thus can improve the continuity at the combination position between the first video segment and the second video segment.
[0121] In actual applications, the combined video can be output, for example, to a user. Taking a medical scene as an example, based on the disease name included in the user input, the corresponding variable text to be processed is determined, and an embodiment of the method shown in FIG. 1B is utilized to obtain a video and provide the video to the user.
[0122] As described above, the video processing method of the embodiment of the present application combines a first video segment corresponding to a template text and a second video segment corresponding to a variable text to be processed. Among them, the first video segment may be a pre-stored video segment, and the second video segment corresponding to the variable text to be processed in the video processing process can be generated. Since the length of the variable text to be processed is smaller than the length of the complete text, the embodiment of the present application can reduce the length of the generated video and the corresponding time cost, and thus can improve the video processing efficiency.
[0123] Furthermore, for the first video segment of the embodiment of the present application, a video sub-segment that has undergone a pause process is set at the boundary position between the template text and the variable text. The above pause process can eliminate the problem of hopping or shaking to a certain extent at the connection position, and thus can improve the continuity at the connection position.
[0124] Embodiment 2 of the method
[0125] As shown in FIG. 2, a flowchart of the video processing method of the embodiment of the present application is shown, and specifically may include the following steps.
[0126] Step 201: Generate a preset video based on the template text, the preset variable text, and the corresponding pause information at the boundary position, where the pause information indicates an audio pause for a predetermined time.
[0127] Step 202: Cut out a first video segment corresponding to the template text from the above preset video and save the first video segment.
[0128] Step 203: Save the preset image parameters at the boundary position of the preset variable text and the information of the preset background image sequence corresponding to the preset variable text based on the information of the preset video.
[0129] Steps 201 to 203 can be used to pre-save information on a first video segment, pre-set image parameters at the boundary positions of pre-set variable text, and a pre-set background image sequence corresponding to the pre-set variable text, based on the generated pre-set video.
[0130] Steps 204 to 211 can be used to generate a second video segment corresponding to the variable text to be processed, and combine the pre-saved first video segment and the second video segment, based on the pre-saved information.
[0131] Step 204: Determine corresponding audio parameters and image parameters for the clause where the variable text to be processed is located.
[0132] Step 205: Extract target audio parameters and target image parameters corresponding to the variable text to be processed from among the above audio parameters and image parameters.
[0133] Step 206: Perform a smoothing process on the target image parameters corresponding to the variable text to be processed, based on the pre-set image parameters.
[0134] Step 207: Generate a target moving image sequence corresponding to the variable text to be processed, based on the target audio parameters and the target image parameters after the smoothing process.
[0135] Step 208: Determine a target background image sequence corresponding to the variable text to be processed, based on the pre-set background image sequence.
[0136] Step 209: Obtain a second video segment corresponding to the variable text to be processed by fusing the target moving image sequence and the target background image sequence.
[0137] Step 210: Perform a smoothing process on the audio sub - segments at the respective boundary positions of the first audio segment in the first video segment and the second audio segment in the second video segment.
[0138] Step 211: Combine the first video segment and the second video segment based on the first audio segment after the smoothing process and the second audio segment after the smoothing process.
[0139] In the application example of the present application, assuming that the preset complete text is the above - mentioned Text A, and the preset variable texts are "<diabetes>", "<fruit>", and "<1800>" in Text A, etc., a preset video can be generated based on Text A and the corresponding pose information, and the first video segment in the preset video, the preset image parameters at the boundary positions of the preset variable texts, and the information of the preset background image sequence corresponding to the preset variable texts can be saved.
[0140] In actual applications, elements such as user input may cause changes in the variable text. For example, in the situation where Text A becomes Text B: "<Regarding coronary heart disease and <vegetables>, I am still researching. I think this <diet advice for coronary heart disease> may also be useful to you, but since it includes recommendations and taboos for about <900> types of ingredients, please click to check>", the variable text to be processed may include "<coronary heart disease>", "<vegetables>", and "<900>" in Text B, etc.
[0141] The embodiments of the present application can generate a second video segment corresponding to the variable text to be processed. For example, first, determine the acoustic parameters and lip features of the phrase where the variable text to be processed is located, and then extract the target acoustic parameters and target lip features corresponding to the variable text to be processed from them, and can generate an audio segment corresponding to the variable text to be processed and a target image sequence respectively. The target image sequence may include a target moving image sequence and a target background image sequence.
[0142] In the process of generating the target moving image sequence, by using step 206 to perform a smoothing process on the features of the target lips, the continuity at the joint position of the lip features can be improved.
[0143] By using step 208 to generate the target background image sequence and realizing the matching at the boundary position between the target background image sequence and the preset background image sequence, the continuity at the joint position of the background image sequence can be improved.
[0144] Before combining the first video segment and the second video segment, first, perform a smoothing process on the audio sub-segments at the above boundary positions of the first audio segment in the first video segment and the second audio segment in the second video segment respectively, and then, based on the smoothed first audio segment and the smoothed second audio segment, the first video segment and the second video segment can be combined.
[0145] As described above, the video processing method of the embodiments of the present application contributes to adding a pause for a preset time at the joint position of the first video segment to eliminate the problem of hopping or shaking at the joint position, and thus the continuity at the joint position can be improved.
[0146] Moreover, in the embodiments of the present application, corresponding voice parameters and image parameters are determined in units of phrases where the variable text to be processed is located. Next, target voice parameters and target image parameters corresponding to the variable text to be processed are extracted from among the voice parameters and image parameters. Since the voice parameters and image parameters corresponding to a phrase have a certain continuity, the target voice parameters and target image parameters corresponding to the variable text to be processed extracted therefrom, and the voice parameters and image parameters corresponding to the template text in the phrase have a certain continuity. Based on this, the continuity between the second video segment corresponding to the variable text to be processed and the first video segment corresponding to the template text in the phrase can be improved, and furthermore, the continuity at the joining position can be further improved.
[0147] In addition, in the embodiments of the present application, smoothing processing is performed on the target image parameters corresponding to the variable text to be processed based on the preset image parameters at the boundary position of the preset variable text. Since the preset image parameters at the boundary position of the preset variable text and the image parameters at the boundary position of the template text have a certain continuity, the above smoothing processing can improve the continuity at the boundary position between the target image parameters after the smoothing processing and the image parameters of the template text. Based on this, the continuity between the second video segment corresponding to the variable text to be processed and the first video segment corresponding to the template text in the phrase can be improved, and furthermore, the continuity at the joining position can be improved.
[0148] Note that in the embodiments of the present application, a target background image sequence is generated based on a preset background image sequence, and by realizing the matching at the boundary position between the target background image sequence and the preset background image sequence, the continuity at the joining position of the background image sequence can be improved.
[0149] Furthermore, before combining the first video segment and the second video segment, the embodiment of the present application performs a smoothing process on the audio sub-segments at the boundary positions of the first audio segment in the first video segment and the second audio segment in the second video segment. The smoothing process can improve the continuity between the first audio segment and the second audio segment after the smoothing process, and thus can improve the continuity at the combination position of the first video segment and the second video segment.
[0150] It should be noted that for the embodiments of the method, for the sake of simplicity of description, they are described as a combination of a series of motion operations. However, those skilled in the art should know that according to the embodiments of the present application, a certain step may use other orders or may be performed simultaneously, so the embodiments of the present application are not limited to the described order of motion operations. Next, those skilled in the art should also know that the embodiments described in the specification all belong to preferred embodiments, and the related motion operations are not necessarily required for the embodiments of the present application.
[0151] Embodiment of the device
[0152] As shown in FIG. 3, a structural block diagram of an embodiment of the video processing device of the present application is shown. Specifically, A providing module 301 used to obtain a first video segment, where the first video segment corresponds to a template text in a first text of a video to be generated, and the first video segment includes a video sub-segment where the audio pauses, and the position of the video sub-segment corresponds to the boundary position between the template text and a variable text to be processed in the first text, the providing module 301; A generating module 302 used to generate a second video segment corresponding to the variable text to be processed; A combining module 303 used to obtain a video corresponding to the first text by combining the first video segment and the second video segment may be included.
[0153] Optionally, the apparatus A preset video generation module used to generate a preset video based on a template text, a preset variable text, and corresponding pose information at the boundary position, where the pose information indicates an audio pose for a predetermined time. The apparatus may further include a cutting module used to cut out a first video segment corresponding to the template text from the preset video.
[0154] Optionally, the generation module 302 A parameter determination module used to determine corresponding audio parameters and image parameters for a phrase with variable text to be processed in the first text, where the image parameters represent state characteristics of a virtual object to appear in the video corresponding to the first text, and the audio parameters represent parameters corresponding to speech synthesis. A parameter extraction module used to extract target audio parameters and target image parameters corresponding to the variable text to be processed from the audio parameters and image parameters. A first segment generation module used to generate a second video segment corresponding to the variable text to be processed based on the target audio parameters and target image parameters may be included.
[0155] Optionally, the generation module 302 A first smoothing processing module used to improve the continuity at the boundary position between the target image parameters and the image parameters of the template text by performing a smoothing process on the target image parameters corresponding to the variable text to be processed based on preset image parameters at the boundary position of the variable text to be processed; It may include a second segment generation module used to generate a second video segment corresponding to the variable text to be processed based on the target image parameters of the variable text to be processed after the smoothing process.
[0156] Optionally, the first video segment may include a first audio segment, and the second video segment may include a second audio segment. The combining module 303 A second smoothing processing module used to perform a smoothing process on the audio sub-segments at the respective combining positions of the first audio segment and the second audio segment; It may include a post-smoothing combining module used to combine the smoothed first audio segment and the smoothed second audio segment.
[0157] Optionally, the image sequence corresponding to the video may include a background image sequence and a moving image sequence. The generation module 302 A moving image sequence generation module used to generate a target moving image sequence corresponding to the variable text to be processed; A background image sequence generation module used to determine a target background image sequence corresponding to the variable text to be processed based on a preset background image sequence; It may include a fusion module used to obtain a second video segment corresponding to the variable text to be processed by fusing the target moving image sequence and the target background image sequence.
[0158] Optionally, the background images at the start and end positions of the target background image sequence match the background images at the start and end positions of the preset background image sequence.
[0159] Optionally, the background image sequence generation module is a first background image sequence generation module used to determine the preset background image sequence as the target background image sequence in a situation where the number of images corresponding to the preset background image sequence matches the number of images corresponding to the target moving image sequence, or is a second background image sequence generation module used to discard the first background image at the middle position from the preset background image sequence in a situation where the number of images corresponding to the preset background image sequence is greater than the number of images corresponding to the target moving image sequence. In a situation where at least two frames of the first background image are discarded, the at least two frames of the first background image are discontinuously distributed in the preset background image sequence, or may include a third background image sequence generation module used to add a second background image to the preset background image sequence in a situation where the number of images corresponding to the preset background image sequence is less than the number of images corresponding to the target moving image sequence.
[0160] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively brief, and the relevant parts may refer to the description of the method embodiments.
[0161] Each embodiment in this specification is described in an advanced way, and the focus of the description of each individual embodiment is different from that of other embodiments. The similar or analogous parts between each embodiment may be referred to each other.
[0162] Regarding the apparatus in the above embodiment, since the specific manner in which each module executes operations is described in detail in the embodiments related to the method, detailed discussion and explanation are omitted here.
[0163] FIG. 4 is a structural block diagram of an apparatus 900 used for video processing shown based on one exemplary embodiment. For example, the apparatus 900 may be a mobile phone, a computer, a digital broadcast terminal, a message transceiver, a game control panel, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0164] As shown in reference to FIG. 4, the apparatus 900 may include one or more units of a processing unit 902, a memory 904, a power unit 906, a multimedia unit 908, an audio unit 910, an input / output (I / O) interface 912, a sensor unit 914, and a communication unit 916.
[0165] The processing unit 902 generally controls the overall operations of the apparatus 900, such as operations leading to display, incoming call origination, data communication, camera operation, and recording operation. The processing element 902 may include one or more processors 920 that complete all or some of the steps of the above method by executing commands. Note that the processing unit 902 may include one or more modules that facilitate the interaction between the processing unit 902 and other units. For example, the processing unit 902 may include a multimedia module that facilitates the interaction between the multimedia unit 908 and the processing unit 902.
[0166] Memory 904 is configured to support operations in device 900 by storing various types of data. Examples of these data include any application programs, or method commands, contact data, phone book data, messages, pictures, and videos, etc. used for operating in device 900. Memory 904 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, for example, static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0167] Power unit 906 provides power to various units of device 900. Power unit 906 may include a power management system, one, or multiple power supplies, and other units related to generating, managing, and distributing power to device 900.
[0168] The multimedia unit 908 includes a screen that provides one output interface between the device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen, thereby receiving input signals from the user. The touch panel includes one or more touch sensors that detect touches, swipes, and gestures on the touch panel. The touch sensors can detect not only the boundaries of touch or slide motion operations, but also the duration and pressure associated with the touch or slide operations. In some embodiments, the multimedia unit 908 includes one front camera and / or a rear camera. When the device 900 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive peripheral multimedia data. Each individual front camera and rear camera may be a fixed optical lens system, or may have a focal length and an optical zoom capability.
[0169] The audio unit 910 is configured to output and / or input audio signals. For example, the audio unit 910 includes one microphone (MIC), and when the device 900 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive peripheral audio signals. The received audio signals may be further stored in the memory 904 or transmitted via the communication unit 916. In some embodiments, the audio unit 910 further includes one speaker used to output audio signals.
[0170] The I / O interface 912 provides an interface between the processing unit 902 and an external interface module, and the external interface module may be, for example, a keyboard, a click wheel, and buttons. These buttons may include, but are not limited to, a home page button, a volume button, a start button, and a lock button.
[0171] The sensor unit 914 includes one or more sensors used to provide various state evaluations for the device 900. For example, the sensor unit 914 can detect the on / off state of the device 900 and the relative positioning of the unit. For example, the unit is the display and the keypad of the device 900, and the sensor unit 914 can further detect a change in the position of the device 900 or a unit of the device 900, the presence or absence of contact between the user and the device 900, the orientation of the device 900, or acceleration / deceleration, and a change in the temperature of the device 900. The sensor unit 914 may include a proximity sensor configured to detect the presence of nearby objects when there is no physical contact. The sensor unit 914 may further include an optical sensor used in an imaging application, such as a CMOS or a CCD image sensor. In some embodiments, the sensor unit 914 may further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0172] The communication unit 916 is configured to facilitate wired or wireless communication between the device 900 and other devices. The device 900 can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, or a combination thereof. In one exemplary embodiment, the communication member 916 receives a broadcast signal or broadcast-related information from a peripheral broadcast management system via a broadcast channel. In one exemplary embodiment, the communication member 916 further includes a Near Field Communication (NFC) module for facilitating short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0173] In an exemplary embodiment, the device 900 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic elements, and is used to execute the above method.
[0174] In an exemplary embodiment, a non-transitory computer-readable storage medium containing commands, such as the memory 904 containing commands, is further provided. The above commands can complete the above method when executed by the processor 920 of the device 900. For example, the non-transitory computer-readable storage medium may be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0175] FIG. 5 is a structural block diagram of a server terminal in some embodiments of the present application. The server terminal 1900 can produce greater differences due to different configurations or performances, and may include one or more central processing units (CPUs) 1922 (for example, one or more processors), a memory 1932, and one or more storage media 1930 (for example, one or more large-capacity storage devices) for storing application programs 1942 or data 1944. Among them, the memory 1932 and the storage media 1930 may store temporarily or permanently. The program stored in the storage media 1930 may include one or more modules (not shown), and each module may include a series of command operations for the server terminal. Further, the central processor 1922 may be set to communicate with the storage media 1930 and execute a series of command operations in the storage media 1930 in the server terminal 1900.
[0176] The server terminal 1900 may include one or more power supplies 1926, one or more wired or wireless network interfaces 1950, one or more input / output interfaces 1958, one or more keyboards 1956, and / or one or more operating systems 1941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0177] A non-transitory computer-readable storage medium, when the commands in the above storage medium are executed by a processor of a device (apparatus, or server terminal), can cause the device to execute a video processing method according to an embodiment of the present application.
[0178] After considering the specification and practicing the invention disclosed herein, those skilled in the art can easily conceive of other embodiments of this application. This application aims to cover any variations, uses, or adaptive changes of this application, and these variations, uses, or adaptive changes comply with the general principles of this application and include common general knowledge or conventional means in the technical field not disclosed in this disclosure. The specification and examples are regarded as merely illustrative, and the actual scope and spirit of this application are defined by the following claims.
[0179] It should be understood that this application is not limited to the exact structure already described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
[0180] The above are only preferred embodiments of this application and are not intended to limit this application. Any modifications, substitutions for equivalents, improvements, etc. made within the spirit and principles of this application should all be included within the protection scope of this application.
[0181] The above has described in detail the video processing method, video processing apparatus, and apparatus used for video processing provided by the embodiments of this application. Specific examples are applied in this specification to discuss the principles and embodiments of this application. The description of the above embodiments is only used to help understand the method and its central idea of this application. Also, for those skilled in the art, according to the idea of this application, any changes can be made in the specific embodiments and application scopes. Therefore, it should be understood that the content of this specification does not limit this application.
Explanation of Reference Numerals
[0182] 301 Providing Module 302 Generating Module 303 Combining Module 900 Apparatus 900 Equipment 904 Memory 906 Power Supply Unit 908 Multimedia Unit 910 Audio Unit 912 I / O Interface 912 Interface 914 Sensor Unit 916 Communication Member 916 Communication Unit 920 Processor 1900 Server Terminal 1922 Central Processor 1926 Power Supply 1930 Memory Medium 1932 Memory 1941 Operating System 1942 Application Program 1944 Data 1950 Wireless Network Interface 1956 Keyboard 1958 Input / Output Interface
Claims
1. A video processing method, which is executed by an electronic device, and the method includes: generating a preset video corresponding to a template text and a preset variable text, a first video segment corresponding to the template text of the preset video includes a video sub-segment where the audio pauses, and a position of the video sub-segment corresponds to a boundary position between the template text and the preset variable text; obtaining the first video segment from the preset video, where the first video segment corresponds to the template text in the first text of the video to be generated, and the variable text to be processed in the first text replaces the preset variable text of the preset video; generating a second video segment corresponding to the variable text to be processed, for a phrase of the variable text to be processed in the first text, determining corresponding audio parameters and image parameters, where the image parameters represent state characteristics of a virtual object that will appear in the video corresponding to the first text, and the audio parameters are used to represent parameters corresponding to audio synthesis; extracting target audio parameters and target image parameters corresponding to the variable text to be processed from the audio parameters and the image parameters; generating a second video segment corresponding to the variable text to be processed based on the target audio parameters and the target image parameters; obtaining a video corresponding to the first text by combining the first video segment and the second video segment.
2. The step of generating a second video segment corresponding to the variable text to be processed includes: A step of generating a preset video based on corresponding pose information at the boundary position, wherein the pose information indicates a voice pose for a predetermined time, and Further comprising: cutting out a first video segment corresponding to the template text from the preset video. The method according to claim 1.
3. In the image of the video sub-segment, the virtual object is in a state of not speaking. The method according to claim 1.
4. The video sub-segment is a sub-segment obtained after pose processing, The pose processing for the video sub-segment is Obtaining an audio signal sub-segment in which the audio becomes a pose by performing a weighting process on the audio signal sub-segment and the mute signal at the combination position corresponding to the boundary position in the first video segment, and Obtaining the image sub-sequence in which the virtual object is in a state of not speaking by performing a weighting process on the image sub-sequence at the combination position in the first video segment and the image sequence of the target state feature, wherein the target state feature indicates a feature in which the virtual object is in a state of not speaking. The method according to claim 1, characterized by including.
5. The step of generating a second video segment corresponding to the variable text to be processed is Improving the continuity at the boundary position between the target image parameter and the image parameter of the template text by performing a smoothing process on the target image parameter corresponding to the variable text to be processed based on a preset image parameter at the boundary position of the variable text to be processed, and Generating a second video segment corresponding to the variable text to be processed based on the target image parameters after the smoothing process, and the method according to claim 1, comprising:
6. The first video segment includes a first audio segment, and the second video segment includes a second audio segment. The step of combining the first video segment and the second video segment includes: Performing a smoothing process on the audio sub-segments at the respective combination positions of the first audio segment and the second audio segment; Combining the first audio segment after the smoothing process and the second audio segment after the smoothing process, and the method according to claim 1, comprising:
7. The image sequence corresponding to the preset video includes a background image sequence and a moving image sequence. The step of generating a second video segment corresponding to the variable text to be processed includes: Generating a target moving image sequence corresponding to the variable text to be processed; Determining a target background image sequence corresponding to the variable text to be processed based on the preset background image sequence; Obtaining a second video segment corresponding to the variable text to be processed by fusing the target moving image sequence and the target background image sequence, and the method according to claim 1, comprising:
8. The background images at the start and end positions of the target background image sequence match the background images at the start and end positions of the preset background image sequence, and the method according to claim 7.
9. The step of determining a target background image sequence corresponding to the variable text to be processed based on the preset background image sequence includes: Determining the preset background image sequence as the target background image sequence in a situation where the number of images corresponding to the preset background image sequence matches the number of images corresponding to the target moving image sequence, or In a situation where the number of images corresponding to the preset background image sequence is greater than the number of images corresponding to the target moving image sequence, a step of discarding a first background image at an intermediate position from among the preset background image sequences, wherein in a situation of discarding at least two frames of the first background image, the at least two frames of the first background image are discontinuously distributed in the preset background image sequence, or The method according to claim 7, comprising adding a second background image to the preset background image sequence in a situation where the number of images corresponding to the preset background image sequence is less than the number of images corresponding to the target moving image sequence.
10. A video processing apparatus, comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and, when executed by one or more processors, implement the method according to any one of claims 1 to 9, and is used for video processing.
11. A computer program configured to cause a processor to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Virtual anchor implementation method and device
CN109637518A
Video generation processing method and device, terminal equipment and storage medium
CN110324709A
Video generation method and device and terminal
CN110381266A
Interactive object driving method, device and equipment and storage medium
CN111460785A
Voice and moving image synthesizing device and voice and moving image data base
JP1999231899A