Video generation method and system based on voice analysis, and storage medium

By performing multimodal feature extraction and contextual correlation model recognition for speech, combined with regional features and sentiment analysis, videos that are more in line with voice content and regional style are generated, which solves the problem of inaccurate video generation in the existing technology and achieves higher emotional expression and user experience.

CN120388579AInactive Publication Date: 2025-07-29SHENZHEN SMART INSURANCE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510464423.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing voice-driven video generation technology lacks in-depth analysis of the implicit situational characteristics, regional accents and multi-dimensional emotions in the voice, resulting in disconnection between the video background, character clothing and voice content or inconsistent emotional expression.

Method used

By performing multimodal feature extraction of the input speech, a pre-trained scenario correlation model and dialect classifier are used to identify regional categories, and combining emotion intensity curves and timing alignment algorithms to generate videos that conform to the speech content and regional style.

Benefits of technology

It improves the accuracy of video generation, makes the video content more consistent with the voice content, and the emotional performance is more vivid, and can optimize the generation process based on user feedback and improve user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388579A_ABST
    Figure CN120388579A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and multimedia, and particularly discloses a video generation method based on voice analysis, and the method comprises the following steps: analyzing an input voice, and extracting a multi-modal voice feature; inputting the multi-modal speech features into a pre-trained scene association model, and outputting a scene label set; on the basis of dialect features in the input voice, region categories are recognized through a dialect classifier, and a corresponding visual element library is loaded from a culture database according to the region categories; selecting a scene template according to scene type labels in the scene label set, and selecting a character action template in combination with emotion type labels and interaction object relation labels; and calculating time sequence distribution of the video elements based on the speech speed change parameters, and rendering the scene template, the character action template and the rhythm of the input speech according to the time sequence distribution through a time sequence alignment algorithm to generate a target video. The method can improve the accuracy of the video generated based on the voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and multimedia technology, and particularly relates to a method, system and storage medium for video generation based on speech analysis. Background Art

[0002] Existing speech-driven video generation technologies are mostly based on text keywords, lacking in-depth analysis of implicit situational features, regional accents, and multi-dimensional emotions in speech. For example, when generating dialect dialogue scenes, traditional methods often result in a disconnection between the video background, character costumes, and speech content due to the neglect of regional cultural elements; or fail to accurately identify emotional fluctuations in speech, making the generated video actions inconsistent with the emotional expression. In addition, existing technologies do not make full use of multi-modal features (such as intonation, speech rate, pauses) in speech, resulting in a mismatch between dynamic elements (such as scene transitions, character expressions) in the generated video and the speech rhythm.

[0003] Therefore, it is necessary to provide a method, system and storage medium for video generation based on speech analysis to improve the accuracy of video generation based on speech analysis, making the generated video more in line with the content, situation, and regional style of the analyzed speech. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, system and storage medium for video generation based on speech analysis, and solve the following technical problems: When generating dialect dialogue scenes, traditional methods often result in a disconnection between the video background, character costumes, and speech content due to the neglect of regional cultural elements; or fail to accurately identify emotional fluctuations in speech, making the generated video actions inconsistent with the emotional expression.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A method for video generation based on speech analysis, the method includes the following steps:

[0007] Parse the input speech to extract multi-modal speech features; the multi-modal speech features include phoneme sequences, emotional intensity curves, speech rate change parameters, and intonation contours;

[0008] Input the multi-modal speech features into a pre-trained situation association model to output a set of situation labels; the set of situation labels includes scene type labels, emotional category labels, and interaction object relationship labels;

[0009] Based on the dialect features in the input speech, identify the regional category through a dialect classifier, and load the corresponding visual element library from the cultural database according to the regional category. The visual element library includes scene templates, character costumes, and regional iconic props;

[0010] Select a scene template according to the scene type tags in the set of scenario tags, and select a character action template in combination with the sentiment category tags and the interaction object relationship tags;

[0011] Calculate the temporal distribution of video elements based on the speech rate change parameter, and render the scene template, the character action template and the rhythm of the input speech according to the temporal distribution through a temporal alignment algorithm to generate a target video.

[0012] As a further solution of the present invention, the pre-trained scenario association model is obtained in the following manner:

[0013] Obtain speech-video alignment samples, each sample including a speech segment and its corresponding video segment, and label the set of scenario tags and the regional tags of the video segment;

[0014] Extract sample speech features from the speech segment, the sample speech features including a sample phoneme sequence, a sample sentiment intensity curve, a sample speech rate change parameter and a sample intonation contour, extract sample video features from the video segment, the sample video features including scene features, character action features and background element features;

[0015] Input the sample speech features and the sample video features into an encoder, and calculate the similarity score of the speech-video pair;

[0016] Construct a contrast loss function based on the similarity score, and optimize the parameters of the encoder through the backpropagation algorithm so that the similarity of the matching speech-video pairs is maximized and the similarity of the non-matching pairs is minimized;

[0017] Use the trained encoder as the scenario association model for real-time scenario tag prediction of the input speech.

[0018] As a further solution of the present invention, the contrast loss function is as shown in the following formula:

[0019]

[0020] where V i represents the sample video feature of the i-th video segment, a i represents the sample speech feature of the i-th speech segment, s(v i ,a i ) represents the cosine similarity between the sample video feature V i and the sample speech feature a i , τ is a temperature coefficient, and its value range is 0.1-1.0, which is used to adjust the smoothness of the similarity distribution, N is the total number of samples in the training batch, j is a sample index variable, and all samples in the batch are traversed (j = 1, 2,..., N).

[0021] As a further solution of the present invention, based on the dialect features in the input speech, identifying the geographical category through a dialect classifier includes:

[0022] Analyzing the dialect features and phoneme distribution of the input speech through the dialect classifier, and outputting a geographical probability distribution;

[0023] Determining multiple candidate geographical categories according to the geographical probability distribution;

[0024] Calculating the fusion weights of multi-geographical visual elements based on the geographical probability distribution and the multiple candidate geographical categories;

[0025] Determining the geographical category based on the fusion weights.

[0026] As a further solution of the present invention, calculating the fusion weights of multi-geographical visual elements based on the geographical probability distribution and the multiple candidate geographical categories includes:

[0027] Calculating the fusion weights of the multi-geographical visual elements based on the geographical probability distribution and the multiple candidate geographical categories through the following formula:

[0028]

[0029] where K is the total number of the multiple candidate geographical categories, P k represents the probability of the k-th geographical category in the geographical probability distribution, γ is a sharpening coefficient, γ is greater than or equal to 1, used to enhance the weight ratio of high-probability regions, m is a geographical category index variable, traversing all K categories, m = 1, 2,..., K.

[0030] As a further solution of the present invention, calculating the temporal distribution of video elements based on the speech rate change parameter includes:

[0031] Segmenting the input speech into multiple speech paragraphs according to the speech rate change parameter;

[0032] Assigning an initial key frame timestamp to each speech paragraph to determine the temporal distribution of the video elements.

[0033] As a further solution of the present invention, rendering the scene template, the character action template and the rhythm of the input speech according to the temporal distribution through a temporal alignment algorithm to generate a target video includes:

[0034] Inserting dynamic special effects within the speech paragraphs where the emotional intensity exceeds a preset threshold based on the emotional intensity curve, and the type of the dynamic special effects is determined by the emotional category label;

[0035] Adjust the initial key frame timestamps of each speech segment in the input speech through a timing alignment algorithm, so that the mouth movement of the character in the character action template is aligned with the time of the phoneme sequence and the scene template.

[0036] As a further solution of the present invention, the method further includes:

[0037] Obtain the user's rating of the target video;

[0038] If the rating is lower than the rating threshold, increase the weight of the emotional intensity curve in the training of the scenario association model.

[0039] A video generation system based on speech analysis, the system includes:

[0040] A speech feature extraction module, configured to parse the input speech and extract multi-modal speech features; the multi-modal speech features include a phoneme sequence, an emotional intensity curve, a speech rate change parameter, and an intonation contour;

[0041] A scenario label acquisition module, configured to input the multi-modal speech features into a pre-trained scenario association model and output a set of scenario labels; the set of scenario labels includes a scene type label, an emotional category label, and an interaction object relationship label;

[0042] A regional category recognition module, configured to identify a regional category based on the dialect features in the input speech through a dialect classifier, and load a corresponding visual element library from a cultural database according to the regional category, where the visual element library includes a scene template, character costumes, and regional iconic props;

[0043] A scene selection module, configured to select a scene template according to the scene type label in the set of scenario labels, and select a character action template in combination with the emotional category label and the interaction object relationship label;

[0044] ]>A target video generation module, configured to calculate the timing distribution of video elements based on the speech rate change parameter, and render the scene template, the character action template, and the rhythm of the input speech according to the timing distribution through a timing alignment algorithm to generate a target video.

[0045] A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the video generation method based on speech analysis as described in any one of the above are implemented.

[0046] Advantages of the present invention: (1) By generating video content based on speech analysis and combining the emotional intensity curve and the scenario association model, the emotional expression of the video is made more accurate and vivid. In particular, generating an emotional intensity curve based on speech analysis can more precisely express emotional changes, thus making the video content more consistent with the emotional responses of the audience and enhancing the appeal of the video; (2) By obtaining the ratings of users for the target video and adjusting the weights of the emotional intensity curve in the scenario association model according to the rating results, the present invention can adaptively optimize the video generation process based on user feedback. If the user rating is lower than the preset threshold, the system will increase the weights of the emotional elements to enhance the emotional expression of the video. This feedback mechanism can ensure the continuous optimization of the video content, thereby improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The present invention will be further described below with reference to the accompanying drawings.

[0048] Figure 1 is a schematic flowchart of the video generation method based on speech analysis of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Please refer to Figure 1 shown in Figure 1 is a schematic flowchart of the video generation method based on speech analysis of the present invention. The present invention is a video generation method based on speech analysis. In some embodiments, Figure 1 the flowchart 100 shown can be executed by an electronic device (such as a processor) with computing capabilities. As Figure 1 shown, the flowchart 100 may include the following operations.

[0051] Step 101, parse the input speech and extract multi-modal speech features; the multi-modal speech features include phoneme sequences, emotional intensity curves, speech rate change parameters, and intonation contours.

[0052] The input speech refers to the speech data containing natural language provided by the user, which can be any speech content such as conversations, narrations, explanations, etc.

[0053] The multi-modal speech features refer to a composite feature set extracted from speech that can reflect multiple aspects of speech content, emotion, structure, etc.

[0054] Phoneme sequences, such as [t], [ʃ], [a], [n], represent the phoneme composition of the word "train". The emotional intensity curve is a graph used to represent the intensity of emotions changing over time, such as the gradual increase of anger. The speech rate change parameter is the pronunciation speed of different segments in speech (unit: syllables per second). The intonation contour refers to the rise and fall of pitch or tone, such as the upward rise at the end of an interrogative sentence.

[0055] In some embodiments, noise removal and frame segmentation can be performed through a speech preprocessing module; a phoneme sequence can be extracted using an audio feature extraction tool (such as OpenSMILE); an emotional intensity curve can be generated by applying an emotion recognition model (such as CNN-LSTM); the speech rate change parameter can be calculated through duration analysis and DTW; and the intonation contour can be obtained using a pitch tracking algorithm (such as YIN).

[0056] Step 102, input the multi-modal speech features into a pre-trained scenario association model, and output a set of scenario labels; the set of scenario labels includes a scene type label, an emotion category label, and an interaction object relationship label.

[0057] The scenario association model is a multi-modal deep neural network, trained through a large amount of speech-scenario paired data, and capable of mapping speech features to scenario labels.

[0058] The set of scenario labels is the set of labels output by the model, used to describe the scenario attributes corresponding to this segment of speech. For example, the scene type label: such as "cafe", "subway", "classroom"; the emotion category label: such as "happy", "angry", "calm"; the interaction object relationship label: such as "between friends", "customer and salesperson", "teacher and student".

[0059] In some embodiments, features such as the extracted phonemes, emotional curves, speech rate, and intonation can be vectorized; input into a multi-modal Transformer or BERT model; and a set of labels is output.

[0060] In some embodiments, the pre-trained scenario association model is obtained in the following manner.

[0061] S10, obtain speech-video alignment samples, each sample including a speech segment and its corresponding video segment, and label the set of scenario labels and geographical labels of the video segment.

[0062] The speech-video alignment sample is a pair of paired samples, which includes a speech segment and the corresponding video segment. The speech and video content are aligned in time, that is, each segment of speech has a corresponding video scene, and vice versa.

[0063] In some embodiments, matching speech-video pairs can be collected from a speech and video database; scenario tags and geographical tags are labeled for each video segment.

[0064] S11. Extract sample speech features from the speech segment, where the sample speech features include a sample phoneme sequence, a sample emotion intensity curve, a sample speech rate change parameter, and a sample intonation contour. Extract sample video features from the video segment, where the sample video features include scene features, character action features, and background element features.

[0065] The sample speech features refer to the feature vectors extracted from the speech segment that describe the speech content, emotion, speech rate, etc.

[0066] In some embodiments, an audio processing tool (such as MFCC or Mel-spectrogram) can be used to extract the phoneme sequence; an emotion analysis algorithm (such as LSTM, CNN) is used to calculate the emotion intensity; the speech rate parameter is calculated by framing the speech segment and analyzing the duration of each frame; the Pitch Tracking algorithm is used to extract the intonation contour.

[0067] The sample video features refer to the features extracted from the video segment that reflect the scene, character actions, and background elements in the video.

[0068] The scene features can include the environment and background information in the video, such as a city, park, beach, etc.

[0069] The character action features can include the behaviors and actions of the characters in the video, such as running, jumping, sitting down, etc.

[0070] The background element features can include the background objects or props in the scene, such as a table, sofa, tree, waves, etc.

[0071] In some embodiments, image processing and video analysis algorithms (such as CNN) can be used to extract the scene features in the video frame; an action recognition algorithm (such as OpenPose) is used to analyze the actions of the characters; the background element features are extracted, and a background separation algorithm is used to extract static or dynamic background objects.

[0072] S12. Input the sample speech features and the sample video features into an encoder to calculate the similarity score of the speech-video pair.

[0073] The encoder is used to map the speech features and video features into a common space for a deep learning model for matching. For example, a two-stream neural network (such as a Transformer Encoder) can process speech and video features simultaneously.

[0074] The similarity score is a metric indicating whether a pair of speech and video match. The larger the value, the higher the degree of match.

[0075] In some embodiments, the speech features and video features can be respectively input into an encoder, and the similarity score can be obtained by calculating the similarity between vectors (such as cosine similarity).

[0076] S13. Construct a contrastive loss function based on the similarity score, and optimize the parameters of the encoder through the backpropagation algorithm, so as to maximize the similarity of matching speech-video pairs and minimize the similarity of non-matching pairs.

[0077] The contrastive loss function is a type of loss function, aiming to optimize the model to make the similarity of matching speech-video pairs higher and the similarity of non-matching pairs lower.

[0078] The backpropagation algorithm is an optimization algorithm commonly used in neural network training. By calculating the gradient and updating the network weights, the loss function is minimized.

[0079] In some embodiments, a contrastive loss function can be constructed based on the similarity score; the loss function can be passed into the backpropagation algorithm; and the parameters of the encoder can be updated through backpropagation to maximize the similarity of matching speech-video pairs.

[0080] In some embodiments, the contrastive loss function is as shown in the following formula (1):

[0081]

[0082] where, V i represents the sample video feature of the i-th video segment, a i represents the sample speech feature of the i-th speech segment, s(v i , a i ) represents the cosine similarity between the sample video feature V i and the sample speech feature a i . τ is the temperature coefficient, and its value range is 0.1 - 1.0, which is used to adjust the smoothness of the similarity distribution. N is the total number of samples in the training batch, and j is the sample index variable, traversing all samples in the batch (j = 1, 2,..., N).

[0083] The goal of the contrastive loss function is to maximize the similarity of matching speech-video pairs (ai and vi), and at the same time minimize the similarity between vi and other non-matching speeches aj (j = i) within the batch.

[0084] S14. Use the trained encoder as the scenario association model for real-time scenario label prediction of the input speech.

[0085] The trained encoder is an optimized encoder model that can generate accurate embedding representations based on the input speech and video features.

[0086] The context association model is an encoder-based model that can predict context labels for new speech input in real time.

[0087] Step 103: Based on the dialect features in the input speech, the regional category is identified through a dialect classifier, and the corresponding visual element library is loaded from the cultural database according to the regional category. The visual element library includes scene templates, character costumes and regional iconic props.

[0088] Dialect features are regional language variant characteristics present in the input speech, including phonetic phonology, intonation, and commonly used vocabulary. For example, in Shanghainese, "I" is pronounced as "wu" and in Cantonese, "yes" is pronounced as "xi".

[0089] A dialect classifier is a machine learning model trained to identify dialect affiliation, with input being a speech signal or its extracted features.

[0090] Regional categories are labels for geographical regions based on dialect classification results. For example, "Guangdong Province", "Shanghai", "Northeast China"

[0091] The cultural database is a structured database used to store regional cultural elements, including clothing, scenes, architecture, props, etc. For example, the database contains data such as "Shanghai-style" cheongsams, alleyway backgrounds, and "Northeastern kangs."

[0092] The visual element library is a collection of images, models and other materials that correspond to regional categories and visually express regional culture.

[0093] In some embodiments, the input speech can be fed into a dialect classifier to identify the region to which it belongs; using the region category as a key, a matching visual element library is retrieved from a cultural database.

[0094] In some embodiments, identifying the regional category by a dialect classifier based on the dialect features in the input speech may include the following operations.

[0095] S20, analyzing the dialect features and phoneme distribution of the input speech by the dialect classifier, and outputting a regional probability distribution.

[0096] A dialect classifier is a specialized machine learning model that analyzes the dialect characteristics of speech input and determines its regional classification. For example, a dialect classification model trained with a convolutional neural network (CNN) or recurrent neural network (RNN) can be used to determine whether a speech is a northern dialect, a southern dialect, or another regional dialect.

[0097] Dialect features are characteristics related to local languages, pronunciation habits, etc. in speech, usually including changes in intonation, timbre, phonemes, etc. of speech.

[0098] Phoneme distribution is the frequency and distribution of each phoneme (such as consonants, vowels, etc.) in the speech sequence in speech, which can reflect the regional characteristics of pronunciation. For example, in Mandarin, "zh", "ch" and "j", "q" often have different pronunciation methods, which may be typical differences between northern and southern dialects.

[0099] Regional probability distribution is a probability value distribution output by the dialect classifier after analyzing dialect features and phoneme distribution, indicating the probabilities that the input speech may belong to different regional categories.

[0100] In some embodiments, the input speech data can be passed to the dialect classifier; the dialect classifier analyzes the dialect features (such as phoneme sequence, intonation) in the speech and generates an analysis result; the classifier outputs a regional probability distribution, indicating the probabilities that the input speech belongs to different regions.

[0101] S21. Determine multiple candidate regional categories according to the regional probability distribution.

[0102] Multiple candidate regional categories are multiple possible regional categories determined according to the regional probability distribution, and these categories are the most likely regions sorted by probability during the classification process.

[0103] In some embodiments, multiple regional categories with higher probabilities can be selected as candidate categories according to the regional probability distribution; these candidate regional categories are passed to subsequent processing steps for further analysis.

[0104] S22. Calculate the fusion weights of multi-regional visual elements based on the regional probability distribution and the multiple candidate regional categories.

[0105] Multi-regional visual elements refer to visual elements related to different regional cultures, such as local architectural styles, clothing, food, etc. For example, the visual elements in the northern region may be snow scenes and northern folk houses, and the visual elements in the southern region may be tropical plants and small bridges over water in the water town.

[0106] The fusion weight is the fusion weight of visual elements of different regional categories calculated according to the regional probability distribution, indicating the influence degree of each regional visual element when generating the video.

[0107] For example, if the regional probability of the southern dialect is 85% and the regional probability of the northern dialect is 10%, then the visual elements of the south may account for 85% of the fusion weight, and the visual elements of the north account for 10%.

[0108] In some embodiments, the fusion weights for each candidate geographical region category can be calculated according to the geographical probability distribution. The weight of each geographical region category is matched with the proportion of its visual elements to generate the fused visual elements.

[0109] For example, for a geographical probability distribution of: 85% for southern dialects and 10% for northern dialects, calculate the fusion weight of the visual elements in the south as 85% and the fusion weight of the visual elements in the north as 10%. Therefore, the generated video footage will tend to show scenes in the south, such as tropical vegetation and the water towns in the south of the Yangtze River.

[0110] In some embodiments, the fusion weights of the multi-geographical visual elements can be calculated based on the geographical probability distribution and the multiple candidate geographical region categories through the following formula (2);

[0111]

[0112] where K is the total number of the multiple candidate geographical region categories, P k represents the probability of the k-th geographical region category in the geographical probability distribution, γ is the sharpening coefficient, γ is greater than or equal to 1, which is used to enhance the weight proportion of the high-probability geographical region, and m is the geographical region category index variable, traversing all K categories, m = 1, 2,..., K.

[0113] The calculated result of the fusion weight W k is used to weightedly fuse the visual elements of different geographical regions. For example, when γ = 2, the weight of the geographical region with the highest probability will be significantly amplified.

[0114] S23. Determine the geographical region category based on the fusion weights.

[0115] In some embodiments, the geographical region category with the highest weight can be selected as the final judged geographical region category according to the fusion weights.

[0116] Step 104. Select a scene template according to the scene type label in the set of scenario tags, and select a character action template in combination with the emotion category label and the interaction object relationship label.

[0117] The scene template is a preset scene model for video background construction, which can dynamically adapt to different scenario semantics. For example, the three-dimensional or two-dimensional video scene frameworks of a coffee shop, a playground, and an office.

[0118] The character action template is a character action model or animation setting that expresses different emotions and interaction methods. For example, shaking hands, jumping, hugging, waving, frowning, etc.

[0119] In some embodiments, the processor may select a matching scene from the scene template library according to the "scene type" in the set of scenario tags, and call the corresponding human actions in the action library according to the "emotion tag" and "interaction object tag".

[0120] Step 105: Calculate the timing distribution of video elements based on the speech rate change parameter, and render the scene template, human action template and the rhythm of the input speech according to the timing distribution through a timing alignment algorithm to generate a target video.

[0121] The timing distribution of video elements is the time period distribution corresponding to each visual element (such as scene switching, human actions, etc.) in the video. For example, from 0 to 3 seconds, "a person walks in" is shown, and from 4 to 7 seconds, "a hug" is shown.

[0122] The timing alignment algorithm is an algorithm used to synchronize and align the speech timeline with the video timeline to ensure that the picture rhythm is consistent with the speech rhythm. For example, the Dynamic Time Warping (DTW) algorithm is used for matching.

[0123] In some embodiments, the processor may analyze the speech rate change parameter, establish a mapping relationship between speech and the time axis; set time anchors for human actions and scene switching; use algorithms such as DTW to match the speech rhythm and the action rhythm; render video elements to the corresponding time periods and output the final target video.

[0124] In some embodiments, calculating the timing distribution of video elements based on the speech rate change parameter includes: segmenting the input speech into multiple speech paragraphs according to the speech rate change parameter; assigning an initial key frame timestamp to each speech paragraph to determine the timing distribution of the video elements.

[0125] The initial key frame timestamp is a time marker assigned to each video paragraph during video generation, used to specify the starting time point of the visual content corresponding to each speech paragraph.

[0126] The timing distribution of video elements is to arrange the visual elements (such as images, video clips, etc.) in the video according to the timing of the speech, in chronological order and according to the timestamps of the speech paragraphs. For example, if the timestamp of a speech paragraph is from 0 seconds to 5 seconds, then within this period of time, video elements (such as a person or the background) will be displayed synchronously.

[0127] In some embodiments, an initial timestamp may be assigned to each speech paragraph according to the duration and content of each paragraph; these timestamps will become the time reference for subsequent video element rendering to ensure that the visual elements are synchronized with the speech in the video.

[0128] In some embodiments, rendering the scene template, the character action template, and the rhythm of the input speech according to the timing distribution through the timing alignment algorithm to generate a target video includes: inserting dynamic special effects within speech paragraphs where the emotional intensity exceeds a preset threshold based on the emotional intensity curve, and the type of the dynamic special effects is determined by the emotional category label; adjusting the initial key frame timestamps of each speech paragraph in the input speech through the timing alignment algorithm so that the mouth movements of the characters in the character action template are aligned with the time of the phoneme sequence and the scene template.

[0129] The preset threshold is a preset emotional intensity value. When the emotional intensity in the speech exceeds this value, certain special effects will be triggered. For example, if the emotional intensity threshold is set to 70, when the emotional intensity of the speech exceeds this value, the system will insert dynamic special effects into the video.

[0130] Dynamic special effects are visual effects in the video used to enhance emotional expression, such as flashing light effects, sparks, background changes, etc.

[0131] The emotional category label is a label attached according to the emotional category recognized in the speech (such as joy, sadness, anger, etc.). For example, if the speech analysis result shows "anger", the emotional category label of this speech paragraph is "anger".

[0132] In some embodiments, it is possible to detect, according to the emotional intensity curve, which speech paragraphs have an emotional intensity exceeding the preset threshold; within these speech paragraphs, select appropriate dynamic special effects according to the emotional category label and insert them into the video.

[0133] The timing alignment algorithm is an algorithm used to adjust the time synchronization of each element (such as character actions, mouth shapes, etc.) in the video with the speech.

[0134] The character action template is a preset sequence of character actions, such as actions of a character speaking, walking, jumping, etc.

[0135] The mouth movements of the characters are the movements of the mouth shapes when the characters are speaking in the video, usually realized through animation techniques.

[0136] In some embodiments, the timing alignment algorithm can be used to analyze the phoneme sequence of each speech paragraph, adjust the synchronization of the mouth shapes of the characters with the speech time; adjust the synchronization of the scene template according to the time of the mouth shapes of the characters so that the visual elements in the video match the speech.

[0137] In some embodiments, the method further includes: obtaining the score of the user for the target video; if the score is lower than the score threshold, increasing the weight of the emotional intensity curve in the training of the scenario association model.

[0138] User rating is the score given by users after evaluating the generated video content, usually based on factors such as video quality and emotional expression. For example, after a user watches a video, the system may ask the user to give a rating from 1 to 5, where 1 means very dissatisfied and 5 means very satisfied.

[0139] The target video is the final video generated by the system. For example, a video about "love story", the system generates the video through speech analysis and the user gives a rating.

[0140] In some embodiments, after a user watches a video, a rating interface can be presented to the user, usually with rating options from 1 to 5. The user selects a rating and the system records the rating.

[0141] The rating threshold is a preset rating standard. When the user's rating is lower than this threshold, it means the video needs further optimization, usually related to video quality, emotional conveyance, etc. For example, assuming the rating threshold is 3 points, if the user gives a video rating of 2 points, it means the video quality is poor and needs optimization.

[0142] Weight is the degree of influence of a certain feature on the final result during model training. A higher weight means that feature has a greater impact on the model prediction result. For example, if the weight of the emotional intensity curve is high, then during the training of the scenario association model, the change of emotion will have a greater impact on the video generation result.

[0143] In some embodiments, the user rating can be obtained and compared with the preset rating threshold; if the rating is lower than the threshold, increase the weight of the emotional intensity curve in the scenario association model training, that is, strengthen the influence of emotional expression on the video generation process.

[0144] The working principle of the present invention: By combining speech analysis, emotional intensity curve, scenario association model and user rating feedback, the present invention can achieve higher emotional expression accuracy, personalized adjustment and dynamic optimization based on user feedback during the video generation process, thereby improving video quality, enhancing user experience and strengthening the emotional resonance of video content.

[0145] The above has described in detail an embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.

Claims

1. A video generation method based on speech analysis, characterized in that The method includes the following steps: Parse the input speech to extract multi-modal speech features; the multi-modal speech features include a phoneme sequence, an emotional intensity curve, a speech rate change parameter, and an intonation contour; Input the multi-modal speech features into a pre-trained scenario association model to output a set of scenario labels; the set of scenario labels includes a scene type label, an emotional category label, and an interaction object relationship label; Based on the dialect features in the input speech, identify the regional category through a dialect classifier, and load the corresponding visual element library from the cultural database according to the regional category. The visual element library includes scene templates, character costumes, and regional iconic props; Select a scene template according to the scene type label in the set of scenario labels, and select a character action template in combination with the emotional category label and the interaction object relationship label; Calculate the temporal distribution of video elements based on the speech rate change parameter, and render the scene template, the character action template, and the rhythm of the input speech according to the temporal distribution through a temporal alignment algorithm to generate a target video.

2. The video generation method based on speech analysis according to claim 1, wherein The pre-trained scenario association model is obtained through the following method: Obtain speech-video alignment samples, each sample including a speech segment and its corresponding video segment, and label the set of scenario labels and regional labels of the video segment; Extract sample speech features from the speech segment. The sample speech features include a sample phoneme sequence, a sample emotional intensity curve, a sample speech rate change parameter, and a sample intonation contour. Extract sample video features from the video segment. The sample video features include scene features, character action features, and background element features; Input the sample speech features and the sample video features into an encoder to calculate the similarity score of the speech-video pair; Construct a contrast loss function based on the similarity score, and optimize the parameters of the encoder through backpropagation algorithm to maximize the similarity of the matching speech-video pairs and minimize the similarity of the non-matching pairs; Use the trained encoder as the scenario association model for real-time scenario label prediction of the input speech.

3. A video generation method based on speech analysis according to claim 2, wherein The contrast loss function is as shown in the following formula: Among them, V i represents the sample video feature of the i-th video segment, a i represents the sample speech feature of the i-th speech segment, s(v i , a i ) represents the cosine similarity between the sample video feature V i and the sample speech feature a i . τ is the temperature coefficient, and its value range is 0.1 - 1.0, which is used to adjust the smoothness of the similarity distribution. N is the total number of samples in the training batch, and j is the sample index variable, traversing all samples in the batch, j = 1, 2,..., N.

4. A video generation method based on speech analysis according to claim 1, characterized in that, The identifying the regional category through a dialect classifier based on the dialect features in the input speech includes: Parse the dialect features and phoneme distribution of the input speech through the dialect classifier to output a regional probability distribution; Determine multiple candidate regional categories according to the regional probability distribution; Calculate the fusion weights of multi-regional visual elements based on the regional probability distribution and the multiple candidate regional categories; Determine the regional category based on the fusion weights.

5. A video generation method based on speech analysis according to claim 4, characterized in that The calculating the fusion weights of multi-regional visual elements based on the regional probability distribution and the multiple candidate regional categories includes: Calculate the fusion weights of the multi-regional visual elements based on the regional probability distribution and the multiple candidate regional categories through the following formula; where K is the total number of the multiple candidate region categories, and P k represents the probability of the k-th region category in the region probability distribution, γ is a sharpening coefficient, γ is greater than or equal to 1 and is used to enhance the weight ratio of high-probability regions, m is a region category index variable that traverses all K categories, and m = 1, 2, …, K.

6. The video generation method based on speech analysis according to claim 1, wherein, The calculating the temporal distribution of video elements based on the speech rate change parameter includes: Segment the input speech into multiple speech paragraphs according to the speech rate change parameter; Assign an initial key frame timestamp to each speech paragraph and determine the temporal distribution of the video elements.

7. The video generation method based on speech analysis according to claim 6, wherein, Rendering the scene template, the character action template and the rhythm of the input speech according to the temporal distribution through a temporal alignment algorithm to generate a target video, including: Based on the emotional intensity curve, insert dynamic special effects within the speech paragraphs where the emotional intensity exceeds a preset threshold, and the type of the dynamic special effects is determined by the emotional category label; Adjust the initial key frame timestamps of each speech paragraph in the input speech through a temporal alignment algorithm so that the lip movement of the character in the character action template aligns with the time of the phoneme sequence and the scene template.

8. The video generation method based on speech analysis according to claim 1, wherein The method further includes: Obtain the user's rating of the target video; If the rating is lower than the rating threshold, increase the weight of the emotional intensity curve in the training of the scenario association model.

9. A video generation system based on speech analysis, characterized in that, The system includes: A speech feature extraction module for parsing the input speech and extracting multi-modal speech features; the multi-modal speech features include a phoneme sequence, an emotional intensity curve, a speech rate change parameter, and an intonation contour; A scenario label acquisition module for inputting the multi-modal speech features into a pre-trained scenario association model and outputting a set of scenario labels; the set of scenario labels includes a scene type label, an emotional category label, and an interaction object relationship label; A regional category recognition module for identifying the regional category based on the dialect features in the input speech through a dialect classifier and loading a corresponding visual element library from a cultural database according to the regional category, and the visual element library includes a scene template, character costumes, and regional iconic props; A scene selection module for selecting a scene template according to the scene type label in the set of scenario labels and selecting a character action template in combination with the emotional category label and the interaction object relationship label; A target video generation module for calculating the temporal distribution of video elements based on the speech rate change parameter and rendering the scene template, the character action template and the rhythm of the input speech according to the temporal distribution through a temporal alignment algorithm to generate a target video.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the video generation method based on speech analysis according to any one of claims 1-8.

Citation Information

Cited By

  • Companion robot emotional intelligence analysis method based on multi-modal data analysis

    CN122527842A