Short video script matching generation method and system combined with big data analysis
By combining big data analysis and multimodal data processing, the problem of lack of real-time analysis and intelligent adjustment in short video script generation technology is solved, and the high consistency between emotions and plots and real-time optimization driven by user feedback is achieved, which improves the personalization and dissemination effect of short videos.
Patent Information
- Application Number
- CN202510054570.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing short video script generation technology lacks real-time analysis and intelligent adjustment capabilities, cannot quickly respond to changes in user preferences and content trends, and only considers single-modal data, neglecting the integrated analysis of multimodal data, resulting in possible inconsistencies between emotional expression and plot design.
By combining the big data analysis method, multimodal data (text, images, sound) of short videos are obtained, consistency analysis and sentiment analysis are performed, consistency coefficients and emotional accuracy coefficients are calculated, script content is automatically adjusted to ensure consistency of emotions and plots, and real-time optimization and adjustment are made based on user feedback data.
Real-time optimization of short video scripts has been achieved, the personalization and intelligence of content has been improved, the high consistency of emotional expression and plot design has been ensured, and the dissemination effect and user satisfaction of short videos have been significantly improved.
Smart Images

Figure CN119988994A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of short video script optimization, and in particular to a short video script matching generation method and system combined with big data analysis. Background Art
[0002] With the rapid development of short video platforms, short videos have become an important form of daily content consumption for Internet users. In order to improve the user appeal and dissemination effect of short videos, content creators are constantly exploring strategies to optimize short video scripts. Short video scripts not only need to consider plot design, content layout and emotional expression, but must also be personalized according to the needs and preferences of different users. However, traditional short video script generation methods mostly rely on manual design, lack flexible real-time optimization mechanisms, and are difficult to quickly respond to changes in user preferences and content trends. Therefore, how to use big data and intelligent algorithms to dynamically optimize short video scripts and improve the personalization and intelligence level of content has become an urgent problem to be solved in the current short video creation field.
[0003] The prior art has the following deficiencies:
[0004] Existing short video script generation technologies usually rely on static creation rules and preset content structures, lacking the ability of real-time analysis and intelligent adjustment. Although some recommendation algorithms based on user data analysis can provide personalized content push, these methods often ignore the dynamic adjustment at the level of short video scripts. Traditional algorithms are usually based on historical data, respond slowly to user feedback, and cannot capture changes in user preferences in a timely manner. In addition, existing short video script generation methods usually only consider single-modal data, lack of integrated analysis of multimodal data, resulting in possible inconsistencies between emotional expression and plot design, affecting the overall effect of short videos. Therefore, how to achieve real-time optimization of short video scripts based on multimodal data and user feedback is a bottleneck problem that needs to be solved in the prior art. The present invention solves this problem by using a dynamic optimization algorithm based on user feedback data and multimodal data, which can respond to changes in user interests in real time, improve the personalization and accuracy of short video scripts, and thus effectively improve the quality and dissemination effect of short video content. Summary of the invention
[0005] The purpose of the present invention is to provide a short video script matching generation method and system combined with big data analysis to solve the problems in the above background.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] The short video script matching generation method combined with big data analysis includes the following steps:
[0008] S1: Acquire multimodal data of a short video and perform preprocessing, wherein the multimodal data of the short video includes text data, image data, and sound data;
[0009] The text data includes the title, description, script content, and user comments of the video, the image data includes the image of each frame in the video, and the sound data includes the audio information in the video;
[0010] S2: performing consistency analysis on the multimodal data and calculating a consistency coefficient, wherein the consistency analysis includes checking whether there are inconsistencies between text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality to evaluate the consistency of the multimodal data content;
[0011] S3: According to the evaluation results, the multimodal data is marked as having consistent content and inconsistent content. On the basis of the consistency of the multimodal data content, the accuracy of the emotional tendency and theme of the short video is evaluated through sentiment analysis and theme classification;
[0012] The sentiment analysis is based on text sentiment, voice tone and pitch to assess sentiment tendency;
[0013] The topic classification matches the topic of the video by analyzing the text content and the image data;
[0014] S4: Based on the evaluation results, extract the accurate emotional tendency and theme of the short video, detect the conflict between emotion and plot, and determine whether the emotion and plot conflict by calculating the matching coefficient between emotion and plot;
[0015] S5: according to the judgment result, if the sentiment analysis result conflicts with the plot design in the video, the generated script content is automatically adjusted to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy;
[0016] S6: Real-time optimization and adjustment based on user feedback data. Dynamic optimization of the script generation process is performed through real-time data collection, including likes, comments, shares, and viewing time.
[0017] As a further solution of the present invention: the consistency analysis of the multimodal data and calculation of the consistency coefficient specifically includes:
[0018] Extract and represent features for each modality including text, image and audio;
[0019] Among them, the text data is processed through a natural language processing model to extract text feature vectors;
[0020] The image data is processed through a convolutional neural network to extract the feature vector of each frame.
[0021] Audio data is extracted from audio feature vectors through Mel spectrum;
[0022] Calculate the joint probability distribution between each pair of modes;
[0023]
[0024] Where P(X,Y) represents the joint probability distribution between each pair of modes, t i represents the modal feature of the t-th i-th mode in each pair of modes, z j represents the jth modal feature of z in each pair of modes, δ represents the matching index function, N represents the total number of each mode, i represents the feature vector in one mode, and j represents the feature vector in another mode;
[0025] The pair of modalities includes: text and image, text and audio, and image and audio;
[0026] The marginal distribution of each mode is obtained by summing the joint probability of each mode;
[0027] Calculate the maximum mutual information between each pair of modes. The calculation expression is:
[0028]
[0029] Where P(x,y) represents the joint probability distribution of two modes X and Y, P(x) is the marginal distribution of mode X, P(y) is the marginal distribution of mode Y, and X and Y represent any two modes;
[0030] According to the maximum mutual information corresponding to each mode, the consistency coefficient is calculated, and the calculation expression is:
[0031]
[0032] In the formula, C total represents the consistency coefficient, I(T,I) represents the maximum mutual information between text and image, I(T,A) represents the maximum mutual information between text and audio, and I(I,A) represents the maximum mutual information between image and audio.
[0033] As a further solution of the present invention: the accuracy assessment of the emotional tendency and theme of the short video through sentiment analysis and theme classification specifically includes:
[0034] Based on the data after consistency verification, sentiment analysis and topic classification are performed, and the sentiment accuracy coefficient and topic accuracy coefficient are calculated. The sentiment accuracy coefficient and topic accuracy coefficient are normalized and comprehensively calculated to obtain the comprehensive accuracy coefficient.
[0035] It is determined whether the comprehensive accuracy coefficient is greater than or equal to a preset threshold value. If so, the emotional tendency and theme of the corresponding short video are accurate; if not, the emotional tendency and theme of the corresponding short video are inaccurate.
[0036] As a further solution of the present invention: the process of obtaining the emotion accuracy coefficient is:
[0037] Obtain multimodal data of short videos and perform feature extraction on the multimodal data;
[0038] Performing multimodal feature fusion on the feature vectors of the extracted multimodal data, including weighted fusion on the feature vectors of the multimodal data to obtain a multimodal joint feature vector, and obtaining a multimodal joint feature vector;
[0039] The feature vector of the multimodal data includes: text features, image features and audio features;
[0040] Based on the multimodal joint feature vector, a fully connected neural network is constructed as a sentiment classification model;
[0041] The cross entropy loss function is used to calculate the difference between the true emotion label and the emotion label predicted by the model. The calculation expression is:
[0042]
[0043] In the formula, M represents the total number of training samples, a represents the sample, b represents the emotion category, and C represents the total number of emotion categories. represents the loss function, represents the predicted probability of the bth emotion of the ath sample, y a,b Indicates the true emotion label of the bth emotion of the ath sample;
[0044] Iteratively update model parameters to optimize the performance of the sentiment classification model;
[0045] Based on the trained sentiment classification model, multimodal data is input to calculate the sentiment classification results and the sentiment accuracy coefficient. The calculation expression is:
[0046]
[0047] In the formula, Z C represents the sentiment accuracy coefficient, represents the matching function, y a represents the true sentiment label of the a-th sample, represents the predicted probability of sentiment of the ath sample;
[0048] Classify short video emotions according to the emotion accuracy coefficient.
[0049] As a further solution of the present invention: the process of obtaining the subject accuracy coefficient is as follows:
[0050] Obtain multimodal data of short videos and perform feature extraction on the multimodal data;
[0051] The dot product of the feature vector of each mode is used to calculate the similarity of each mode. The calculation expression is:
[0052]
[0053] In the formula, S(X,Y) represents the similarity of each mode, E X represents the eigenvector in mode X, E Y represents the eigenvector in mode Y;
[0054] Calculate the comprehensive similarity score between text, image and audio by weighted fusion of multimodal similarity scores;
[0055] For each frame of the image and the corresponding audio clip in the short video, calculate the comprehensive similarity score of each frame;
[0056] Based on the average of the multimodal similarity scores, the topic accuracy coefficient of the short video is calculated, and the calculation expression is:
[0057]
[0058] In the formula, T represents text features, I represents image features, A represents audio features, d represents each frame, H represents the total number of each frame, and Z D Represents the topic accuracy coefficient, S(T,I,A) d Represents the comprehensive similarity score between each frame and the corresponding audio segment.
[0059] As a further solution of the present invention: the process of obtaining the matching coefficient is:
[0060] Based on the text data of short videos, the text Transformer encoder in the VATT model is used to extract the emotional feature vector E of the text. t and the plot feature vector C t ;
[0061] Based on each frame of the short video, the visual Transformer encoder in the VATT model is used to extract the emotional feature vector E of the image. b and the plot feature vector C v ;
[0062] Based on the audio data of the short video, the spectrogram is generated by short-time Fourier transform, and the audio Transformer encoder in the VATT model is used to extract the audio's emotional feature vector E.a and the plot feature vector C a ;
[0063] Based on text, image and audio modalities, the similarity between the emotional feature vector and the plot feature vector is calculated respectively. The calculation expression is:
[0064]
[0065] In the formula, i represents each mode, E i represents the emotional feature vector of modality i, C i Represents the plot feature vector of modality i, which includes text, image, and audio, ∥E i ∥ is the Euclidean norm of the sentiment feature vector, which is calculated by the Euclidean distance norm calculation expression. ∥C i ∥ is the Euclidean norm of the plot feature vector, which is calculated by the Euclidean distance norm calculation expression;
[0066] The similarity results of each modality are weighted and fused to calculate the matching coefficient between emotion and plot. The calculation expression is:
[0067]
[0068] Where M score Represents the matching coefficient, W i Represents the preset weight of the i-th mode.
[0069] As a further solution of the present invention: if the sentiment analysis result conflicts with the plot design in the video, the generated script content is automatically adjusted to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy, specifically including:
[0070] Extracting emotional features of the short video script content through sentiment analysis, and extracting plot features of the video plot design using the video plot feature generation model;
[0071] According to the matching coefficient between emotional features and plot features, determine whether there is a conflict between the script content and the video plot design;
[0072] Extract emotional keywords K from short video scripts through emotional feature vectors e ;
[0073] Extracting plot keywords K from video plots through plot feature vectors c ;
[0074] To judge the semantic conflict between emotional keywords and plot keywords, the cosine similarity calculation formula based on keyword embedding vector is:
[0075]
[0076] In the formula, Conflict(k e ,k c ) represents the semantic conflict coefficient. If the semantic conflict coefficient is greater than the preset threshold, there is a conflict between the emotional keyword and the plot keyword;
[0077] According to the conflict analysis results, a new script content is generated that matches the video plot keywords, where the emotional keywords in the new script content meet the optimization constraints. The calculation expression is:
[0078]
[0079] Among them, Sim(K ′ e ,K c ) represents the semantic similarity between the optimized emotional keywords and plot keywords, ensuring the consistency between the script content and the plot design, k ′ e represents the optimized emotional keyword set, k c represents the set of plot keywords, k ′ e ∈K ′ e represents any emotional keyword in the emotional keyword set, k c ∈K c represents any plot keyword in the plot keyword set, max represents the maximum value, k ′ e ·k c Represents the vector dot product of sentiment keywords and plot keywords, k ′ e ∥ represents the norm of the sentiment keyword vector, ∥k c ∥ represents the norm of the plot keyword vector;
[0080] Regenerate the short video script based on the optimized script content, update the title, description and emotional tendency information in the script, and output the optimized short video script file.
[0081] The short video script matching generation system combined with big data analysis includes:
[0082] A data acquisition module, wherein the data acquisition module obtains and preprocesses multimodal data of the short video, wherein the multimodal data of the short video includes text data, image data, and sound data;
[0083] A multimodal data consistency assessment module, which is used to perform consistency analysis on the multimodal data and calculate a consistency coefficient, wherein the consistency analysis includes checking whether there are inconsistencies between text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality, which is used to evaluate the consistency of the multimodal data content;
[0084] An accuracy assessment module, wherein the accuracy assessment module marks the multimodal data as having consistent content or inconsistent content according to the assessment result, and performs an accuracy assessment on the emotional tendency and theme of the short video through sentiment analysis and theme classification on the basis of the consistency of the multimodal data content;
[0085] A conflict analysis module, which extracts the accurate emotional tendency and theme of the short video according to the evaluation results, detects the conflict between the emotion and the plot, and determines whether the emotion and the plot conflict by calculating the matching coefficient between the emotion and the plot;
[0086] A script optimization module, wherein, based on the judgment result, if the sentiment analysis result conflicts with the plot design in the video, the script optimization module automatically adjusts the generated script content to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy;
[0087] A user feedback and real-time adjustment module performs real-time optimization and adjustment based on user feedback data, and dynamically optimizes the script generation process through real-time data collection, including likes, comments, shares, and viewing time.
[0088] Beneficial effects of the present invention:
[0089] (1) The present invention constructs an efficient multimodal consistency evaluation method through comprehensive collection and in-depth analysis of multimodal data of short videos. By introducing the calculation mechanism of consistency coefficient, the system can accurately identify and quantify the content relevance and logical consistency between text, image and audio data, avoiding the problem of script design deviation caused by conflict or mismatch between multimodal data. On this basis, by using sentiment analysis and topic classification technology, by integrating text sentiment tendency analysis, image visual content recognition and audio sentiment feature detection, the accuracy coefficient of sentiment and theme is comprehensively calculated to ensure that the generation of short video script content can highly meet the actual emotional expression and theme requirements of video content. In the scenario where there is a conflict between emotional tendency and plot design, the present invention uses deep neural network to perform semantic analysis of multimodal features, combines comparative analysis of emotional features and plot features, accurately judges conflict points and automatically adjusts script content, and optimizes the semantic matching between emotional keywords and plot keywords. This mechanism can make the generated script content more natural and smooth and the emotional logic highly consistent, ensuring the high unity of short videos in emotional expression and narrative structure, thereby significantly improving the accuracy, rationality and practicality of script generation, making it more valuable for dissemination and attractive to users.
[0090] (2) By analyzing user feedback (including likes, comments, shares, and viewing time) in real time and combining it with a dynamic optimization algorithm, the system can generate personalized optimization strategies and dynamically adjust key parameters in the script generation process (such as plot arrangement and emotional expression) based on these strategies. Specifically, the system acquires and processes user behavior data in real time through an iterative optimization process, evaluates user preferences and content trends, and thus achieves precise adjustment of script content. During the implementation process, information from multimodal data sources is first collected to ensure the diversity and comprehensiveness of the data. Next, the system analyzes the consistency between multimodal data to ensure the consistency of different data sources in terms of themes and emotions, and based on this, evaluates the accuracy of the emotional tendency and theme of the short video. Through the dynamic optimization strategy driven by feedback data, the system can respond to user needs in real time and adjust the structure and content of the short video to improve its dissemination effect and user engagement. Finally, through continuous iterative optimization, the system not only improves the personalization and attractiveness of the short video script, but also significantly enhances the responsiveness of the short video content to user needs, thereby significantly improving the intelligence level of the content generation process and user satisfaction. This innovative real-time optimization mechanism based on user feedback breaks through the limitations of traditional content creation and brings more efficient and accurate script generation and optimization solutions to short video production. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] The present invention will be further described below in conjunction with the accompanying drawings.
[0092] Figure 1It is a flowchart of the specific steps of the short video script matching generation method combined with big data analysis of the present invention;
[0093] Figure 2 It is a flowchart of the short video script matching generation system combined with big data analysis in the present invention. DETAILED DESCRIPTION
[0094] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0095] See also Figure 1 As shown, the present invention is a short video script matching generation method combined with big data analysis, comprising the following steps:
[0096] S1: Acquire multimodal data of a short video and perform preprocessing, wherein the multimodal data of the short video includes text data, image data, and sound data;
[0097] The text data includes the title, description, script content, and user comments of the video, the image data includes the key frame images in the video, and the sound data includes the audio information in the video;
[0098] S2: performing consistency analysis on the multimodal data and calculating a consistency coefficient, wherein the consistency analysis includes checking whether there is any inconsistency between the text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality, so as to evaluate the consistency of the content of the multimodal data;
[0099] S3: According to the evaluation results, the multimodal data is marked as having consistent content and inconsistent content. On the basis of the consistency of the multimodal data content, the accuracy of the emotional tendency and theme of the short video is evaluated through sentiment analysis and theme classification;
[0100] The sentiment analysis is based on text sentiment, voice tone and pitch to assess sentiment tendency;
[0101] The topic classification matches the topic of the video by analyzing the text content and the image data;
[0102] S4: Based on the evaluation results, extract the accurate emotional tendency and theme of the short video, detect the conflict between emotion and plot, and determine whether the emotion and plot conflict by calculating the matching coefficient between emotion and plot;
[0103] S5: according to the judgment result, if the emotion analysis result conflicts with the plot design in the video, automatically adjusting the generated script content to ensure the consistency of emotion and plot, thereby optimizing the script generation strategy;
[0104] S6: Real-time optimization and adjustment based on user feedback data. Through real-time data collection (including likes, comments, shares, and viewing time), the script generation process is dynamically optimized so that the generated short video script can be continuously improved according to user needs and content trends.
[0105] In S1, multimodal data of a short video is obtained and preprocessed, wherein the multimodal data of the short video includes text data, image data, and sound data, specifically including:
[0106] The collection of text data includes extracting titles, descriptions, and script content from the metadata of short videos, and capturing comment text related to short videos through the user comment interface;
[0107] The image data collection uses a frame extraction algorithm to parse the short video content into continuous frame images, ensuring that each frame image can fully reflect the dynamic change process of the video;
[0108] The collection of sound data is based on audio separation technology, which extracts the audio tracks in short videos into independent audio files and associates them with timestamps;
[0109] During the acquisition process, in order to ensure the integrity and accuracy of multimodal data, the module uses a time synchronization mechanism to align text, image and audio data;
[0110] The text data is preprocessed through the natural language processing interface to remove noise information, the image data is preprocessed through the image quality detection algorithm to remove low resolution and blurry frames, and the sound data is preprocessed through denoising and normalization to improve the audio quality;
[0111] All data is stored in the data warehouse in a structured manner, providing a basis for subsequent analysis and processing.
[0112] In S2: performing consistency analysis on the multimodal data and calculating a consistency coefficient, the consistency analysis includes checking whether there is inconsistency between text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality, which is used to evaluate the consistency of the multimodal data content, specifically including:
[0113] The performing consistency analysis on the modal data and calculating the consistency coefficient specifically includes:
[0114] Extract and represent features for each modality including text, image and audio;
[0115] Among them, the text data is processed through a natural language processing model to extract text feature vectors;
[0116] The image data is processed through a convolutional neural network to extract the feature vector of each frame.
[0117] Audio data is extracted from audio feature vectors through Mel spectrum;
[0118] Calculate the joint probability distribution between each pair of modes;
[0119]
[0120] Where P(X,Y) represents the joint probability distribution between each pair of modes, t i represents the modal feature of the t-th i-th mode in each pair of modes, z j represents the jth modal feature of z in each pair of modes, δ represents the matching index function, N represents the total number of each mode, i represents the feature vector in one mode, and j represents the feature vector in another mode;
[0121] The pair of modalities includes: text and image, text and audio, and image and audio;
[0122] The marginal distribution of each mode is obtained by summing the joint probability of each mode;
[0123] Calculate the maximum mutual information between each pair of modes. The calculation expression is:
[0124]
[0125] Where P(x,y) represents the joint probability distribution of two modes X and Y, P(x) is the marginal distribution of mode X, P(y) is the marginal distribution of mode Y, and X and Y represent any two modes;
[0126] According to the maximum mutual information corresponding to each mode, the consistency coefficient is calculated, and the calculation expression is:
[0127]
[0128] In the formula, C total represents the consistency coefficient, I(T,I) represents the maximum mutual information between text and image, I(T,A) represents the maximum mutual information between text and audio, and I(I,A) represents the maximum mutual information between image and audio;
[0129] It is determined whether the consistency coefficient is greater than or equal to a preset threshold. If so, it means that the corresponding multimodal data content is consistent and is marked as consistent in content. If not, it means that the corresponding multimodal data content is inconsistent and is marked as inconsistent in content.
[0130] In S3, based on the consistency of the multimodal data content, the accuracy of the emotional tendency and theme of the short video is evaluated through sentiment analysis and theme classification, specifically including:
[0131] Based on the data after consistency verification, sentiment analysis and topic classification are performed, and the sentiment accuracy coefficient and topic accuracy coefficient are calculated. The sentiment accuracy coefficient and topic accuracy coefficient are normalized and comprehensively calculated to obtain the comprehensive accuracy coefficient.
[0132] Determine whether the comprehensive accuracy coefficient is greater than or equal to a preset threshold value. If so, the emotional tendency and theme of the corresponding short video are accurate; if not, the emotional tendency and theme of the corresponding short video are inaccurate;
[0133] The acquisition process of the emotion accuracy coefficient is as follows:
[0134] Obtain multimodal data of short videos and perform feature extraction on the multimodal data;
[0135] Performing multimodal feature fusion on the feature vectors of the extracted multimodal data, including weighted fusion on the feature vectors of the multimodal data to obtain a multimodal joint feature vector, and obtaining a multimodal joint feature vector;
[0136] The feature vector of the multimodal data includes: text features, image features and audio features;
[0137] Based on the multimodal joint feature vector, a fully connected neural network is constructed as a sentiment classification model;
[0138] The cross entropy loss function is used to calculate the difference between the true emotion label and the emotion label predicted by the model. The calculation expression is:
[0139]
[0140] In the formula, M represents the total number of training samples, a represents the sample, b represents the emotion category, and C represents the total number of emotion categories. represents the loss function, represents the predicted probability of the bth emotion of the ath sample, y a,b Indicates the true emotion label of the bth emotion of the ath sample;
[0141] Iteratively update model parameters to optimize the performance of the sentiment classification model;
[0142] Based on the trained sentiment classification model, multimodal data is input to calculate the sentiment classification results and the sentiment accuracy coefficient. The calculation expression is:
[0143]
[0144] In the formula, Z C represents the sentiment accuracy coefficient, represents the matching function, y a represents the true sentiment label of the a-th sample, represents the predicted probability of sentiment of the ath sample;
[0145] Classify short video emotions according to the emotion accuracy coefficient;
[0146] The process of obtaining the topic accuracy coefficient is as follows:
[0147] Obtain multimodal data of short videos and perform feature extraction on the multimodal data;
[0148] The dot product of the feature vector of each mode is used to calculate the similarity of each mode. The calculation expression is:
[0149]
[0150] In the formula, S(X,Y) represents the similarity of each mode, E X represents the eigenvector in mode X, E Y represents the eigenvector in mode Y;
[0151] Calculate the comprehensive similarity score between text, image and audio by weighted fusion of multimodal similarity scores;
[0152] For each frame of the image and the corresponding audio clip in the short video, calculate the comprehensive similarity score of each frame;
[0153] Based on the average of the multimodal similarity scores, the topic accuracy coefficient of the short video is calculated, and the calculation expression is:
[0154]
[0155] In the formula, T represents text features, I represents image features, A represents audio features, d represents each frame, H represents the total number of each frame, and Z D Represents the topic accuracy coefficient, S(T,I,A) d Represents the comprehensive similarity score between each frame and the corresponding audio segment;
[0156] The calculation expression of the comprehensive accuracy coefficient is:
[0157]
[0158] In the formula, H Z represents the comprehensive accuracy coefficient, g1 and g2 represent the preset proportional coefficients, and both g1 and g2 are greater than 0, Z C represents the sentiment accuracy coefficient, Z DRepresents the topic accuracy coefficient.
[0159] It should be noted that the comprehensive accuracy coefficient reflects the accuracy of the emotional tendency and theme of the short video, and the larger the value of the comprehensive accuracy coefficient, the higher the accuracy of the emotional tendency and theme of the corresponding short video.
[0160] In S4, based on the evaluation results, the emotional tendency and theme of the short video are accurately extracted, the conflict between the emotion and the plot is detected, and the matching coefficient between the emotion and the plot is calculated to determine whether the emotion and the plot conflict, including:
[0161] Based on the text data of short videos, the text Transformer encoder in the VATT model is used to extract the emotional feature vector E of the text. t and the plot feature vector C t ;
[0162] Based on each frame of the short video, the visual Transformer encoder in the VATT model is used to extract the emotional feature vector E of the image. v and the plot feature vector C v ;
[0163] Based on the audio data of the short video, the spectrogram is generated by short-time Fourier transform, and the audio Transformer encoder in the VATT model is used to extract the audio's emotional feature vector E. a and the plot feature vector C a ;
[0164] Based on text, image and audio modalities, the similarity between the emotional feature vector and the plot feature vector is calculated respectively. The calculation expression is:
[0165]
[0166] In the formula, i represents each mode, E i represents the emotional feature vector of modality i, C i Represents the plot feature vector of modality i, which includes text, image, and audio, ∥E i ∥ is the Euclidean norm of the sentiment feature vector, which is calculated by the Euclidean distance norm calculation expression. ∥C i ∥ is the Euclidean norm of the plot feature vector, which is calculated by the Euclidean distance norm calculation expression;
[0167] The similarity results of each modality are weighted and fused to calculate the matching coefficient between emotion and plot. The calculation expression is:
[0168]
[0169] Where Mscore Represents the matching coefficient, W i represents the preset weight of the i-th mode;
[0170] It is determined whether the matching coefficient is greater than or equal to a preset threshold. If so, it means that the matching degree between the corresponding modal emotion and the plot is high. If not, it means that the matching degree between the corresponding modal emotion and the plot is low.
[0171] In S5, according to the judgment result, if the sentiment analysis result conflicts with the plot design in the video, the generated script content is automatically adjusted to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy, specifically including:
[0172] Extracting emotional features of the short video script content through sentiment analysis, and extracting plot features of the video plot design using the video plot feature generation model;
[0173] According to the matching coefficient between emotional features and plot features, determine whether there is a conflict between the script content and the video plot design;
[0174] Extract emotional keywords K from short video scripts through emotional feature vectors e ;
[0175] Extracting plot keywords K from video plots through plot feature vectors c ;
[0176] To judge the semantic conflict between emotional keywords and plot keywords, the cosine similarity calculation formula based on keyword embedding vector is:
[0177]
[0178] In the formula, Conflict(k e ,k c ) represents the semantic conflict coefficient. If the semantic conflict coefficient is greater than the preset threshold, there is a conflict between the emotional keyword and the plot keyword;
[0179] According to the conflict analysis results, a new script content is generated that matches the video plot keywords, where the emotional keywords in the new script content meet the optimization constraints. The calculation expression is:
[0180]
[0181] Among them, Sim(K ′ e ,K c ) represents the semantic similarity between the optimized emotional keywords and plot keywords, ensuring the consistency between the script content and the plot design, k ′ e represents the optimized emotional keyword set, kc represents the set of plot keywords, k ′ e ∈K ′ e represents any emotional keyword in the emotional keyword set, k c ∈K c represents any plot keyword in the plot keyword set, max represents the maximum value, k ′ e ·k c Represents the vector dot product of sentiment keywords and plot keywords, k ′ e ∥ represents the norm of the sentiment keyword vector, ∥k c ∥ represents the norm of the plot keyword vector;
[0182] Regenerate the short video script based on the optimized script content, update the title, description and emotional tendency information in the script, and output the optimized short video script file.
[0183] In S6: Real-time optimization and adjustment based on user feedback data. Through real-time data collection, including likes, comments, shares, and viewing time, the script generation process is dynamically optimized, including:
[0184] The system uses the IoT collection module to obtain real-time feedback data such as users’ likes, comments, shares, and viewing time on short videos, and stores it in the feedback database;
[0185] The system first cleans and preprocesses these data, including removing noise data and outliers, and assigning weights to different types of feedback data. For example, it sets higher weights for viewing time and sharing behavior to more comprehensively reflect users' true preferences for short video content.
[0186] On the basis of data preprocessing, the system uses a dynamic optimization algorithm based on deep learning to analyze the feedback data and generate optimization strategies. Through training models, the system identifies the emotional tendencies and thematic directions of user preferences, and combines the multimodal data of short videos to dynamically adjust key parameters in script generation, including content length, plot arrangement, emotional expression, etc.
[0187] The system continuously iterates and optimizes the script generation strategy through real-time updates of feedback data, so that the generated short video content can be consistent with user needs and content trends, thereby improving the content appeal and dissemination effect of the script.
[0188] See also Figure 2 As shown, the short video script matching generation system combined with big data analysis includes:
[0189] A data acquisition module, wherein the data acquisition module obtains and preprocesses multimodal data of the short video, wherein the multimodal data of the short video includes text data, image data, and sound data;
[0190] A multimodal data consistency assessment module, which is used to perform consistency analysis on the multimodal data and calculate a consistency coefficient, wherein the consistency analysis includes checking whether there are inconsistencies between text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality, which is used to evaluate the consistency of the multimodal data content;
[0191] An accuracy assessment module, wherein the accuracy assessment module marks the multimodal data as having consistent content or inconsistent content according to the assessment result, and performs an accuracy assessment on the emotional tendency and theme of the short video through sentiment analysis and theme classification on the basis of the consistency of the multimodal data content;
[0192] A conflict analysis module, which extracts the accurate emotional tendency and theme of the short video according to the evaluation results, detects the conflict between the emotion and the plot, and determines whether the emotion and the plot conflict by calculating the matching coefficient between the emotion and the plot;
[0193] A script optimization module, wherein, based on the judgment result, if the sentiment analysis result conflicts with the plot design in the video, the script optimization module automatically adjusts the generated script content to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy;
[0194] User feedback and real-time adjustment module, which performs real-time optimization and adjustment based on user feedback data, dynamically optimizes the script generation process through real-time data collection, including likes, comments, shares, and viewing time
[0195] The working principle of the present invention is as follows: the multimodal data (including text, image and audio data) of the short video is obtained and preprocessed to ensure the integrity and accuracy of the data; secondly, the consistency analysis of the multimodal data is performed, the consistency of each modal content is evaluated by calculating the consistency coefficient, and sentiment analysis and theme classification are carried out based on the consistency results, and the comprehensive accuracy coefficient is calculated to judge the accuracy of the emotion and theme; then, the emotional tendency and theme of the short video are extracted, the conflict between emotion and plot is detected, and the script content is analyzed and optimized by the matching coefficient to ensure the consistency of emotion and plot; finally, combined with user feedback data (including likes, comments, sharing, viewing time, etc.), the script generation process is adjusted and optimized in real time, so that the script content can dynamically respond to user needs and content trends, thereby improving the dissemination effect and user satisfaction of short videos. This method can effectively improve the generation quality and dynamic adaptability of short video scripts, and provides an efficient and intelligent solution for short video content production.
[0196] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0197] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0198] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0199] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0200] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A short video script matching generation method combined with big data analysis, characterized in that: The following steps are involved: S1: Acquire multimodal data of a short video and perform preprocessing, wherein the multimodal data of the short video includes text data, image data, and sound data; The text data includes the title, description, script content, and user comments of the video, the image data includes the image of each frame in the video, and the sound data includes the audio information in the video; S2: performing consistency analysis on the multimodal data and calculating a consistency coefficient, wherein the consistency analysis includes checking whether there are inconsistencies between text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality to evaluate the consistency of the multimodal data content; S3: According to the evaluation results, the multimodal data is marked as having consistent content and inconsistent content. On the basis of the consistency of the multimodal data content, the accuracy of the emotional tendency and theme of the short video is evaluated through sentiment analysis and theme classification; The sentiment analysis is based on text sentiment, voice tone and pitch to assess sentiment tendency; The topic classification matches the topic of the video by analyzing the text content and the image data; S4: According to the evaluation results, extract the accurate emotional tendency and theme of the short video, detect the conflict between emotion and plot, and determine whether the emotion and plot conflict by calculating the matching coefficient between emotion and plot; S5: according to the judgment result, if the sentiment analysis result conflicts with the plot design in the video, automatically adjusting the generated script content to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy; S6: Real-time optimization and adjustment based on user feedback data. Dynamic optimization of the script generation process is performed through real-time data collection, including likes, comments, shares, and viewing time.
2. The short video script matching generation method combined with big data analysis according to claim 1 is characterized in that: The performing consistency analysis on the multimodal data and calculating the consistency coefficient specifically includes: Extract and represent features for each modality including text, image and audio; Among them, the text data is processed through a natural language processing model to extract text feature vectors; The image data is processed through a convolutional neural network to extract the feature vector of each frame. Audio data is extracted from audio feature vectors through Mel spectrum; Calculate the joint probability distribution between each pair of modes; Where P(X,Y) represents the joint probability distribution between each pair of modes, t i represents the modal feature of the t-th i-th mode in each pair of modes, z j represents the jth modal feature of z in each pair of modes, δ represents the matching index function, N represents the total number of each mode, i represents the feature vector in one mode, and j represents the feature vector in another mode; The pair of modalities includes: text and image, text and audio, and image and audio; The marginal distribution of each mode is obtained by summing the joint probability of each mode; Calculate the maximum mutual information between each pair of modes. The calculation expression is: Where P(x,y) represents the joint probability distribution of two modes X and Y, P(x) is the marginal distribution of mode X, P(y) is the marginal distribution of mode Y, and X and Y represent any two modes; According to the maximum mutual information corresponding to each mode, the consistency coefficient is calculated, and the calculation expression is: In the formula, C total represents the consistency coefficient, I(T,I) represents the maximum mutual information between text and image, I(T,A) represents the maximum mutual information between text and audio, and I(I,A) represents the maximum mutual information between image and audio.
3. The short video script matching generation method combined with big data analysis according to claim 1 is characterized in that: The accuracy assessment of the emotional tendency and theme of the short video through sentiment analysis and theme classification specifically includes: Based on the data after consistency verification, sentiment analysis and topic classification are performed, and the sentiment accuracy coefficient and topic accuracy coefficient are calculated. The sentiment accuracy coefficient and topic accuracy coefficient are normalized and comprehensively calculated to obtain the comprehensive accuracy coefficient. It is determined whether the comprehensive accuracy coefficient is greater than or equal to a preset threshold value. If so, the emotional tendency and theme of the corresponding short video are accurate; if not, the emotional tendency and theme of the corresponding short video are inaccurate.
4. The short video script matching generation method combined with big data analysis according to claim 3 is characterized in that: The acquisition process of the emotion accuracy coefficient is as follows: Obtain multimodal data of short videos and perform feature extraction on the multimodal data; Performing multimodal feature fusion on the feature vectors of the extracted multimodal data, including weighted fusion on the feature vectors of the multimodal data to obtain a multimodal joint feature vector, and obtaining a multimodal joint feature vector; The feature vector of the multimodal data includes: text features, image features and audio features; Based on the multimodal joint feature vector, a fully connected neural network is constructed as a sentiment classification model; The cross entropy loss function is used to calculate the difference between the true emotion label and the emotion label predicted by the model. The calculation expression is: In the formula, M represents the total number of training samples, a represents the sample, b represents the emotion category, and C represents the total number of emotion categories. represents the loss function, represents the predicted probability of the bth emotion of the ath sample, y a,b Indicates the true emotion label of the bth emotion of the ath sample; Iteratively update model parameters to optimize the performance of the sentiment classification model; Based on the trained sentiment classification model, multimodal data is input to calculate the sentiment classification results and the sentiment accuracy coefficient. The calculation expression is: In the formula, Z C represents the sentiment accuracy coefficient, represents the matching function, y a represents the true sentiment label of the a-th sample, represents the predicted probability of sentiment of the ath sample; Classify short video emotions according to the emotion accuracy coefficient.
5. The short video script matching generation method combined with big data analysis according to claim 3 is characterized in that: The process of obtaining the topic accuracy coefficient is as follows: Obtain multimodal data of short videos and perform feature extraction on the multimodal data; The dot product of the feature vector of each mode is used to calculate the similarity of each mode. The calculation expression is: In the formula, S(X,Y) represents the similarity of each modality, E X represents the eigenvector in mode X, E Y represents the eigenvector in mode Y; Calculate the comprehensive similarity score between text, image and audio by weighted fusion of multimodal similarity scores; For each frame of the image and the corresponding audio clip in the short video, calculate the comprehensive similarity score of each frame; Based on the average of the multimodal similarity scores, the topic accuracy coefficient of the short video is calculated, and the calculation expression is: In the formula, T represents text features, I represents image features, A represents audio features, d represents each frame, H represents the total number of each frame, and Z D Represents the topic accuracy coefficient, S(T,I,A) d Represents the comprehensive similarity score between each frame and the corresponding audio segment.
6. The short video script matching generation method combined with big data analysis according to claim 1 is characterized in that: The process of obtaining the matching coefficient is as follows: Based on the text data of short videos, the text Transformer encoder in the VATT model is used to extract the emotional feature vector E of the text. t and the plot feature vector C t ; Based on each frame of the short video, the visual Transformer encoder in the VATT model is used to extract the emotional feature vector E of the image. v and the plot feature vector C v ; Based on the audio data of the short video, the spectrogram is generated by short-time Fourier transform, and the audio Transformer encoder in the VATT model is used to extract the audio's emotional feature vector E. a and the plot feature vector C a ; Based on text, image and audio modalities, the similarity between the emotional feature vector and the plot feature vector is calculated respectively. The calculation expression is: In the formula, i represents each mode, E i represents the emotional feature vector of modality i, C i represents the plot feature vector of modality i, which includes text, image and audio, ||E i || is the Euclidean norm of the sentiment feature vector, which is calculated by the Euclidean distance norm calculation expression, ||C i || is the Euclidean norm of the plot feature vector, which is calculated by the Euclidean distance norm calculation expression; The similarity results of each modality are weighted and fused to calculate the matching coefficient between emotion and plot. The calculation expression is: Where M score Represents the matching coefficient, W i Represents the preset weight of the i-th mode.
7. The short video script matching generation method combined with big data analysis according to claim 1 is characterized in that: If the sentiment analysis result conflicts with the plot design in the video, the generated script content is automatically adjusted to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy, specifically including: Extracting emotional features of the short video script content through sentiment analysis, and extracting plot features of the video plot design using the video plot feature generation model; According to the matching coefficient between emotional features and plot features, determine whether there is a conflict between the script content and the video plot design; Extract emotional keywords K from short video scripts through emotional feature vectors e ; Extracting plot keywords K from video plots through plot feature vectors c ; To judge the semantic conflict between emotional keywords and plot keywords, the cosine similarity calculation formula based on keyword embedding vector is: In the formula, Conflict(k e ,k c ) represents the semantic conflict coefficient. If the semantic conflict coefficient is greater than the preset threshold, there is a conflict between the emotional keyword and the plot keyword; According to the conflict analysis results, a new script content is generated that matches the video plot keywords, where the emotional keywords in the new script content meet the optimization constraints. The calculation expression is: Among them, Sim(K′ e ,K c ) represents the semantic similarity between the optimized emotional keywords and plot keywords, ensuring the consistency between the script content and the plot design, k′ e represents the optimized emotional keyword set, k c represents the set of plot keywords, k′ e ∈K′ e represents any emotional keyword in the emotional keyword set, k c ∈K c represents any plot keyword in the plot keyword set, max represents the maximum value, k′ e ·k c Represents the vector dot product of sentiment keywords and plot keywords, k′ e || represents the norm of the sentiment keyword vector, ||k c || represents the norm of the plot keyword vector; Regenerate the short video script based on the optimized script content, update the title, description and emotional tendency information in the script, and output the optimized short video script file.
8. A short video script matching generation system combined with big data analysis, characterized in that: A short video script matching generation method combined with big data analysis as described in any one of claims 1 to 7, comprising: A data acquisition module, wherein the data acquisition module obtains and pre-processes multimodal data of the short video, wherein the multimodal data of the short video includes text data, image data, and sound data; A multimodal data consistency assessment module, which is used to perform consistency analysis on the multimodal data and calculate a consistency coefficient, wherein the consistency analysis includes checking whether there are inconsistencies between text, image and sound data; and calculating a consistency coefficient based on the correlation between each modality, which is used to evaluate the consistency of the multimodal data content; An accuracy assessment module, wherein the accuracy assessment module marks the multimodal data as having consistent content or inconsistent content according to the assessment result, and performs an accuracy assessment on the emotional tendency and theme of the short video through sentiment analysis and theme classification on the basis of the consistency of the multimodal data content; A conflict analysis module, which extracts the accurate emotional tendency and theme of the short video according to the evaluation results, detects the conflict between the emotion and the plot, and determines whether the emotion and the plot conflict by calculating the matching coefficient between the emotion and the plot; A script optimization module, wherein, based on the judgment result, if the sentiment analysis result conflicts with the plot design in the video, the script optimization module automatically adjusts the generated script content to ensure the consistency of the sentiment and the plot, thereby optimizing the script generation strategy; A user feedback and real-time adjustment module performs real-time optimization and adjustment based on user feedback data, and dynamically optimizes the script generation process through real-time data collection, including likes, comments, shares, and viewing time.
Citation Information
Cited By
Self-media content streaming matching method and system based on dynamic semantic analysis
CN120578967A
Client mining method and system based on short video
CN121901453A