A multimodal video content effectiveness feedback visual analytics method and system
By employing a visual analysis method for the effectiveness feedback of multimodal video content, this method quantifies the multimodal data features in videos, calculates effectiveness factors based on domain requirements, generates feedback results, and recommends reference videos. This solves the problem of lacking comprehensive analysis in existing technologies and achieves efficient and visual guidance for video content improvement.
Patent Information
- Application Number
- CN202310976858.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-04
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2043-08-04
AI Technical Summary
Existing technologies lack comprehensive analysis of multimodal content in video content effectiveness analysis, making it difficult to provide targeted feedback and improvement suggestions, and relying on manual analysis consumes a lot of human resources.
A visual analysis method for the effectiveness feedback of multimodal video content is adopted. By quantifying the characteristics of multimodal data and combining them with domain requirements to determine effectiveness factors, calculate relevance, generate feedback results and recommend reference videos, and display the results in a visual form.
It enables multi-level, multi-modal effectiveness feedback analysis of video content, allowing users to quickly understand areas for improvement, providing reference videos, reducing labor costs, and improving analysis efficiency.
Smart Images

Figure CN116910302B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of information technology and visualization technology, and particularly relates to a multi-modal video content effectiveness feedback visual analysis method and system. BACKGROUND
[0002] In recent years, the amount of multimedia video has been growing rapidly. A video contains a large amount of information in various modalities such as images, audio, and text, which embodies the information transmission and thought expression intention of the video author. The expression manner of information in different modalities in the video content is closely related to the expression effect. Speech is a common expression form in life. The expression, action, and tone of the speaker during the speech will have an important effect on the communication of the speech content, and can make the audience have different understanding experience and resonance feeling, thereby achieving better speech effect. The speech process of the speaker is usually recorded as a video to record the speech results in different occasions such as practice and formal speech, and to support subsequent analysis and dissemination. Similar to other video authors, the speaker needs to know the specific feedback on a particular speech, and obtain suggestions for improvement and reference objects for learning.
[0003] Currently, in the analysis and feedback of video content effectiveness, manual analysis by trainers is usually relied on, which needs to rely on experience and consumes a large amount of human cost. In recent years, there are also works in related fields trying to conduct automatic analysis. Some speech training software supports extraction and analysis of voice data, but the analysis means is mostly basic and cannot comprehensively analyze various speech skills in the speech. Chinese patent application CN113743271A discloses a video content effectiveness visual analysis method and system based on multi-modal emotion, which mainly relies on multi-modal emotion information and fails to cover other aspects of video content. At the same time, it mainly explores the existing video database and analyzes the effectiveness rules, and cannot form the effectiveness feedback of a specific video, nor can it obtain the reference object for potential adjustment. SUMMARY
[0004] The present application aims to provide a multi-modal video content effectiveness feedback visual analysis method and system.
[0005] The video content effectiveness in the application refers to the correlation between the multi-modal content in the video and the expression effect thereof, and the effectiveness evaluation mode is determined in combination with the actual field, including but not limited to the relationship between the method of developing a speech and the speech performance in a speech video, the relationship between the method of teaching a course and the course effect in a teaching video, the relationship between the entertainment content display method and the audience experience in an entertainment video, and the like. Taking the speech video as an example, the speech skill in the speech is quantified in the application, the effectiveness analysis of the speech video is introduced, and the effectiveness feedback of a specific speech video and the speech context relationship are obtained for experts, beginners, judges and the like, so as to recommend other speech segments according to certain rules for user reference and analysis.
[0006] The technical solution adopted in the application is as follows:
[0007] A multi-modal video content effectiveness feedback visual analysis method, and the steps thereof include:
[0008] Collecting a certain type of video and the label of the objective index of effectiveness thereof;
[0009] Quantifying and extracting the multi-modal data features of the content concerned in the video;
[0010] On the basis of the extracted multi-modal data features, the effectiveness factors are determined in combination with the actual demand of the field, and the numerical values of the effectiveness factors of different contents are calculated;
[0011] The correlation between the effectiveness factors and the objective index of effectiveness is analyzed, and the correlation result of the effectiveness factors is obtained;
[0012] The correlation between the effectiveness factors and the objective index of effectiveness is utilized to extract the effectiveness feedback result of the video to be analyzed;
[0013] The recommended video result is generated in combination with the data of the video to be analyzed for user reference;
[0014] The effectiveness feedback result of the video to be analyzed and the multi-modal data context thereof are displayed in different visual forms for user hierarchical exploration of the effectiveness feedback result.
[0015] Further, the certain type of video includes a speech video, a teaching video, a sales video, an entertainment video and the like, and the label of the objective index of effectiveness includes a play count, a ranking, a score, a transaction volume and the like.
[0016] Further, the multi-modal data sources include a video, an image, a sound, a text and the like, and the multi-modal data features include facial expressions, body movements, eye gaze, positions, voice tones, rhythm pauses and the like of a person in the video, as well as backgrounds, color tones and background sounds of the video picture.
[0017] Further, the effectiveness factors are determined according to the actual needs of the field, including: establishing factors affecting the effectiveness of a specific field according to the theory and needs of the field corresponding to the specific type of video, which correspond to the skills, methods, etc. of the specific field, and have an influence on the performance effect in the specific field.
[0018] Further, the effectiveness factors include at least one of the following: emotional proportion, emotional average level, emotional change degree, emotional diversity, action amplitude, action diversity, eye range, eye change speed, position change amplitude, position change speed, tone change amplitude, rhythm speed, pause, background type, color tone light and dark, etc.
[0019] Further, the correlation between the effectiveness factors and the effectiveness objective indicators is analyzed, including: establishing the correlation between the effectiveness factors and the effectiveness objective indicators, such as analyzing the positive and negative correlation and the degree of correlation between the two.
[0020] Further, the correlation between the effectiveness factors and the effectiveness objective indicators is utilized to extract the effectiveness feedback results of the video to be analyzed, including: extracting the multi-modal data features of the video to be analyzed, calculating the effectiveness factor values, and predicting the effectiveness feedback results of the video to be analyzed according to the correlation between the effectiveness factors and the objective indicators.
[0021] Further, the recommended video results are generated in combination with the data of the video to be analyzed, wherein the data of the video to be analyzed includes but is not limited to multi-modal data features, effectiveness factor values and effectiveness feedback results, the recommendation method includes but is not limited to similarity retrieval from a video database, the recommendation basis includes but is not limited to the overall and local features of the video, and the recommendation object granularity includes but is not limited to the overall and segment of the video.
[0022] Further, the hierarchical exploration of the effectiveness feedback results supports the following functions from the overall to the local: effectiveness factor feedback function, video context understanding function, time interval distribution understanding function, data summary and similarity recommendation function.
[0023] A multi-modal video content effectiveness feedback visual analysis system, comprising:
[0024] A data collection module is responsible for collecting a certain type of video and its effectiveness objective indicators;
[0025] A data feature extraction module is responsible for quantitatively extracting multi-modal data features of the content concerned in the video;
[0026] An effectiveness factor calculation module is responsible for determining effectiveness factors based on multi-modal data features and combining actual needs of the field to calculate the values of different effectiveness factors.
[0027] An effectiveness analysis prediction module is responsible for establishing the correlation between the effectiveness factors and the objective indicators of effectiveness, and using the correlation results obtained by analysis to predict the effectiveness in the video to be analyzed.
[0028] A reference video recommendation module is responsible for recommending reference videos (the recommended videos can be similar or different from the video to be analyzed, which can be determined according to different needs) from the database according to the video to be analyzed, using specified video content and related parameters.
[0029] A visual analysis module is responsible for integrating the functions and data of the above modules to display the data and the results generated by the modules in different visual forms, and to present the results in a complete interface, so that the user can understand the effectiveness feedback results of the video to be analyzed through the interface, and support deeper exploration.
[0030] Through the visual analysis method and system of the present application, the user can understand the effectiveness feedback results of the specific video content to find the specific aspects that can be improved, understand the time distribution of the factor effectiveness to find the improvement position in the video content, understand the effectiveness factors in combination with the multi-modal context in the video for in-depth understanding of the video performance, obtain reference video instances for adjustment and improvement, and summarize the multi-modal effectiveness factors in the video for quick understanding and comparison.
[0031] Compared with the prior art, the present application has the following advantages and positive effects:
[0032] 1. The present application proposes a processing and analysis process of effectiveness feedback of multi-modal content in a video, and provides a full-process solution of visual analysis of video content effectiveness feedback. Compared with the prior art, the present application can better support the user to understand the effectiveness feedback results and conduct targeted exploration.
[0033] 2. The present application proposes an interactive visual analysis system for displaying, recommending, analyzing and exploring multi-modal content in the user's video, which allows the user to quickly understand the effectiveness feedback of different effectiveness factors in the video, supports the user to analyze in detail according to the context of the video, helps the user to quickly find reference video samples through recommendation, and supports targeted and detailed exploration of the interested video samples to support understanding of possible improvements of the video to be analyzed.
[0034] 3.The present application is based on the feedback of video content effectiveness, combined with a variety of visualization forms, proposes a visual analysis method and system based on multi-modal video content effectiveness feedback, which can be used to analyze the effectiveness feedback and possible improvement of the content expression in the video. With the help of visualization method, the video content effectiveness feedback is analyzed, and the effectiveness feedback, enhanced video content, effectiveness time slice and recommended video are displayed through the visualization system. It has advantages in forming video content effectiveness feedback and supporting users to form improved insights. Through intuitive and effective visualization methods and interactive ways, users can quickly understand and form in-depth insights. Therefore, multi-modal video content effectiveness feedback visual analysis is regarded as the main form of video analysis in the present application, and is not limited to specific fields and specific visualization methods. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 is the overall flow of the method of the present application and the layout of the multi-modal video content effectiveness feedback visual analysis system.
[0036] Figure 2 is the interface diagram of the multi-modal video content effectiveness feedback visual analysis system of an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to better understand the present application for those skilled in the art, the multi-modal emotion-based video content effectiveness visual analysis method and system provided by the present application are described in detail below in combination with the drawings, but do not constitute a limitation on the present application.
[0038] The present application mainly includes the following contents (wherein the presentation field is described, the present application can also be applied to teaching videos, entertainment videos and other video types):
[0039] 1.Multi-modal data acquisition and processing flow
[0040] The multi-modal data acquisition and processing flow mainly includes the following steps for specific fields: 1) data collection, 2) data feature extraction, 3) effectiveness factor calculation, 4) effectiveness analysis and prediction, 5) reference video recommendation, 6) visualization result generation. Multi-modal data includes three modalities of image, sound and text. As shown in the following, the presentation video is taken as an example for description. Figure 1
[0041] 1) Data collection: Through web crawlers, videos of the World Public Speaking Championship published on public platforms such as YouTube and related description information (i.e. the label of the effectiveness objective indicator) are crawled. The speeches are divided into different levels such as finals, semi-finals, large areas, medium areas, small areas and clubs, which are used as the measuring standard of the effectiveness of the speeches, that is, the higher the level of the competition, the higher the level of the speaker and the more effective the speech. In order to ensure the effectiveness of the correlation analysis, the number of speech videos of each level should be roughly equal. In addition to the level information, the speaker's name, region, speech topic, duration and other information are also collected, which are displayed in the visualization system as needed.
[0042] 2) Data feature extraction: In order to obtain the multi-modal emotional data in the video, it is necessary to extract the image frames, speech audio and speech text from the video, and all modalities are aligned with the text timestamp. The following introduces the feature extraction algorithm and tool used in the invention from different modalities:
[0043] a. Facial expression: Face localization and face recognition from image frames and clustering of faces using DBSCAN (Reference: M. Ester, H.-P. Kriegel, J. Sander, and X. Xu. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, KDD’96, p. 226-231. AAAI Press, 1996.) to find all the speaker’s face pictures appearing in the video. Then use AffectNet (Reference: A. Mollahosseini, B. Hasani, and M. H. Mahoor. Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Trans. Affect. Comput., 10(1):18-31, Jan. 2019. doi: 10.1109 / TAFFC. 2017.2740923) to extract continuous arousal and valence data from the face and use open-source methods on the web (Reference: O. Arriaga, M. Valdenegro-Toro, and P. Plöger. Real-time convolutional neural networks for emotion and gender classification.) for discrete emotion class recognition.
[0044] b. Gaze: The gaze direction of both eyes is estimated using the OpenFace toolkit (Reference: T. Baltrusaitis, A. Zadeh, Y. C. Lim, and L.-P. Morency. Openface 2.0: Facial behavior analysis toolkit. In 2018 13th IEEE International Conference on Automatic Face Gesture Recognition (FG 2018), pp. 59-66, 2018. and E. Wood, T. Baltruaitis, X. Zhang, Y. Sugano, P. Robinson, and A. Bulling. Rendering of eyes for eye-shape registration and gaze estimation. In 2015 IEEE International Conference on Computer Vision (ICCV), pp. 3756-3764, 2015.) The angle of the speaker looking at the camera is defined by the angle between the direction of the eye coordinates relative to the camera and the gaze direction.
[0045] c. Body Pose: The body pose skeleton is estimated by the MMPose toolkit (Reference: M. Contributors. Openmmlab pose estimation toolbox and benchmark.) and the skeleton of the speaker is filtered out by setting rules. Further, the skeleton energy (Reference: R. Niewiadomski, M. Mancini, and S. Piana. Human and virtual agent expressive gesture quality analysis and synthesis. Coverbal Synchrony in Human-Machine Interaction, pp. 269-292, 2013.) and the pose diversity in the interval are calculated. The latter is calculated by computing the cosine distance between all aligned and normalized pose skeletons and the pose of the first frame in the interval, and then computing the standard deviation of the distance matrix.
[0046] d. Stage usage: The distance to the camera is estimated by the OpenFace toolkit after estimating the speaker's head position. If it is an online speech, the speaker's position is defined by the center of the bounding box on the picture; if it is an offline speech, the speaker's position is obtained according to the actual position of the speaker calculated by the camera.
[0047] e. Volume pitch: The loudness of the speaker's speech is calculated as the volume value and the frequency as the pitch value by the Praat toolbox (Reference: P. Boersma and D. Weenink. Praat: doing phonetics by computer [Computer program]. Version 6.1.38, retrieved 2 January 2021.).
[0048] f. Speech rate pause: The pause includes the interval time between words and sentences, and the duration of each word syllable is calculated to obtain the speech rate value. The estimation of word syllables can be completed by the NLTK language toolkit (Reference: S. Bird, E. Klein, and E. Loper. Natural language processing with Python: analyzing text with the natural language toolkit. “O'Reilly Media, Inc.", 2009.).
[0049] g. Text content: The audio part in the video is converted into text using the audio-to-text service provided by Microsoft Azure, and the text and sentences and corresponding timestamps can be obtained.
[0050] 3) Effectiveness factor calculation: The features extracted in the previous step are time-varying features and cannot directly reveal the data trend and the direct influence factors of the speech effectiveness. Based on this, the present application extracts different effectiveness factors based on the multi-modal feature data, and the calculation method and corresponding multi-modal features are as follows:
[0051] Diversity: For emotional category feature representation, the emotional categories contained in a specific speech video and the relative proportion are calculated by , where represents the number of emotional categories, represents the proportion of a certain emotional category.
[0052] Average: represents the average value of a multimodal feature, which is calculated by the mean of the time series. It is suitable for facial continuous affect data (arousal and valence), skeleton energy of body posture, volume, pitch, speech rate, pause, etc.
[0053] Volatility: represents the degree of time series change of a multimodal feature, which is calculated by the complexity of the CID algorithm (reference: G. E. Batista, E. J. Keogh, O. M. Tataw, and V. de Souza. Cid: an efficient complexity-invariant distance for time series. Data Mining and Knowledge Discovery, 28(3):634-669, 2014.). It is suitable for facial continuous affect data (arousal and valence), gaze direction, distance from the camera, speaker position, skeleton energy of body posture, volume, pitch, speech rate, pause, etc.
[0054] Dispersion: represents the amplitude of change of a multimodal feature, which is calculated by the coefficient of variation obtained by dividing the standard deviation of the time series by the average value. It is suitable for gaze direction, distance from the camera, speaker position, etc.
[0055] Ratio: represents the proportion of a certain state, which is obtained by calculating the ratio of the state to all states. It is suitable for facial discrete affect types, whether the gaze is directed at the camera lens, etc.
[0056] 4) Effectiveness analysis prediction: In order to calculate the correlation between effectiveness factors and speech effectiveness, the collected videos are labeled with the level of the competition they belong to (final, semi-final, large area, middle area, small area, club), and they are marked as 6, 5, 4, 3, 2, 1, respectively. Such labels can be regarded as ordinal variables, that is, there is a certain order relationship between discrete labels. For such problems, the invention uses the method of multiclass ordinal regression (reference: P. A. Gutiérrez, M. Perez-Ortiz, J. Sanchez-Monedero, F. Fernandez-Navarro, and C. Hervas-Martinez. Ordinal regression methods: survey and experimental study. IEEE Transactions on Knowledge and Data Engineering, 28(1):127-146, 2015.) for analysis and processing, which can obtain the p value between each effectiveness factor and the level label, where p represents the probability of the hypothesis in hypothesis testing, p <0.05 is significant, p <0.01 is very significant, and is used as the importance of the effectiveness factor. For the video to be analyzed, based on the effectiveness correlation calculated based on the existing data set, the effectiveness feedback result of the video to be analyzed can be predicted.
[0057] 5) Reference video recommendation: According to the speech video to be analyzed and the data and related parameters in the speech video database, the recommended video result is generated.
[0058] 6) Visualization result generation: Combine the data and analysis results generated by the above process, select the appropriate form to generate the visualization result according to the data characteristics and actual needs.
[0059] Through the above process, multi-modal data can be automatically obtained through video input, the relationship between multi-modal effectiveness factors and speech effectiveness can be mined, and reference videos can be recommended to provide data support for visual analysis methods and systems.
[0060] 2. Multi-modal video content effectiveness feedback visual analysis system
[0061] As Figure 1The right visual analytics module is shown, according to the reading habit from left to right, from top to bottom, the system interface is divided into four functions: A. Effectiveness factor feedback (presentation factor panel), B. Video context understanding (speaker panel), C. Time interval distribution understanding (time period slice panel), D. Data summary and similarity recommendation (mirror panel). The four main functions can be used together to help users explore the effectiveness feedback of the video to be analyzed and find the possibility of improvement.
[0062] A. The effectiveness factor feedback function shows the effectiveness feedback of the video to be analyzed in the form of a chart, which shows the effectiveness rule diagram based on the data set and the distribution of the video to be analyzed in the data set. This function also includes the function of selecting one or more effectiveness factors, and the result of the selection will affect the results of other functions, and specific effectiveness factors can be explored and analyzed.
[0063] B. The video context understanding function provides a multi-modal data situation based on the video player to understand the video context content, which can enhance the understanding of multi-modal effectiveness feedback in the video while watching the video, and further provide data views through interaction to support in-depth exploration by users.
[0064] C. The time interval distribution understanding function provides video effectiveness feedback distribution display and video interval selection function, and visualizes the multi-modal effectiveness factor distribution, multi-modal data and text content in the selected interval in chronological order. It can support users to explore and analyze more finely according to the selected effectiveness factors and corresponding time intervals.
[0065] D. The data summary and similarity recommendation function provides a multi-modal data summary of the selected segment in C, and according to the user's selected needs, uses the reference video recommendation module to form a recommendation result and displays the result on the system interface to assist users to understand the video objects that can be used for reference.
[0066] In this part, the focus of the invention is on the arrangement of functions and the capabilities that should be provided, and no specific visualization form is limited. Any visualization form that can assist users in analyzing the effectiveness of the speech can be included in the system.
[0067] 3. Video multi-level exploration method with effectiveness feedback as the core
[0068] Just showing data is far from enough, the system proposed in 2 provides a video multi-level exploration method with effectiveness feedback as the core. Figure 2 The system interface diagram of an embodiment of the invention, where A, B, C, D are functions A~D described below.
[0069] Function A provides the effectiveness factor feedback results of the video to be analyzed, as well as the effectiveness rules and data distribution display. The feedback results (the color bar shown by A1 maps the effectiveness feedback results) and the distribution (the panel shown by A2 displays further results) of different effectiveness factors can be intuitively understood in function A. In A, different effectiveness factors can be clicked, and functions C and D will change accordingly.
[0070] Function B forms an understanding of the video context in the form of a video player. The user can understand the effectiveness factors in the video content while watching the video, thereby enhancing the understanding of the effectiveness factors and the correlation between the effectiveness factors and the video context. In this function, the key content in the video can be highlighted (B2-B5). The highlighting can be achieved by superimposing visualizations on the video or by superimposing interactive functions on the video. The data view (B1) can be triggered under the condition of mouse hovering, helping the user to have a deeper understanding of the multi-modal data and its effectiveness.
[0071] Function C displays the distribution of the video content effectiveness feedback with the video time, as well as the multi-modal data display and effectiveness factor display in the selected interval. This function maps the effectiveness feedback results on the video progress bar (C1), supports the user to select the speech interval for detailed exploration, and divides the time axis into multiple parts to display the changes of the effectiveness factors in each time slice interval and the corresponding multi-modal data (C2). The text under the mapping of the effectiveness feedback results (C3) facilitates the user to intuitively understand the text content in each interval and the corresponding multi-modal effectiveness feedback. Function C helps the user to have a deeper understanding of the speech effectiveness and multi-modal data.
[0072] Function D displays the multi-modal summary of the selected video interval and displays the reference video recommendation results set by the user, so as to help the user quickly understand and compare the speech situation and find cases for learning. The multi-modal video summary displays important data features to help the user quickly understand the video content (D1). The user can configure the reference video recommendation options (D3), and then obtain the reference video recommendation results. The recommendation results can also be displayed in the form of a multi-modal video summary, which facilitates the user to find suitable reference sources (D2). The effectiveness factor comparison panel (D4) can be triggered under the condition of mouse hovering, which facilitates the understanding of the differences between the reference video and the video to be analyzed. Clicking a recommendation result can focus on the video, and the data of the video can be displayed in detail in functions B and C.
[0073] Based on the same inventive concept, another embodiment of the present application provides a multi-modal video content effectiveness feedback visual analysis system, characterized in that it comprises:
[0074] a data collection module, responsible for collecting a certain type of video and its effectiveness objective index label;
[0075] a feature extraction module, responsible for collecting emotional data from images, texts, sounds and other modalities in the video, quantifying and extracting multi-modal data features of the content of interest in the video, including facial expressions, body movements, eye gaze, position, background, voice tone, speech speed, background sound, text content and other data features;
[0076] an effectiveness factor calculation module, responsible for determining effectiveness factors based on multi-modal data features and combining actual needs of the field to calculate the values of different effectiveness factors;
[0077] an effectiveness analysis and prediction module, responsible for establishing the correlation between the effectiveness factors and the effectiveness objective index, and using the correlation results obtained by analysis to predict the effectiveness of the video to be analyzed;
[0078] a reference video recommendation module, responsible for recommending reference videos similar to the video to be analyzed from the database according to the specified video content and related parameters;
[0079] a visual analysis module, responsible for integrating the functions and data of the above modules, displaying the results generated by the data and modules in different visual forms, and presenting the results in a complete interface, so that users can understand the effectiveness feedback of the specific video to be analyzed through the interface, and support deeper exploration.
[0080] The specific implementation process of each module is described in the foregoing method of the present application.
[0081] Based on the same inventive concept, another embodiment of the present application provides an electronic device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present application.
[0082] Based on the same inventive concept, another embodiment of the present application provides a computer readable storage medium (such as ROM / RAM, magnetic disk, optical disk), which stores a computer program, and the computer program is executed by a computer to realize each step of the method of the present application.
[0083] The multi-modal video content effectiveness feedback visual analysis method and system of the present application are described in detail above, but obviously the specific implementation form of the present application is not limited to this. For those skilled in the art, various obvious changes made to it without departing from the spirit of the method of the present application and the scope of the claims are within the scope of protection of the present application.
Claims
1. A visual analysis method for multimodal video content validity feedback, characterized in that, Includes the following steps: Collect tags for the videos to be analyzed and their objective effectiveness indicators; Quantitatively extract multimodal data features of content of interest in a video; the multimodal data includes video, images, audio, and text; Based on the extracted multimodal data features, effectiveness factors are determined in combination with the actual needs of the domain, and the effectiveness factor values for different content are calculated. The correlation between effectiveness factors and objective effectiveness indicators was analyzed to obtain the correlation results of effectiveness factors; By leveraging the correlation between effectiveness factors and objective effectiveness indicators, effectiveness feedback results are extracted from the videos to be analyzed. The method of extracting effectiveness feedback results from the video to be analyzed by utilizing the correlation between effectiveness factors and objective effectiveness indicators includes: extracting multimodal data features of the video to be analyzed, calculating the values of effectiveness factors, and predicting the effectiveness feedback results of the video to be analyzed according to the correlation between effectiveness factors and objective indicators; the effectiveness feedback results are used to discover specific aspects of the video to be analyzed that can be improved, understand the temporal distribution of factor effectiveness, obtain reference video examples, and summarize the multimodal effectiveness factors in the video. The system combines data from the videos to be analyzed to generate recommended videos for user reference; the data from the videos to be analyzed includes multimodal data features, validity factor values, and validity feedback results. The validity feedback results of the video to be analyzed and its multimodal data context are displayed in different visualization forms, allowing users to explore the validity feedback results in a hierarchical manner.
2. The method according to claim 1, characterized in that, The videos to be analyzed include one of the following: speech videos, teaching videos, sales videos, and entertainment videos. The tags for the objective effectiveness indicators include play count, ranking, score, and sales volume.
3. The method according to claim 1, characterized in that, The multimodal data features include facial expressions, body movements, eye contact, location, tone of voice, rhythmic pauses, background, color tone, and background sound of the people in the video.
4. The method according to claim 1, characterized in that, The determination of effectiveness factors based on the actual needs of the field includes: establishing factors that affect the effectiveness of the corresponding field based on the theories and needs of the field to be analyzed. These factors correspond to the skills and methods in that field and have an impact on the performance effect in that field. The effectiveness factors include at least one of the following: emotional proportion, average emotional level, degree of emotional change, emotional diversity, range of motion, variety of motion, eye contact range, speed of eye contact change, range of position change, speed of position change, range of tone change, tempo, number of pauses, background type, and color tone.
5. The method according to claim 1, characterized in that, The analysis of the correlation between effectiveness factors and objective effectiveness indicators is to establish the relationship between effectiveness factors and objective effectiveness indicators, including analyzing the positive and negative correlations and the degree of correlation between the two.
6. The method according to claim 1, characterized in that, The method of generating recommended video results by combining the data of the video to be analyzed includes similarity retrieval from the video database, and the recommendation criteria include the overall and local features of the video. The granularity of the recommended objects includes the whole video and segments.
7. The method according to claim 1, characterized in that, The hierarchical exploration of the effectiveness feedback results supports the following functions of joint analysis and expression from the whole to the part: effectiveness factor feedback function, video context understanding function, time interval distribution understanding function, data summarization and similarity recommendation function.
8. A multimodal video content validity feedback visual analysis system employing the method described in any one of claims 1 to 7, characterized in that, include: The data collection module is responsible for collecting tags for the videos to be analyzed and their objective effectiveness indicators; The data feature extraction module is responsible for quantifying and extracting multimodal data features of the content of interest in the video; The effectiveness factor calculation module is responsible for determining effectiveness factors based on multimodal data characteristics and actual domain needs, and calculating the values of different effectiveness factors. The effectiveness analysis and prediction module is responsible for establishing the correlation between effectiveness factors and objective effectiveness indicators, and using the correlation results obtained from the analysis to predict the effectiveness in the video to be analyzed. The reference video recommendation module is responsible for recommending reference videos from the database based on the video to be analyzed, using specified video content and related parameters; The visual analytics module is responsible for integrating the functions and data of the above modules, displaying the data and the results generated by each module in different visualization forms, and presenting them in a complete interface so that users can understand the effectiveness feedback results of the video to be analyzed and support deeper exploration.
9. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video content validity visual analysis method and system based on multi-modal emotion
CN113743271A
Lecture evaluation system and method
JP6993745B1
Video recommendation system and method thereof
US20130259399A1