Video interaction method and related device
By analyzing video multi-modal emotional data and collecting user multi-dimensional emotional data, personalized emotional responses and adjusting video content, the problem that existing technology cannot be intelligently adjusted is solved and efficient emotional interaction is achieved.
Patent Information
- Application Number
- CN202510251821.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video interaction technology cannot intelligently adjust according to user mood fluctuations, resulting in the inadequate satisfaction of user emotional interaction needs.
By conducting a comprehensive analysis of the multi-modal emotional data of the video, the emotional fluctuation points and areas in the video are determined. When the video is played to these areas, the user's multi-dimensional emotional data is obtained, personalized emotional response is generated, and the video content is adjusted according to the user's emotional fluctuation trend.
Realize instant and diverse emotional feedback, enhance the real-time and diversity of emotional interactions, and meet the user's emotional interaction needs.
Smart Images

Figure CN120075493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video interaction technology, and more particularly, to a video interaction method and related device. Background Art
[0002] With the rapid development of video platforms, users' demand for the viewing experience of videos has gradually shifted from simple video playback to more interactive and personalized emotional experiences. Most of the existing emotional interactions of videos rely on traditional user feedback forms, such as comments, bullet screens, etc., and cannot be intelligently adjusted according to users' emotional fluctuations, resulting in the failure to fully meet users' emotional interaction needs. Summary of the Invention
[0003] In view of this, the present invention discloses a video interaction method and related device to enable users to generate personalized emotional responses on the video based on multi-dimensional emotional data when watching the video during video interaction, and to adjust the video content according to the trend of users' emotional fluctuations, so as to achieve instant and diverse emotional feedback, enhance the real-time and diversity of emotional interaction, and thus meet users' emotional interaction needs.
[0004] A video interaction method includes:
[0005] Comprehensively analyzing the video based on video multi-modal emotional data to determine the emotional fluctuation points in the video and determine the emotional fluctuation region including the emotional fluctuation points;
[0006] When the video plays to the emotional fluctuation region, obtaining multi-dimensional emotional data of the user when watching the video, and generating a personalized emotional response on the video based on the multi-dimensional emotional data;
[0007] Predicting the trend of users' emotional fluctuations based on the multi-dimensional emotional data, and adjusting the video content based on the trend of users' emotional fluctuations.
[0008] Optionally, the comprehensively analyzing the video based on video multi-modal emotional data to determine the emotional fluctuation points in the video and determine the emotional fluctuation region including the emotional fluctuation points includes:
[0009] Performing sentiment classification on the text in the video, identifying the sentiment words in the text, and determining the text emotion value related to emotional fluctuations based on the sentiment words;
[0010] Performing visual emotion analysis on the video based on computer vision technology to obtain the visual emotion value related to emotional fluctuations;
[0011] Analyzing the audio data in the video based on audio signal processing technology to obtain the audio emotion value related to emotional fluctuations;
[0012] Based on the text emotion value, visual emotion value, and audio emotion value at each time point of the video, determine the comprehensive video emotion value at each time point of the video, mark the position where each comprehensive video emotion value appears in the video as an emotion fluctuation point, and determine the emotion fluctuation region containing the emotion fluctuation point.
[0013] Optionally, performing sentiment classification on the text in the video, identifying sentiment words in the text, and determining the text emotion value related to emotion fluctuation based on the sentiment words, including:
[0014] Extract all the text information from the video to obtain the text;
[0015] Perform sentiment analysis on the text to determine the sentiment words in the text;
[0016] Determine the emotion score and weight of each sentiment word;
[0017] Based on the emotion scores and weights corresponding to each sentiment word, obtain the text emotion value.
[0018] Optionally, performing visual emotion analysis on the video based on computer vision technology to obtain the visual emotion value related to emotion fluctuation, including:
[0019] Based on face recognition technology, identify the people and their expressions in the video, and use the facial expression recognition algorithm to determine the emotions of the people;
[0020] Use a deep learning model to analyze the shot transitions and scene changes of the video, and detect the visual signals of emotion fluctuations;
[0021] Take the people, their expressions, their emotions, and the visual signals as visual features, and determine the emotion score and weight corresponding to each visual feature;
[0022] Based on the emotion scores and weights corresponding to each visual feature, obtain the visual emotion value.
[0023] Optionally, performing analysis on the audio data in the video based on audio signal processing technology to obtain the audio emotion value related to emotion fluctuation, including:
[0024] Extract audio features from the audio data of the video based on audio signal processing technology;
[0025] Determine the emotion score and weight corresponding to each audio feature;
[0026] Based on the emotion scores and weights corresponding to each audio feature, obtain the audio emotion value.
[0027] Optionally, when the video plays to the emotional fluctuation area, obtaining multi-dimensional emotional data of the user while watching the video, and generating a personalized emotional response on the video based on the multi-dimensional emotional data, including:
[0028] When the video plays to the emotional fluctuation area, obtaining multi-dimensional emotional data of the user while watching the video;
[0029] Performing emotional data fusion on the multi-dimensional emotional data, and generating the personalized emotional response on the video according to the result of the emotional data fusion.
[0030] Optionally, predicting the user's emotional fluctuation trend based on the multi-dimensional emotional data, and adjusting the video content based on the user's emotional fluctuation trend, including:
[0031] Obtaining video content feature data and user interaction behavior data, and inputting the multi-dimensional emotional data, the video content feature data, and the user interaction behavior data into a set emotional fluctuation prediction model to predict the user's emotional fluctuation trend;
[0032] Adapting and adjusting the video content based on the user's emotional fluctuation trend.
[0033] Optionally, further including:
[0034] Extracting user emotional features from the multi-dimensional emotional data;
[0035] Determining the user's emotional type based on the user's emotional features;
[0036] Recommending emotional interaction methods and / or videos with a high matching degree to the user's emotional type.
[0037] Optionally, further including:
[0038] Providing a personalized emotional interaction method for the user based on the user's emotional fluctuation trend.
[0039] A video interaction device, including:
[0040] A determination unit, configured to perform comprehensive analysis on the video based on video multi-modal emotional data, determine the emotional fluctuation points in the video, and determine the emotional fluctuation area including the emotional fluctuation points;
[0041] An acquisition unit, configured to, when the video plays to the emotional fluctuation area, obtain multi-dimensional emotional data of the user while watching the video, and generate a personalized emotional response on the video based on the multi-dimensional emotional data;
[0042] An adjustment unit for predicting the user's emotional fluctuation trend based on the multi-dimensional emotional data and adjusting the video content based on the user's emotional fluctuation trend.
[0043] A computer storage medium storing at least one instruction, which when executed by a processor implements the above-mentioned video interaction method.
[0044] An electronic device, comprising: a memory and a processor;
[0045] The memory is used for storing at least one instruction;
[0046] The processor is used for executing the at least one instruction to implement the above-mentioned video interaction method.
[0047] As can be seen from the above technical solutions, the present invention discloses a video interaction method and related devices, which comprehensively analyze a video based on video multi-modal emotional data, determine the emotional fluctuation points in the video, and determine the emotional fluctuation area including the emotional fluctuation points. When the video plays to the emotional fluctuation area, multi-dimensional emotional data of the user when watching the video is obtained, and a personalized emotional response is generated on the video based on the multi-dimensional emotional data. The user's emotional fluctuation trend is predicted based on the multi-dimensional emotional data, and the video content is adjusted based on the user's emotional fluctuation trend. Through the collection and analysis of video multi-modal emotional data in this application, when the user conducts video interaction, a personalized emotional response can be generated on the video according to the multi-dimensional emotional data of the user when watching the video, and the video content can be adjusted according to the user's emotional fluctuation trend, realizing instant and diversified emotional feedback, enhancing the real-time and diversity of emotional interaction, and thus meeting the user's emotional interaction needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the disclosed drawings without creative efforts.
[0049] Figure 1 It is a flowchart of a video interaction method disclosed in an embodiment of the present invention;
[0050] Figure 2 It is a schematic structural diagram of a video interaction device disclosed in an embodiment of the present invention;
[0051] Figure 3 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Most of the existing emotional interactions in videos rely on traditional user feedback forms such as comments and bullet screens. Although these methods can reflect users' immediate emotions, there are also certain limitations in actual use. For example:
[0053] (1) Limitations of bullet screen interaction. As the most common emotional expression method, although bullet screens can allow users to express their emotions immediately during the viewing process, the frequent appearance of bullet screens often interferes with the user's viewing experience. Especially at moments when the content is more intense or emotional, the bullet screens may interrupt the audience's immersion.
[0054] (2) Delayed emotional feedback. Most of the existing video interaction methods rely on explicit input from users, such as the bullet screens or comments sent by users. These methods often have a certain delay and cannot timely feedback the emotional fluctuations of users.
[0055] (3) High interaction threshold. When watching videos, users hope to express their emotions in a more direct and convenient way. However, the existing interaction means usually rely on expressions such as text and bullet screens, which may pose an operation threshold for some users and it is difficult to quickly establish an emotional connection with the content.
[0056] From the above limitations, it can be seen that the existing methods cannot make intelligent adjustments according to users' emotional fluctuations, resulting in the failure to fully meet users' emotional interaction needs.
[0057] Based on this, the embodiments of the present invention disclose a video interaction method and related device. By collecting and analyzing multi-modal emotional data of videos, when users perform video interaction, personalized emotional responses can be generated on the video according to the multi-dimensional emotional data of users when watching videos, and the video content can be adjusted according to the trend of users' emotional fluctuations, realizing immediate and diverse emotional feedback, enhancing the real-time and diversity of emotional interaction, and thus meeting users' emotional interaction needs.
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] See Figure 1 , a flowchart of a video interaction method disclosed in the embodiments of the present application. The method includes:
[0060] Step S101: Comprehensively analyze the video based on the video multi-modal emotion data, determine the emotion fluctuation points in the video, and determine the emotion fluctuation regions containing the emotion fluctuation points.
[0061] In practical applications, multi-modal information (such as text, vision, audio) can be used to comprehensively analyze the video, predict the possible emotion fluctuation positions when the user watches the video, mark the emotion fluctuation points in the video in advance, and identify potential emotion fluctuation regions (such as excitement, depression, tension, joy, etc.), providing data support for subsequent emotion interactions.
[0062] For example, a sudden horror scene in a video may trigger the user's frightened emotion, and the system can identify this potential emotion fluctuation region in advance through emotion analysis. At this time, an emotion feedback interaction module can be provided in the potential emotion fluctuation region, or the maximization of the emotion value can be triggered at an appropriate time.
[0063] Step S102: When the video plays to the emotion fluctuation region, obtain the multi-dimensional emotion data of the user when watching the video, and generate a personalized emotion response on the video based on the multi-dimensional emotion data.
[0064] Among them, the multi-dimensional emotion data includes but is not limited to: text emotion value, voice emotion value, and physiological emotion value.
[0065] The personalized emotion response includes but is not limited to: personalized emotion interaction methods, personalized content recommendations, etc.
[0066] For example, when the user's emotion reaches a climax, the system can guide the user to participate in a more interactive emotional expression session, such as posting humorous or touching bullet screens.
[0067] For example, personalized content recommendations can recommend interactive content that matches the user's emotional state to the user.
[0068] When the video plays to the emotion fluctuation region, the emotion feedback interaction module is used to collect and process the user's emotion feedback in real time during the video viewing process, capture the user's emotion changes, so as to respond to the user's emotion expression in a timely manner, thereby optimizing the user experience and increasing interactivity. The multi-dimensional emotion data is not limited to text, but also includes various data such as voice, facial expressions, gestures, heart rate, etc. By comprehensively analyzing these feedback information, the system can generate personalized emotion responses or interaction behaviors. This enables the user's emotion expression to not only be real-time feedback to the platform, but also be dynamically processed by the platform system, thereby realizing personalized emotion regulation and enhancing the viewing experience.
[0069] Step S103: Predict the user's emotion fluctuation trend based on the multi-dimensional emotion data, and adjust the video content based on the user's emotion fluctuation trend.
[0070] During the process of predicting the user's emotional fluctuation trend and adjusting the video content, based on the multi-dimensional emotional data of the user collected in real time, the system predicts the user's emotional fluctuation trend and adjusts the video content accordingly to enhance the user's immersion and viewing experience. Through accurate prediction of emotional fluctuations, the system can identify the possible emotional ups and downs of the user in advance and dynamically adjust the video content (such as adjusting the picture rhythm, sound effects, plot settings, etc.) to match the user's emotional needs, achieving emotional resonance with the user.
[0071] In practical applications, adjusting the video content based on the user's emotional fluctuation trend may include: adjusting the emotion-driven content, multi-modal content, and interactive content in the video.
[0072] Adjusting the video content based on the user's emotional fluctuation trend may also include: adjusting the video recommendation content. When the user's mood is high, the system recommends more intense and exciting content of a similar nature; when the user's mood is low, the system recommends relaxing and pleasant content to ensure that the user's emotions are properly guided and regulated.
[0073] In summary, this application discloses a video interaction method, which comprehensively analyzes the video based on the video multi-modal emotional data, determines the emotional fluctuation points in the video, and determines the emotional fluctuation area including the emotional fluctuation points. When the video plays to the emotional fluctuation area, it obtains the multi-dimensional emotional data of the user when watching the video, generates a personalized emotional response on the video based on the multi-dimensional emotional data, predicts the user's emotional fluctuation trend based on the multi-dimensional emotional data, and adjusts the video content based on the user's emotional fluctuation trend. By collecting and analyzing the video multi-modal emotional data, this application enables the user to generate a personalized emotional response on the video according to the multi-dimensional emotional data when watching the video during video interaction, and can adjust the video content according to the user's emotional fluctuation trend, realizing instant and diverse emotional feedback, enhancing the real-time and diversity of emotional interaction, and thus meeting the user's emotional interaction needs.
[0074] In one embodiment, step S101 may specifically include:
[0075] (1) Classify the emotions of the text in the video, identify the emotional words in the text, and determine the text emotion value related to emotional fluctuations based on the emotional words.
[0076] Among them, the text in the video includes: dialogues, subtitles, voiceovers, and other texts.
[0077] Use the natural language processing (NLP) emotion analysis method to classify the emotions of the text in the video and identify the emotional words in the text.
[0078] Specifically, all text information (such as subtitles, character dialogues, etc.) is extracted from the video to obtain the text; sentiment analysis is performed on the text to determine the sentiment words in the text, and the emotion scores and weights of each sentiment word are determined; based on the emotion scores and weights corresponding to each of the sentiment words, the text emotion value is obtained.
[0079] Among them, the process of determining the emotion scores and weights of the sentiment words in each text is as follows:
[0080] (1) Determination of the emotion scores of sentiment words
[0081] The score of each sentiment word is initially determined by its sentiment tendency in the sentiment dictionary, usually including sentiment directions such as positive, negative, and neutral. The emotion scores of these sentiment words in the text are based on the output results of the sentiment analysis model, and the model will analyze the sentiment polarity (positive / negative) and intensity (for example, a score from -1 to +1) of the words. For example, the emotion score of "happy" may be +0.8, and the emotion score of "sad" is -0.7.
[0082] (2) Determination of the weights of sentiment words
[0083] The weight of a sentiment word is calculated according to its context and importance in the specific context. Here, the present application adopts a dynamic weight adjustment mechanism based on the context. Specifically
[0084] Context sensitivity: The same sentiment word may have different weights in different situations. For example, "victory" may be a positive sentiment word when describing a game, but in a negative situation, it may carry a sarcastic or other emotional color. Therefore, the present application will dynamically adjust the weight of this sentiment word according to the context.
[0085] Sentence structure and sentiment intensity: Through syntactic analysis and grammar rules, the present application will determine the role of the sentiment word in the sentence. For example, the weight of the word "happy" in "very happy" will be higher than the ordinary weight of the word "happy", and the sentiment intensity of the former is stronger, so the score will also increase accordingly.
[0086] The present application uses machine learning and deep learning models (such as LSTM, BERT, etc.) to analyze the weights of sentiment words in sentences, and appropriately adjusts the weights based on the analysis results of the context.
[0087] Rather than statically assigning a fixed emotion score to each emotional word, this application dynamically adjusts the scores and weights of emotional words based on user feedback, text changes, and the system's learning process. This enables the system of this application to adapt to the emotional responses of various different situations and user groups, thereby improving the accuracy and real-time performance of emotion analysis.
[0088] In practical applications, emotional words in the text can be determined through an emotion dictionary or an emotion analysis model, such as BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), etc., such as happy, sad, angry, fearful, etc., and a corresponding emotion score and weight are attached to each emotional word. Among them, based on the weight of the emotional word and the emotion speculation in the context, the emotion fluctuation points in the video content can be marked.
[0089] Among them, the expression of the text emotion value is as follows:
[0090] (1);
[0091] In the formula, represents the text emotion value of the video at time point t, represents the weight of the i-th emotional word, represents the emotion score of the i-th emotional word, and N represents the number of emotional words.
[0092] (2) Perform visual emotion analysis on the video based on computer vision technology to obtain a visual emotion value related to emotion fluctuation.
[0093] Among them, based on computer vision technology, the facial expressions of people in the video, scene changes, color tones, camera angles, etc. can be analyzed to obtain a visual emotion value related to emotion fluctuation.
[0094] Specifically, identify the people and their expressions in the video based on face recognition technology, and use a facial expression recognition algorithm to determine the emotions of the people, such as smiling, frowning, etc.; use a deep learning model to analyze the shot transitions and scene changes of the video to detect visual signals of emotion fluctuation, for example, tense background music and rapid shot transitions may trigger anxiety emotions; use the people, their expressions, their emotions, and the visual signals as visual features, and determine the emotion score and weight corresponding to each visual feature; based on the emotion scores and weights corresponding to each visual feature, obtain a visual emotion value.
[0095] The determination of the emotion scores and weights corresponding to each visual feature is based on multi-dimensional analysis and machine learning algorithms, including the following main factors:
[0096] (1) Emotion scores of facial expressions
[0097] Each expression corresponds to an emotional state, and the system assigns an emotion score to each expression through a facial expression recognition algorithm. For example, "smile" corresponds to an emotion score of +0.8 (indicating a strong happy emotion), "frown" corresponds to an emotion score of -0.7 (indicating a strong confused or unhappy emotion), and "anger" corresponds to an emotion score of -1.0 (indicating an angry emotion).
[0098] These scores are not fixed values, but are set based on an emotion dictionary and specific emotion scoring results obtained during model training, and will be fine-tuned according to the context in actual applications.
[0099] (2) Emotion scores of shot and scene changes
[0100] The system analyzes factors such as the rhythm of scene transitions and the emotional characteristics of music. For example, when the camera switches quickly, the system assigns a negative emotion score based on the resulting sense of tension. Similarly, the rhythm and style of background music also affect the emotion score. For example, tense background music adds a negative value (e.g., -0.8) to the emotion score, while relaxing background music may make the emotion score positive (e.g., +0.5).
[0101] (3) Comprehensive weights of emotion scores
[0102] The emotion scores and weights of each visual feature (including expressions, shots, scenes, background music, etc.) are dynamically calculated based on their impact on the overall emotion. When multiple visual signals are superimposed, the system calculates the "importance weight" of each feature. This step is based on deep learning algorithms (such as convolutional neural networks) to learn and dynamically optimize the feature weights, taking into account the importance of features in a specific scene. For example, in a climax part of a plot, the facial expressions of the characters may have a greater impact on the emotion, while the rapid switching of shots may have a smaller impact on the emotional fluctuations. In this case, the present application assigns a higher weight to the facial expressions of the characters.
[0103] Dynamic adjustment of emotion scores and weights
[0104] The system of this application does not simply assign fixed emotion scores and weights to visual features, but adopts a real-time learning and context-aware mechanism. Each time a user watches a video, the system adjusts the emotion scores and weights of visual features according to the user's real-time feedback, emotional fluctuations and other information. For example, among different user groups, certain emotional expressions may be assigned different weights (for example, users with different cultural backgrounds may have different emotional responses to certain expressions).
[0105] By continuously accumulating the user's feedback data and emotional responses, the system dynamically adjusts the emotion scores and weights of visual features to optimize the accuracy of the overall emotional response.
[0106] In practical applications, emotional classification can be performed on visual features to mark the emotions of each video frame or segment, and create a video emotion fluctuation map.
[0107] Among them, the expression of the visual emotion value is as follows:
[0108] (2);
[0109] In the formula, represents the visual emotion value of the video at time t, represents the weight of the jth visual feature, represents the emotion score of the visual feature, represents the number of visual features.
[0110] (3) Analyze the audio data in the video based on audio signal processing technology to obtain an audio emotion value related to emotion fluctuation.
[0111] Analyze the audio data in the video (such as character dialogues, background sounds, etc.), audio, pitch changes, etc., and extract audio emotion information related to emotion fluctuation from it to determine the audio emotion value.
[0112] Specifically, audio features are extracted from the audio data of the video (such as the changes in the pitch, speed, pause, volume, and intonation of the speech) based on audio signal processing technology; the emotion scores and weights corresponding to each audio feature are determined; based on the emotion scores and weights corresponding to each audio feature, the audio emotion value is obtained.
[0113] The determination process of the emotion score and weight corresponding to each audio feature is as follows:
[0114] The determination of the emotion score and weight of each audio feature is based on a speech emotion analysis model, combined with machine learning technology, and trained and optimized through a large amount of speech data.
[0115] (1) The emotion score of the pitch (high or low) of the speech
[0116] Relationship between pitch changes and emotions: High pitch is usually associated with positive emotions such as excitement and happiness, while low pitch may be associated with negative emotions such as sadness and calmness.
[0117] This application will set corresponding emotion scores for each pitch range:
[0118] High pitch (e.g., >300 Hz): +0.7 (indicating positive emotions such as excitement);
[0119] Medium pitch (e.g., 150 Hz - 300 Hz): +0.3 (indicating neutral or slightly positive emotions);
[0120] Low pitch (e.g., <150 Hz): -0.6 (indicating negative emotions such as sadness).
[0121] These scores are obtained through training based on emotional speech data and will be fine-tuned according to the context in actual applications.
[0122] (2) Emotion scores for speech rate
[0123] Relationship between speech rate and emotional fluctuations: A faster speech rate usually indicates urgency, nervousness or excitement, while a slower speech rate indicates contemplation, frustration or sadness.
[0124] This application sets emotion scores for speech rate:
[0125] Fast speech rate (e.g., >150 words per minute): +0.8 (anxiety, eagerness, etc.);
[0126] Normal speech rate (e.g., 120 - 150 words per minute): +0.2 (stable, calm, etc.);
[0127] Slow speech rate (e.g., <120 words per minute): -0.5 (contemplation, fatigue, sadness, etc.).
[0128] These scores are also based on the learning and analysis of a large amount of emotional speech data.
[0129] (3) Emotion scores for pauses and tone changes
[0130] The frequency and duration of pauses can reflect emotional fluctuations. For example, longer pauses are usually associated with thinking, hesitation or nervousness:
[0131] Short pause time (e.g., <1 second): +0.2 (stable emotion);
[0132] Long pause time (e.g., >1 second): -0.6 (indicating nervousness, thinking or emotional fluctuations).
[0133] Changes in tone (such as a rising tone) also affect emotion scores:
[0134] Rising tone: +0.7 (indicating doubt, surprise, anxiety);
[0135] Falling tone: -0.5 (indicating contemplation, helplessness).
[0136] (4)Emotional score of volume
[0137] Changes in volume are usually related to the level of emotion. Volume increases when emotions are high and decreases when emotions are low.
[0138] This application sets an emotional score for volume:
[0139] High volume: +0.8 (indicating excitement, agitation, etc.);
[0140] Normal volume: 0 (no significant emotion);
[0141] Low volume: -0.5 (indicating sadness, contemplation, etc.).
[0142] (5)Emotional score of intonation
[0143] The rise or fall of intonation can also reflect changes in emotion:
[0144] Rising intonation (such as a rising tone): +0.6 (indicating surprise, doubt, excitement);
[0145] Falling intonation (such as a falling tone): -0.4 (indicating determination, helplessness, etc.).
[0146] The dynamic adjustment of emotional scores and weights. The audio emotion value is not fixed but is dynamically adjusted according to the context and user feedback. As the user's viewing experience and emotional responses accumulate, the system will adjust the weights and scores of each audio feature, so as to provide more personalized and accurate emotion prediction and feedback.
[0147] Among them, audio features can also be classified into emotion labels based on a speech emotion analysis model.
[0148] In practical applications, the impact of audio changes on the emotions of the audience can be analyzed, the relationship between audio fluctuations and video emotion changes can be marked, and it can be fed back to the system in real time.
[0149] Among them, the expression of the audio emotion value is as follows:
[0150] (3);
[0151] In the formula, represents the audio emotion value of the video at time t, represents the weight of the k-th audio feature, represents the emotional score of the audio feature, Represents the number of audio features.
[0152] (4) Based on the text emotion value, visual emotion value, and audio emotion value at each time point of the video, determine the comprehensive video emotion value of the video at each time point, mark the position where each comprehensive video emotion value appears in the video as an emotion fluctuation point, and determine the emotion fluctuation region containing the emotion fluctuation point.
[0153] Specifically, combine the text emotion value, visual emotion value, and audio emotion value to determine the comprehensive video emotion value of the video at each time point, that is, the comprehensive emotion state of the video content at this event point; use the sliding window algorithm or time series analysis method, and mark the position where each comprehensive video emotion value appears in the video as an emotion fluctuation point, and draw an emotion fluctuation curve based on each emotion fluctuation point. According to the emotion fluctuation curve, the system can predict the user's emotional reaction in advance and provide data support for subsequent user interaction and personalized recommendation.
[0154] Among them, the expression of the comprehensive video emotion value is as follows:
[0155] (4);
[0156] In the formula, represents the comprehensive video emotion value of the video at time point t, represents the weight of the text emotion value, represents the weight of the visual emotion value, represents the weight of the audio emotion value.
[0157] In one embodiment, in step S102, when the video plays to the emotion fluctuation region, obtain multi-dimensional emotion data of the user when watching the video, and generate a personalized emotion response on the video based on the multi-dimensional emotion data. Specifically, it may include:
[0158] (1) When the video plays to the emotion fluctuation region, obtain multi-dimensional emotion data of the user when watching the video.
[0159] Among them, the multi-dimensional emotion data includes but is not limited to: text emotion value, voice emotion value, and physiological emotion value.
[0160] Specifically, (a) The process of obtaining the text emotion value is as follows:
[0161] When the user watches the video, they express their emotions by inputting bullet screens, comments, etc. The system analyzes the text input by the user to identify their emotional state (such as happy, angry, surprised, etc.), combines natural language processing technology, determines the emotional vocabulary in the text, and attaches the corresponding emotion score and weight to each emotional vocabulary.
[0162] The system can also convert the user's emotional expressions into corresponding emotion tags and update the emotional fluctuation graph in real time.
[0163] The expression formula for text emotion value is as follows:
[0164] (5);
[0165] In the formula, represents the text emotion value of the user's text at time t, represents the weight of the i-th emotional word, represents the emotion score of the i-th emotional word, and N represents the number of emotional words.
[0166] (b)The process of obtaining the voice emotion value is as follows:
[0167] Through intelligent speech recognition technology, the system can monitor the user's voice input in real time (such as voice comments, interactive voice commands, etc.). Through voice emotion analysis, the system can identify the audio features in the user's voice (such as excited, calm, nervous, etc.), generate corresponding emotion feedback, and attach corresponding emotion scores and weights to each audio feature. Combining parameters such as tone, speech rate, and speech intensity, the voice emotion analysis uses an emotion analysis model to classify the voice and obtain the emotion type and intensity.
[0168] The expression formula for voice emotion value is as follows:
[0169] (6);
[0170] In the formula, represents the voice emotion value of the user's voice at time t, represents the weight of the k-th audio feature, represents the emotion score of the k-th audio feature, represents the number of audio features.
[0171] (c)The process of obtaining the physiological emotion value is as follows:
[0172] With the help of intelligent hardware devices (such as smart watches, earphones, etc.), the system can collect the user's physiological data in real time, such as heart rate, skin conductance response, etc. These physiological data can reflect the user's emotional fluctuations during the video viewing process, that is, physiological emotions. For example, an accelerated heart rate or an increased skin conductance response usually indicates that the user is in a tense or excited state.
[0173] Through the analysis of sensor data and physiological signals, the system can further refine the user's emotional response and adjust the presentation of video content according to these physiological change data.
[0174] (2)Fuse the multi-dimensional emotion data, and generate the personalized emotion response on the video according to the result of the emotion data fusion.
[0175] (a)Based on the text emotion value, voice emotion value, and physiological emotion value at each time point when the user watches the video, determine the comprehensive emotion value of the user at each time point.
[0176] The system integrates the emotion data (text, voice, physiological signals, etc.) collected through different channels through a weighted fusion algorithm to integrate the information feedback by each emotion feedback and generate a comprehensive user emotion state, that is, the comprehensive emotion value of the user. This fusion algorithm can minimize the inaccuracy of a single data source and improve the accuracy and comprehensiveness of emotion analysis.
[0177] The expression of the comprehensive emotion value of the user is as follows:
[0178] (7);
[0179] In the formula, represents the comprehensive emotion value of the user at time point t, represents the weight of the text emotion value, represents the weight of the voice emotion value, represents the weight of the physiological emotion value.
[0180] (b)Generate a personalized emotion response on the video based on the comprehensive emotion value of the user.
[0181] According to the comprehensive emotion value of the user at each time point, the real-time emotion fluctuation of the user can be determined to generate a personalized emotion response. These emotion responses can be content adjustments (such as video recommendations, plot modifications, bullet screen interactions, etc.), or triggers for user interactions (such as the appearance of bullet screens, emotion feedback animations, etc.). In addition, the system can also adjust the next recommended video or insert emotion enhancement elements in a timely manner according to the user's emotion fluctuation to provide a more immersive viewing experience for the user.
[0182] In practical applications, customized emotion feedback can be provided based on real-time user emotion data. For example, when the user is in a pleasant mood, the system can push more relaxing and funny content; if the user is in a low mood, the system may encourage the user to improve their mood through interactive mini-games, emotionally warm pictures, etc. This application can also achieve interactive feedback between the user and the video. For example, at the peak of emotion fluctuation, the system may trigger an instant interactive session, such as interacting with video characters, or adjusting the screen brightness, volume, etc. to enhance the depth of the emotion experience.
[0183] In one embodiment, step S103 predicts the user's emotional fluctuation trend based on the multi-dimensional emotional data and adjusts the video content based on the user's emotional fluctuation trend, which specifically includes:
[0184] (1) Obtain video content feature data and user interaction behavior data, and input the multi-dimensional emotional data, video content feature data, and user interaction behavior data into a set emotional fluctuation prediction model to predict the user's emotional fluctuation trend.
[0185] Among them, the data required for training the emotional fluctuation prediction model includes: the user's historical multi-dimensional emotional data, video content feature data, external environmental factors (such as time period, device type, etc.), and user interaction behavior data (such as click, skip, pause, etc.). According to the user's historical multi-dimensional emotional data, including information such as text, voice, and physiological data, a user emotion time series is constructed to provide a historical reference for predicting the user's emotional fluctuation.
[0186] In practical applications, based on deep learning models, such as Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), etc., by analyzing the user's multi-dimensional emotional data, combining the user's behavior data and video content features, the user's emotional fluctuation trend can be predicted.
[0187] The LSTM model can handle the dependency relationships of sequential data, so it is very suitable for the prediction task of emotional fluctuations.
[0188] The expression of the user's emotion prediction value is as follows:
[0189] (8);
[0190] In the formula, represents the user's emotion prediction value at the next time point t + 1, represents the user's multi-dimensional emotional data at the current time point t, represents the video content feature data (such as plot nodes, camera switching, etc.), represents the user interaction behavior data (such as fast forward, pause, etc.).
[0191] (2) Adaptively adjust the video content based on the user's emotional fluctuation trend.
[0192] Among them, adjusting the video content includes but is not limited to: emotion-driven content adjustment, multi-modal content adjustment, interactive content adjustment, etc.
[0193] Specifically, the process of emotion-driven content adjustment includes:
[0194] Based on the user's emotional fluctuation trend, the video content can be adjusted in real time. Emotion-driven content adjustment includes, but is not limited to, aspects such as video images, sound effects, camera switching, and plot advancement. For example, if it is predicted that the user is about to enter an emotional climax or a tense state, the system can switch to more impactful images or sound effects in advance to enhance the emotional experience. If it is predicted that the user's emotion is about to enter a tense state, the system can automatically increase the volume of the background music in the video or add some visual effects of rapid camera switching; if it is predicted that the user is about to enter a relaxed state, the system can increase the comfort level by adjusting the video image tone and warm background sound.
[0195] The process of multi-modal content adjustment includes:
[0196] In addition to video images and sound effects, the system can also adjust other multi-modal content, such as the intelligent generation or pushing of bullet screen content, the insertion of interactive plots, etc. For example, when it is predicted that the user is in a low mood, the system can automatically generate bullet screens with humorous elements to boost the atmosphere, or guide the user to interact with the video content to relieve the user's negative emotions.
[0197] The process of interactive content adjustment includes:
[0198] In some specific situations (such as interactive dramas, selective storylines), the user's emotional fluctuations will affect the plot direction. The system predicts the emotional fluctuations and automatically selects the plot path that matches the user's emotional state, enabling the user to participate in the personalized adjustment of the content, thereby improving the immersion and engagement.
[0199] It should be noted that the process of adaptively adjusting video content based on the user's emotional fluctuation trend is not a one-time process. It is a real-time dynamic process. As the user's emotions change continuously, the system will make adaptive adjustments according to the real-time emotional feedback. For example, when the user's emotional fluctuations do not follow the prediction, the system will re-predict the user's emotional fluctuation trend based on the new emotional feedback data and make a secondary adjustment to the video content according to the adjusted emotional prediction value.
[0200] The expression of the adjusted emotional prediction value is as follows:
[0201] (9);
[0202] In the formula, represents the adjusted emotional prediction value, represents the feedback adjustment coefficient, indicating the degree of correction of the current real-time emotion to the prediction value, represents the multi-dimensional emotional data of the user at the current time point t, represents the user's emotional prediction value at the time point t.
[0203] In one embodiment, the video recommendation content can also be dynamically adjusted according to the user's emotional fluctuation trend. For example, when the user is in a high mood, the system recommends more intense content with similar plots; when the user is in a low mood, the system recommends relaxing and pleasant content to ensure that the user's emotions are properly guided and regulated, thus achieving personalized content recommendation.
[0204] In one embodiment, based on the user's emotional data and user behavior characteristics, the video recommendation content, interaction methods, and personalized emotional expression paths can be dynamically adjusted. Through emotion recognition and user behavior feature analysis, customized recommendation content and interaction methods are provided for different user emotional states, thereby enhancing the emotional connection between the user and the video content and improving user engagement and retention rate.
[0205] In one embodiment, the video interaction method may further include:
[0206] Extract the user's emotional characteristics from the multi-dimensional emotional data of the user;
[0207] Determine the user's emotional type based on the user's emotional characteristics;
[0208] Recommend emotional interaction methods and / or videos with a high degree of matching with the user's emotional type.
[0209] Specifically, based on data such as real-time emotional feedback, user viewing history, and social media behavior, a user emotional profile is constructed. The user emotional profile includes characteristics such as the user's emotional fluctuation pattern, user video viewing preferences, and user emotional tendencies. These characteristics will be used as input data for the personalized recommendation model.
[0210] The user's emotions are monitored in real time through multi-modal data (such as text, voice, video behavior, etc.) to obtain the user's multi-dimensional emotional data, and the user's emotional characteristics are extracted from the user's multi-dimensional emotional data. The user's emotional characteristics can be the intensity of positive emotions, negative emotions, or neutral emotions, etc.
[0211] Based on the user portrait of the user's emotional characteristics, the user is classified into emotional types through machine learning techniques (such as clustering analysis, decision trees, support vector machines, etc.), such as: emotionally stable type, emotionally volatile type, emotionally sensitive type, etc. Corresponding emotional interaction methods are provided according to the user's emotional type, and corresponding videos can be recommended to the user according to the user's emotional type.
[0212] Provide corresponding emotional interaction methods according to the user's emotional type. For example, optimize the interaction design according to the emotional response types of different users. For users with large emotional fluctuations, the system can provide more emotional feedback options, such as bullet screen input, voice expressions, gesture control, etc., to promote users to express emotions more naturally in the video. When the user's emotion is highly excited, the system can analyze the user's emotional tendency through speech recognition and natural language processing technologies and recommend interactive content that matches their emotional state. For example, when the user expresses anger, the system can push relevant plots or bullet screen content to relieve their emotion.
[0213] (2)Emotion-driven content recommendation algorithm
[0214] Adopt a deep learning recommendation system to perform personalized content recommendation by combining user emotional characteristics and user viewing behavior data. By predicting the user's emotional trend, push content that meets the current emotional needs in advance.
[0215] The recommendation formula is as follows:
[0216] (10);
[0217] In the formula, represents the i-th video content recommended to the user at time point t, represents the user's emotional characteristics, represents the emotional attribute of the video content, represents the environmental and user behavior context information.
[0218] The system not only makes recommendations based on the user's historical data, but also adjusts according to the user's real-time emotional fluctuations. For example, if the system detects that the user is in a low mood, it will give priority to recommending positive video content; if the user is excited, it will recommend some content with intense plots to maintain the continuity of the viewing experience and emotional matching.
[0219] Among them, the process of recommending emotional interaction methods and / or videos with a high degree of matching with the user's emotional type is as follows:
[0220] It is possible to score the emotional matching degree of each recommended content, and the content with a higher matching degree is more in line with the user's current emotional needs. The system calculates the matching degree of each content with the current emotional state through an emotion prediction model and a content analysis module, so as to give priority to recommending content with a high emotional fit.
[0221] The emotional matching degree formula is as follows:
[0222] (11);
[0223] In the formula, represents the matching degree of the current content with the user's emotion, Represents the user's emotional characteristics of the current user, Represents the emotional attribute of the i-th content, Represents the number of contents.
[0224] In practical applications, based on the user's emotion type, the system can also calculate the content preference model in different emotional states. For example, users with a pleasant emotion may prefer light and humorous content, while users with an anxious emotion tend to seek tense and exciting content. Through this personalized emotion-content matching, the recommendation results are more in line with the user's emotional needs, improving user satisfaction.
[0225] In one embodiment, the video interaction method may further include:
[0226] Providing a personalized emotion interaction method based on the user's emotional fluctuation trend.
[0227] For example, when the user's emotion reaches a climax, the system can guide the user to participate in more interactive emotional expression sessions, such as posting humorous or touching bullet comments, or participating in plot selection-based interactions to enhance the sense of immersion.
[0228] It can also intelligently push corresponding bullet comment content by analyzing the relevance between the user's emotion and the video plot, or set and automatically generate interactive emotional feedback content according to the plot and the user's emotion, such as selective plots, character interactions, etc.
[0229] This application can enhance the user's sense of immersion, participation, and emotional connection by capturing the user's emotional state in real time, analyzing emotional fluctuations, and adjusting the video content and interaction method according to these emotional feedbacks. Through the dynamic adjustment of emotional feedback, the user's viewing experience is improved in multiple aspects such as vision, audio, and interaction, enabling the audience to better resonate with the content, and through refined emotion regulation, allowing the user to immerse in a more real and emotional interaction environment.
[0230] The system monitors the user's emotional fluctuation state in real time and uses emotion recognition technologies (such as multi-modal perception technologies like facial expression analysis, speech emotion analysis, eye movement tracking, action recognition, etc.) to capture the user's emotional reaction. According to the intensity and type of the user's current emotion, the system will automatically adjust the content presentation method of the video, including the picture tone, music melody, plot push, etc., to enhance the audience's emotional resonance and strengthen the sense of immersion in viewing. Based on the time series data of the user's emotional fluctuations, a "emotional fluctuation curve" of the user is generated. When the user's emotion enters a climax or a trough, the system will adjust the video content according to the preset emotional curve. Especially for the peak (excited, pleasant) and trough (sad, fearful, etc.) states of the emotion, the system enhances by changing the visual and audio elements of the video.
[0231] The expression of the emotional fluctuation curve formula is as follows:
[0232] (12);
[0233] In the formula, represents the emotion enhancement effect at time point t, represents the i-th emotional feature of the user, represents the i-th emotional trigger point in the video content, Indicates the number of different emotion points.
[0234] In one embodiment, step S103 may specifically include:
[0235] The audio of the video content and / or the visual of the video content are adjusted based on the trend of the user's emotional fluctuations.
[0236] Specifically, when adjusting the video audio, the system can automatically adjust the background music or sound effects of the video based on emotion recognition technology to match the rhythm, pitch, volume, etc. of the audio with the user's emotional state. For example, when the user is depressed, the system can push soothing background music to enhance their emotional recovery; when the user is emotionally excited, the system can enhance the sound effects or music with a strong sense of rhythm to drive the user's emotions to a climax.
[0237] The expression for audio matching is as follows:
[0238] (13);
[0239] In the formula, Indicates the matching degree between audio and emotion, Indicates the user's emotional characteristics. Represents the emotion adjustment parameters of the audio of the video content.
[0240] When adjusting the video visually, the system adjusts the hue, brightness, contrast and other visual elements in the video in real time according to the user's emotional fluctuations. For example, when the user is nervous or excited, the system can enhance the saturation and contrast of the picture and use more vivid colors; when the user is depressed, the system can use softer tones and gradient backgrounds to create a soothing effect. Through this visual adjustment, the system can accurately guide the user's emotional experience, allowing the user to feel a visual atmosphere that matches the emotional state when watching the video, further enhancing the sense of immersion.
[0241] In practical applications, this application can also achieve role emotion guidance and user interaction.
[0242] Specifically, based on the user's emotional state, the system can guide the user to interact emotionally with the virtual characters in the video. For example, when the user feels lonely or sad, the system can recommend warm, supportive or comforting character interactions to help the user relieve the emotion; when the user is in an excited or cheerful state, the system can recommend humorous or passionate character interactions to further stimulate the user's emotions. Users can choose to interact with the characters in the video for emotional feedback. For example, when watching an emotionally rich plot, users can choose to like, express emotional feelings, or choose to interact with the characters to enhance the emotional resonance between the audience and the content. The system continuously adjusts the interactive content based on the user's emotions and interactive feedback to maintain a high degree of emotional matching.
[0243] In one embodiment, the system can adjust the interactive content, plot direction and video display method according to the user's emotional state, thereby enhancing the user's immersion. For example, when the user is watching a horror or suspense video, the system can timely enhance the tension of the sound and picture to mobilize the user's emotional response; while when watching romantic and warm content, the system can reduce the tension and enhance the comfortable and warm emotional experience.
[0244] The expression of the emotion matching feedback formula is as follows:
[0245] (14);
[0246] In the formula, represents the immersion score, Indicates the user's emotional characteristics. Indicates the immersion index of the video content. and Represents the weight coefficient.
[0247] Corresponding to the above method embodiment, the present application also discloses a video interaction device.
[0248] See also Figure 2 , a schematic diagram of the structure of a video interaction device disclosed in an embodiment of the present application, the device may include:
[0249] The determination unit 201 is used to perform a comprehensive analysis on the video based on the multimodal emotion data of the video, determine the emotion fluctuation points in the video, and determine the emotion fluctuation area containing the emotion fluctuation points.
[0250] In practical applications, multimodal information (such as text, vision, and audio) can be used to conduct a comprehensive analysis of videos, predict the locations where users may experience emotional fluctuations when watching videos, mark the emotional fluctuation points in the video in advance, and identify potential emotional fluctuation areas (such as highs, lows, tension, joy, etc.), providing data support for subsequent emotional interactions.
[0251] For example, a sudden terrifying scene in a video may trigger a user's startled emotion, and the system can identify this potential emotional fluctuation area in advance through emotion analysis. At this time, an emotion feedback interaction module can be provided in the potential emotional fluctuation area, or the maximization of emotional value can be triggered at an appropriate time.
[0252] An acquisition unit 202, configured to obtain multi-dimensional emotion data of the user when watching the video when the video is played to the emotional fluctuation area, and generate a personalized emotion response on the video based on the multi-dimensional emotion data.
[0253] Among them, the multi-dimensional emotion data includes but is not limited to: text emotion value, voice emotion value, and physiological emotion value.
[0254] The personalized emotion response includes but is not limited to: personalized emotion interaction methods, personalized content recommendations, etc.
[0255] For example, when the user's emotion reaches a climax, the system can guide the user to participate in a more interactive emotional expression session, such as posting humorous or touching bullet comments.
[0256] For example, personalized content recommendation is to recommend interactive content that matches the user's emotional state to the user.
[0257] When the video is played to the emotional fluctuation area, the emotion feedback interaction module is used to collect and process the user's emotion feedback in real time during the video viewing process, capture the user's emotion changes, so as to respond to the user's emotion expression in a timely manner, thereby optimizing the user experience and increasing interactivity. The multi-dimensional emotion data is not limited to text, but also includes various data such as voice, facial expressions, gestures, heart rate, etc. By comprehensively analyzing these feedback information, the system can generate personalized emotion responses or interaction behaviors. So that the user's emotion expression can not only be real-time fed back to the platform, but also be dynamically processed by the platform system, and then realize personalized emotion regulation and enhance the movie viewing experience.
[0258] An adjustment unit 203, configured to predict the user's emotion fluctuation trend based on the multi-dimensional emotion data, and adjust the video content based on the user's emotion fluctuation trend.
[0259] During the prediction of the user's emotion fluctuation trend and the adjustment of the video content, during the user's movie viewing process, based on the real-time collected multi-dimensional emotion data of the user, predict the user's emotion fluctuation trend, and accordingly adjust the video content to enhance the user's immersion and movie viewing experience. Through accurate emotion fluctuation prediction, the system can identify the possible emotion fluctuations of the user in advance, and dynamically adjust the video content (such as adjusting the picture rhythm, sound effects, plot settings, etc.) to match the user's emotion needs, so as to achieve emotional resonance with the user.
[0260] In practical applications, adjusting video content based on the user's emotional fluctuation trend may include: adjusting the emotion-driven content, multi-modal content, and interactive content in the video, etc.
[0261] Adjusting video content based on the user's emotional fluctuation trend may also include: adjusting video recommendation content. When the user's mood is high, the system recommends more intense and exciting content of a similar nature; when the user's mood is low, the system recommends relaxing and pleasant content to ensure that the user's emotions are properly guided and regulated.
[0262] In summary, the present application discloses a video interaction method, which comprehensively analyzes the video based on video multi-modal emotion data, determines the emotional fluctuation points in the video, and determines the emotional fluctuation area containing the emotional fluctuation points. When the video plays to the emotional fluctuation area, multi-dimensional emotion data of the user when watching the video is obtained, and a personalized emotional response is generated on the video based on the multi-dimensional emotion data. The user's emotional fluctuation trend is predicted based on the multi-dimensional emotion data, and the video content is adjusted based on the user's emotional fluctuation trend. Through the collection and analysis of video multi-modal emotion data, the present application enables, when the user conducts video interaction, a personalized emotional response to be generated on the video according to the multi-dimensional emotion data of the user when watching the video, and the video content can be adjusted according to the user's emotional fluctuation trend, achieving instant and diverse emotional feedback, enhancing the real-time and diversity of emotional interaction, and thus meeting the user's emotional interaction needs.
[0263] In one embodiment, the determination unit 201 may include:
[0264] A text emotion value determination subunit, configured to perform sentiment classification on the text in the video, identify the sentiment words in the text, and determine the text emotion value related to emotional fluctuation based on the sentiment words;
[0265] A visual emotion value determination subunit, configured to perform visual emotion analysis on the video based on computer vision technology to obtain the visual emotion value related to emotional fluctuation;
[0266] An audio emotion value determination subunit, configured to analyze the audio data in the video based on audio signal processing technology to obtain the audio emotion value related to emotional fluctuation;
[0267] An emotional fluctuation area determination subunit, configured to determine the comprehensive video emotion value of the video at each time point based on the text emotion value, visual emotion value, and audio emotion value of the video at each time point, mark the position where each comprehensive video emotion value appears in the video as an emotional fluctuation point, and determine the emotional fluctuation area containing the emotional fluctuation point.
[0268] In one embodiment, the text emotion value determination subunit may specifically be configured to:
[0269] Extract all the text information from the video to obtain the text;
[0270] Perform sentiment analysis on the text to determine the sentiment words in the text;
[0271] Determine the emotion scores and weights of each sentiment word;
[0272] Based on the emotion scores and weights corresponding to each sentiment word, obtain the text emotion value.
[0273] In one embodiment, the visual emotion value determination subunit may specifically be used for:
[0274] Based on face recognition technology, identify the people and their facial expressions in the video, and use facial expression recognition algorithms to determine the emotions of the people;
[0275] Use a deep learning model to analyze the shot transitions and scene changes of the video, and detect visual signals of emotional fluctuations;
[0276] Take the people, their facial expressions, their emotions, and the visual signals as visual features, and determine the emotion scores and weights corresponding to each visual feature;
[0277] Based on the emotion scores and weights corresponding to each visual feature, obtain the visual emotion value.
[0278] In one embodiment, the audio emotion value determination subunit may specifically be used for:
[0279] Based on audio signal processing technology, extract audio features from the audio data of the video;
[0280] Determine the emotion scores and weights corresponding to each audio feature;
[0281] Based on the emotion scores and weights corresponding to each audio feature, obtain the audio emotion value.
[0282] In one embodiment, the obtaining unit 202 may specifically be used for:
[0283] When the video plays to the emotional fluctuation area, obtain multi-dimensional emotion data of the user while watching the video;
[0284] Perform emotion data fusion on the multi-dimensional emotion data, and generate the personalized emotion response on the video according to the emotion data fusion result.
[0285] In one embodiment, the adjustment unit 203 may specifically be used for:
[0286] Obtain video content feature data and user interaction behavior data, and input the multi-dimensional emotion data, the video content feature data, and the user interaction behavior data into a set emotion fluctuation prediction model to predict the user emotion fluctuation trend;
[0287] Adaptive adjust the video content based on the user emotion fluctuation trend.
[0288] In one embodiment, the video interaction device may further include:
[0289] An extraction unit, configured to extract user emotion features from the multi-dimensional emotion data;
[0290] An emotion type determination unit, configured to determine the user emotion type based on the user emotion features;
[0291] A recommendation unit, configured to recommend emotion interaction methods and / or videos with a high matching degree to the user emotion type.
[0292] In one embodiment, the video interaction device may further include:
[0293] An interaction method providing unit, configured to provide a personalized emotion interaction method for the user based on the user emotion fluctuation trend.
[0294] It should be noted that for the specific working principles of the components in the device embodiment, please refer to the corresponding parts of the method embodiment, which will not be elaborated here.
[0295] Corresponding to the above embodiment, the present application also discloses a computer storage medium, and the computer storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the steps shown in the video interaction method embodiment are implemented.
[0296] The computer storage medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may be a machine-readable signal medium or a machine-readable storage medium. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0297] Corresponding to the above embodiments, as Figure 3 shown, the present invention further provides an electronic device, which may include: a processor 1 and a memory 2;
[0298] Wherein, the processor 1 and the memory 2 complete mutual communication through a communication bus 3;
[0299] The processor 1 is configured to execute at least one instruction;
[0300] The memory 2 is configured to store at least one instruction;
[0301] The processor 1 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0302] The memory 2 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0303] Wherein, the processor executes at least one instruction to implement the steps shown in the video interaction method embodiment.
[0304] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0305] The various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0306] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video interaction method, characterized in that: include: Performing a comprehensive analysis on the video based on the multimodal emotion data of the video, determining the emotion fluctuation points in the video, and determining the emotion fluctuation area containing the emotion fluctuation points; When the video is played to the emotional fluctuation area, multi-dimensional emotional data of the user when watching the video is obtained, and a personalized emotional response is generated on the video based on the multi-dimensional emotional data; The user's emotion fluctuation trend is predicted based on the multi-dimensional emotion data, and the video content is adjusted based on the user's emotion fluctuation trend.
2. The video interaction method according to claim 1, characterized in that: The method of comprehensively analyzing the video based on the multimodal emotion data of the video, determining the emotion fluctuation points in the video, and determining the emotion fluctuation area containing the emotion fluctuation points, includes: Performing sentiment classification on the text in the video, identifying sentiment words in the text, and determining text sentiment values related to sentiment fluctuations based on the sentiment words; Performing visual emotion analysis on the video based on computer vision technology to obtain visual emotion values related to emotion fluctuations; Analyzing the audio data in the video based on audio signal processing technology to obtain audio emotion values related to emotion fluctuations; Based on the text emotion value, visual emotion value and audio emotion value of the video at each time point, the comprehensive emotion value of the video at each time point is determined, the position where each comprehensive emotion value appears in the video is marked as an emotion fluctuation point, and the emotion fluctuation area containing the emotion fluctuation point is determined.
3. The video interaction method according to claim 2, characterized in that: The emotional classification of the text in the video, identifying the emotional words in the text, and determining the emotional value of the text related to emotional fluctuations based on the emotional words, includes: Extracting all text information from the video to obtain the text; Performing sentiment analysis on the text to determine sentiment words in the text; Determining a sentiment score and a weight for each of the sentiment words; The text emotion value is obtained based on the emotion scores and weights corresponding to each of the emotion words.
4. The video interaction method according to claim 2, characterized in that: The performing visual emotion analysis on the video based on computer vision technology to obtain a visual emotion value related to emotion fluctuations includes: Identify people and their expressions in the video based on face recognition technology, and determine people's emotions using facial expression recognition algorithms; Use deep learning models to analyze the shot transitions and scene changes in the video to detect visual signals of emotional fluctuations; Taking the character, the character expression, the character emotion and the visual signal as visual features, and determining an emotion score and a weight corresponding to each visual feature; The visual emotion value is obtained based on the emotion scores and weights corresponding to the visual features.
5. The video interaction method according to claim 2, characterized in that: The step of analyzing the audio data in the video based on the audio signal processing technology to obtain the audio emotion value related to the emotion fluctuation includes: Extracting audio features from the audio data of the video based on audio signal processing technology; Determine the emotion score and weight corresponding to each of the audio features; The audio emotion value is obtained based on the emotion scores and weights corresponding to the respective audio features.
6. The video interaction method according to any one of claims 1 to 5, characterized in that: When the video is played to the emotional fluctuation area, multi-dimensional emotional data of the user when watching the video is obtained, and a personalized emotional response is generated on the video based on the multi-dimensional emotional data, including: When the video is played to the emotion fluctuation area, obtaining multi-dimensional emotion data of the user when watching the video; Emotion data fusion is performed on the multi-dimensional emotion data, and the personalized emotion response is generated on the video according to the emotion data fusion result.
7. The video interaction method according to claim 1, characterized in that: The predicting the user's emotion fluctuation trend based on the multi-dimensional emotion data, and adjusting the video content based on the user's emotion fluctuation trend, includes: Acquire video content feature data and user interaction behavior data, and input the multi-dimensional emotion data, the video content feature data and the user interaction behavior data into a set emotion fluctuation prediction model to predict the user emotion fluctuation trend; The video content is adaptively adjusted based on the emotion fluctuation trend of the user.
8. The video interaction method according to claim 1, characterized in that: Also includes: Extracting user emotion features from the multi-dimensional emotion data; Determining a user emotion type based on the user emotion characteristics; Recommend emotional interaction methods and / or videos that have a high degree of match with the user's emotional type.
9. The video interaction method according to claim 1, characterized in that: Also includes: Based on the user's emotional fluctuation trend, a personalized emotional interaction method is provided for the user.
10. A video interactive device, characterized in that: include: A determination unit, configured to perform a comprehensive analysis on the video based on the multimodal emotion data of the video, determine the emotion fluctuation points in the video, and determine the emotion fluctuation area containing the emotion fluctuation points; an acquisition unit, configured to acquire multi-dimensional emotion data of the user when watching the video when the video is played to the emotion fluctuation area, and generate a personalized emotion response on the video based on the multi-dimensional emotion data; An adjustment unit is used to predict the user's emotion fluctuation trend based on the multi-dimensional emotion data, and adjust the video content based on the user's emotion fluctuation trend.
11. A computer storage medium, characterized in that: The computer storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, the video interaction method as described in any one of claims 1 to 9 is implemented.
12. An electronic device, characterized in that: The electronic device comprises: a memory and a processor; The memory is used to store at least one instruction; The processor is used to execute the at least one instruction to implement the video interaction method as described in any one of claims 1 to 9.