An interactive method and system based on video content analysis

By combining video parsing, interactive analysis, and processing modules, and dynamically adjusting interaction methods and emotion recognition, the problem of insufficient user interaction in video content analysis is solved, and the intelligence and flexibility of personalized video recommendations are realized.

CN120856917BActive Publication Date: 2026-03-20BEIJING IACTIVE NETWORK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511354103.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-03-20
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing video content analysis technologies lack the ability to interact with users in real time, and cannot fully incorporate dynamic user feedback to adjust video recommendations, resulting in insufficient intelligence and flexibility in the recommendation results.

Method used

The video parsing module obtains login information and compares it with the historical identity database to determine the interaction strategy or generation strategy; the interaction analysis module dynamically adjusts the interaction method based on ambient lighting and sound; the interaction processing module accurately perceives the emotion type through convolutional neural networks and deep learning models; and the interaction recommendation module recommends videos based on the matching differences between playback type and emotion type.

Benefits of technology

It achieves targeted and adaptable interaction methods, improves the accuracy of emotion recognition and the intelligence level of recommendation results, ensures the fit between video content and user interaction, and meets personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856917B_ABST
    Figure CN120856917B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video interaction, and discloses an interactive method and system based on video content analysis, which comprises the following steps: a video analysis module acquires the playing type of a real-time playing video, acquires login information and compares the login information in a historical identity library; when a historical interaction strategy is identified, an interaction analysis module judges whether the interaction mode is to be re-determined based on ambient light and ambient sound; an interaction processing module judges whether shooting parameters are to be adjusted based on ambient water mist concentration and ambient particle concentration, determines an emotion type based on real-time facial data and a convolutional neural network model, and determines emotion audio fluctuation based on an audio signal and a deep learning model; and an interaction recommendation module judges whether there is a matching difference according to the playing type and the emotion type, and recommends a playing video according to a judgment result. Through matching of video content analysis and an interaction process, the application ensures the intelligent degree and flexibility of a recommendation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video interaction technology, and more specifically, to an interaction method and system based on video content analysis. Background Technology

[0002] With the rapid development of artificial intelligence and machine learning technologies, video content analysis technology has been widely used in education, entertainment, advertising, and other fields. By analyzing video content, it is possible to uncover the themes and emotions of videos. However, with the dramatic increase in video data, users face certain challenges in finding videos that match their interests and needs from a massive amount of video data. On the one hand, video content analysis lacks real-time interactive capabilities with users; on the other hand, it cannot fully incorporate dynamic user feedback to adjust the content of video recommendations. As a result, the recommended videos fail to meet users' personalized needs, leading to insufficient intelligence and flexibility in the recommendation results.

[0003] Therefore, how to provide an interactive method and system based on video content analysis is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the present invention proposes an interactive method and system based on video content analysis, aiming to solve the problems that video content analysis lacks real-time interaction capabilities with users, cannot fully integrate dynamic user feedback to adjust video recommendation content, and the recommended videos fail to meet users' personalized needs, resulting in insufficient intelligence and flexibility in the recommendation results.

[0005] In one aspect, the present invention proposes an interactive system based on video content analysis, comprising:

[0006] The video parsing module is configured to obtain the playback type of the real-time video, obtain login information and compare the login information with the historical identity database, and determine the historical interaction strategy or interaction generation strategy based on the comparison result.

[0007] The interaction analysis module is configured to determine the interaction method based on ambient light and ambient sound when the interaction generation strategy is identified, and to determine whether to re-determine the interaction method based on ambient light and ambient sound when the historical interaction strategy is identified. The interaction method includes touch screen interaction, gesture interaction and voice interaction.

[0008] The interaction processing module is configured to, when determining the touch screen interaction or the gesture interaction, acquire an environmental water mist concentration and an environmental particle concentration, judge whether to adjust a shooting parameter based on the environmental water mist concentration and the environmental particle concentration, determine facial real-time data, and determine an emotion type based on the facial real-time data and a convolutional neural network model, when determining the voice interaction, acquire an audio signal and convert the audio signal into an emotional text, determine an emotional audio fluctuation based on the audio signal and a deep learning model, and determine the emotion type according to the emotional text and the emotional audio fluctuation.

[0009] The interaction recommendation module is configured to judge whether there is a matching difference according to the playing type and the emotion type, and recommend a playing video according to a judgment result.

[0010] Further, when acquiring login information and comparing the login information in a historical identity library, and determining a historical interaction strategy or an interaction generation strategy according to a comparison result, the method comprises:

[0011] The historical identity library comprises a plurality of historical login information and a plurality of historical interaction modes, and each historical login information corresponds to a historical interaction mode.

[0012] When the historical identity library comprises historical login information identical to the login information, the video analysis module determines the historical interaction strategy.

[0013] When the historical identity library does not comprise historical login information identical to the login information, the video analysis module determines the interaction generation strategy.

[0014] Further, when recognizing the interaction generation strategy, and determining an interaction mode based on environmental illumination and environmental sound, the method comprises:

[0015] When the environmental illumination is within a standard environmental illumination range, and the environmental sound is not within a standard environmental sound range, the interaction analysis module determines that the interaction mode is the gesture interaction.

[0016] When the environmental illumination is not within the standard environmental illumination range, and the environmental sound is within the standard environmental sound range, the interaction analysis module determines that the interaction mode is the voice interaction.

[0017] When the environmental illumination is not within the standard environmental illumination range, and the environmental sound is not within the standard environmental sound range, the interaction analysis module determines that the interaction mode is the touch screen interaction.

[0018] Further, when recognizing the historical interaction strategy, and judging whether to redetermine the interaction mode based on the environmental illumination and the environmental sound, the method comprises:

[0019] The interaction analysis module acquires a historical interaction mode corresponding to the same historical login information as the login information, and records it as a to-be-determined interaction mode;

[0020] When the to-be-determined interaction mode does not meet the corresponding standard environmental condition, the to-be-determined interaction mode is cancelled, and the interaction mode is re-determined;

[0021] When the to-be-determined interaction mode meets the corresponding standard environmental condition, the interaction mode is determined according to the to-be-determined interaction mode;

[0022] The standard environmental condition is the standard environmental light range and the standard environmental sound range.

[0023] Further, when the touch screen interaction or the gesture interaction is determined, the environmental water mist concentration and the environmental particle concentration are acquired, and whether to adjust the shooting parameter is determined based on the environmental water mist concentration and the environmental particle concentration. When the face real-time data is determined, it includes:

[0024] The interaction processing module acquires the environmental water mist concentration and the environmental particle concentration, and performs data standardization on the environmental water mist concentration and the environmental particle concentration to determine the target environmental water mist concentration and the target environmental particle concentration, and determines the environmental visibility according to the target environmental water mist concentration and the target environmental particle concentration;

[0025] When the environmental visibility is greater than or equal to an environmental visibility threshold, it is determined that the shooting parameter is not adjusted, and the current shooting parameter is used to determine the face real-time data;

[0026] When the environmental visibility is less than the environmental visibility threshold, it is determined that the shooting parameter is adjusted, the adjustment is to increase the exposure compensation / increase the aperture, and the face real-time data is determined according to the adjusted shooting parameter.

[0027] Further, when the emotion type is determined based on the face real-time data and the convolutional neural network model, it includes:

[0028] The interaction processing module acquires a face sample set, and samples the face sample set to determine a training set and a test set;

[0029] An initial convolutional neural network model is preselected, the initial convolutional neural network model is iteratively trained according to the training set, and the initial convolutional neural network model after iterative training is verified according to the test set;

[0030] If the ternary loss of the initial convolutional neural network model after the current iterative training is greater than the ternary loss of the initial convolutional neural network model after the previous iterative training, the initial convolutional neural network model after the current iterative training is gradient clipped, and the iterative training is continued.

[0031] If the ternary loss of the initial convolutional neural network model after the current iteration training is less than or equal to the ternary loss of the initial convolutional neural network model after the previous iteration training, the iteration training is stopped, and the initial convolutional neural network model after the current iteration training is determined as the convolutional neural network model, and the face real-time data is substituted into the convolutional neural network model to determine the emotion type.

[0032] Further, when determining the voice interaction, acquiring the audio signal and converting it into emotional text, comprising:

[0033] The interaction processing module acquires the initial audio signal and performs noise reduction to determine the audio signal;

[0034] The interaction processing module converts the audio signal into emotional text and determines the emotional category of each noun and the emotional category of each adjective in the emotional text, the emotional category including positive emotion, negative emotion and neutral emotion;

[0035] Determine the positive number ratio of the total positive emotion number to the total emotion number, determine the negative number ratio of the total negative emotion number to the total emotion number, and determine the neutral number ratio of the total neutral emotion number to the total emotion number;

[0036] Arrange the positive number ratio, negative number ratio and neutral number ratio in descending order, and determine the emotional orientation of the emotional text according to the arrangement result.

[0037] Further, when determining the emotional audio fluctuation based on the audio signal and the deep learning model, and determining the emotion type according to the emotional text content and the emotional audio fluctuation, comprising:

[0038] The interaction processing module determines the time domain image based on the audio signal and the deep learning model, and acquires all adjacent amplitude intervals in the time domain image;

[0039] Statistical first amplitude interval number of amplitude interval greater than amplitude interval threshold, and statistical second amplitude interval number of amplitude interval less than or equal to amplitude interval threshold;

[0040] When the first amplitude interval number is greater than the second amplitude interval number, it is determined that the emotional audio fluctuation is negative fluctuation;

[0041] When the first amplitude interval number is less than the second amplitude interval number, it is determined that the emotional audio fluctuation is positive fluctuation;

[0042] When the first amplitude interval number is equal to the second amplitude interval number, it is determined that the emotional audio fluctuation is neutral fluctuation;

[0043] The amplitude interval threshold is determined according to the speech speed of the audio signal, and an emotion type is determined based on the emotion orientation and the emotion audio fluctuation.

[0044] Further, when judging whether there is a matching difference according to the playing type and the emotion type, and recommending playing a video according to the judgment result, it comprises:

[0045] When the matching of the playing type and the emotion type has a difference, a video is recommended to be played according to the emotion type;

[0046] When the matching of the playing type and the emotion type has no difference, a video is recommended to be played according to the playing type.

[0047] Compared with the prior art, the beneficial effects of the present application are that: the video analysis module matches the historical interaction strategy for old users and generates the interaction strategy for new users by comparing the login information with the historical identity library, ensuring the pertinence of the determined interaction mode, the interaction analysis module dynamically determines or adjusts the interaction mode based on the environmental light and the environmental sound, switches to voice interaction when the light is dim, and switches to gesture interaction in a noisy scene, ensuring the adaptability of the interaction mode to the real-time playing video, thereby improving the compatibility of the system to the environment, the interaction processing module accurately perceives the emotion type for different interaction modes, in touch screen interaction or gesture interaction, the shooting parameters under the influence of the environmental water mist concentration and the environmental particle concentration are adjusted to ensure the accuracy of the real-time facial data, and in voice interaction, the emotion type is determined through the dual dimensions of emotional text and emotional audio fluctuation, realizing the dynamic capture of the real-time state of different interaction modes, laying a data foundation for recommending playing a video, thereby avoiding the risk of feedback lag, recommending playing a video according to the matching difference between the emotion type and the playing type, ensuring the degree of fit between the video analysis form and the interaction process, and ensuring the intelligent level and flexibility of the recommendation result.

[0048] On the other hand, the present application also provides an interaction method based on video content analysis, which is used for the above-mentioned interaction system based on video content analysis, comprising:

[0049] Obtaining the playing type of the real-time playing video, obtaining the login information and comparing the login information in the historical identity library, and determining the historical interaction strategy or the interaction generation strategy according to the comparison result;

[0050] When the interaction generation strategy is recognized, the interaction mode is determined based on the environmental light and the environmental sound, and when the historical interaction strategy is recognized, it is judged whether to determine the interaction mode again based on the environmental light and the environmental sound, the interaction mode comprising touch screen interaction, gesture interaction and voice interaction;

[0051] When the touch screen interaction or the gesture interaction is determined, the ambient water mist concentration and the ambient particle concentration are acquired, it is determined whether to adjust the shooting parameter based on the ambient water mist concentration and the ambient particle concentration, the real-time facial data is determined, and the emotion type is determined based on the real-time facial data and the convolutional neural network model; when the voice interaction is determined, the audio signal is acquired and converted into emotion text, the emotion audio fluctuation is determined based on the audio signal and the deep learning model, and the emotion type is determined according to the emotion text and the emotion audio fluctuation;

[0052] It is determined whether there is a matching difference according to the playing type and the emotion type, and the video is recommended to be played according to the determination result.

[0053] It can be understood that the above-mentioned interactive method and system based on video content analysis have the same beneficial effects, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0054] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not intended to limit the scope of the application. Moreover, the same reference numerals are used throughout the same figures. In the drawings:

[0055] Figure 1 A functional block diagram of an interactive system based on video content analysis provided for an embodiment of the present application;

[0056] Figure 2 A flowchart of an interactive method based on video content analysis provided for an embodiment of the present application. DETAILED DESCRIPTION

[0057] Exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0058] In some embodiments of the present application, referring to Figure 1 An interactive system based on video content analysis, as shown in the figure, comprises:

[0059] The video analysis module is configured to acquire a playing type of a real-time playing video, acquire login information and compare the login information with a historical identity library, and determine a historical interaction strategy or an interaction generation strategy according to a comparison result.

[0060] The interaction analysis module is configured to determine an interaction mode based on ambient light and ambient sound when the interaction generation strategy is identified, and determine whether to re-determine the interaction mode based on the ambient light and the ambient sound when the historical interaction strategy is identified, the interaction mode including touch screen interaction, gesture interaction and voice interaction.

[0061] The interaction processing module is configured to acquire ambient water mist concentration and ambient particle concentration when the touch screen interaction or the gesture interaction is determined, determine whether to adjust shooting parameters based on the ambient water mist concentration and the ambient particle concentration, determine real-time facial data, and determine an emotion type based on the real-time facial data and a convolutional neural network model, acquire an audio signal and convert the audio signal into an emotional text when the voice interaction is determined, determine emotional audio fluctuation based on the audio signal and a deep learning model, and determine the emotion type according to the emotional text and the emotional audio fluctuation.

[0062] The interaction recommendation module is configured to determine whether there is a matching difference according to the playing type and the emotion type, and recommend the playing video according to a determination result.

[0063] Specifically, the video analysis module first obtains the playing type of the real-time playing video, which is the basic prerequisite for subsequent accurate recommendation of playing videos. Only by clearly defining the content field to which the real-time playing video belongs, such as education, current affairs, dance, etc., can the scope for recommendation direction be determined. At the same time, the video analysis module obtains the login information of the user and compares it with the historical identity library. The login information is the identity ID of the user login. For users without historical interaction records, the system uses an interaction generation strategy to ensure that new users can obtain adaptive interaction modes in a timely manner. For users with historical interaction records, a historical interaction strategy is adopted to make the interaction mode preferentially adapt to the user's past environmental interaction habits, thereby reducing the user's adaptation cost and avoiding interaction deviation caused by "one-size-fits-all". Whether the interaction generation strategy or the historical interaction strategy is used, the interaction analysis module determines the interaction mode based on the environmental light and the environmental sound. It includes touch screen interaction, gesture interaction, and voice interaction. Environmental factors will directly affect the user's interaction experience, and environmental factors usually cannot be directly changed. For example: the display case in the mall, environmental sound such as mall public broadcasting, noisy crowd sound, environmental light such as the brightness of the fixed spotlight at the top of the display case area, etc. These environmental factors will affect the recognition of the interaction mode. For example, gesture recognition is easily disturbed in a strong light environment, while touch screen interaction is more stable. In a noisy environment, the voice of the interaction is easily misidentified, while touch screen interaction or gesture interaction is more stable. For the historical interaction strategy, whether to determine the interaction mode again is determined according to the real-time environmental factors. Considering that the user's use environment may change, such as switching from indoors to outdoors, it ensures that the interaction mode is always dynamically adapted to the current environment. The interaction processing module processes differently according to different interaction modes. When it is determined to be touch screen interaction or gesture interaction, the interaction processing module obtains the environmental water mist concentration and the environmental particle concentration. This is because these two factors will affect the quality of the real-time facial data captured by high-definition cameras and other miniature image acquisition devices. For example, water mist causes the picture to be blurred, and particles cause image noise. Based on these data, the shooting parameters such as aperture size are adjusted to effectively ensure the clarity of the real-time facial data. Then the facial emotion features are extracted through a convolutional neural network model. The convolutional neural network model is good at image detail recognition and can adapt to the analysis needs of facial expressions to finally determine the emotion type. When it is determined to be voice interaction, the interaction processing module converts the audio signal into emotional text to obtain semantic-level emotional information. At the same time, the deep learning model analyzes the emotional audio fluctuations, such as rhythm features such as speech speed. The combination of the two can avoid the emotional misjudgment caused by single dependence on text or audio, effectively making up for the defects of video content interaction in emotion recognition, thereby achieving comprehensive emotion judgment.The interactive recommendation module takes the playing type and the emotional type as the core, judges the matching difference according to the playing type and the emotional type, for example, the user watches a serious educational video but the emotional type is irritability, there is a matching difference between the two, the recommendation direction needs to be adjusted, the recommendation is dynamically adjusted combined with the real-time emotional state of the user, the risk of traditional video content analysis and interaction being split is avoided, and the intelligent degree and flexibility of the recommendation result are ensured.

[0064] It can be understood that by comparing the login information with the historical identity library, the new user can obtain the adaptive interaction mode without complex setting, and the old user can continue the familiar interaction habit, ensuring the reliability of the interaction. The interaction analysis module dynamically adjusts the interaction mode based on environmental factors, avoiding the problem that the fixed interaction mode fails in special environments, such as gesture recognition failure in dim environment and voice non-response in noisy environment, improving the stability of video content and interaction adaptation. The interaction processing module adjusts the shooting parameters and multi-dimensional emotional judgment, thereby improving the emotional recognition accuracy, providing a reliable basis for recommended video, and the interaction recommendation module matches the emotional type of the user with the playing type of the playing video, effectively meeting the personalized needs of the user, thereby finding the video content that meets the interest and emotion in the current interaction in the massive video, and improving the intelligent degree of the recommendation result.

[0065] In some embodiments of the present application, when the login information is obtained and compared with the historical identity library, the historical interaction strategy or the interaction generation strategy is determined according to the comparison result, including: the historical identity library includes a plurality of historical login information and a plurality of historical interaction modes, and each historical login information and a historical interaction mode correspond to each other, when there is a historical login information same as the login information in the historical identity library, the video analysis module determines the historical interaction strategy, and when there is no historical login information same as the login information in the historical identity library, the video analysis module determines the interaction generation strategy.

[0066] Specifically, the historical identity library stores a plurality of historical login information and a plurality of historical interaction modes, and each historical login information forms a unique corresponding relationship with a historical interaction mode, that is, the user's past login record is bound to the interaction mode he has used. When the user logs in, the video analysis module first obtains the login information, and then compares it with all the historical login information in the historical identity library one by one. If there is historical login information in the historical identity library that is completely consistent with the current login information, it means that the user has a past record. The video analysis module determines to use the historical interaction strategy, that is, to use the historical interaction mode bound to the historical login information. If there is no historical login information in the historical identity library that is the same as the current login information, it means that the user is logging in for the first time or has no past record. The video analysis module determines to use the interaction generation strategy and generates a new interaction mode for the user, thereby improving the response efficiency of the system and ensuring the accuracy of the interaction and identity matching.

[0067] In some embodiments of the present application, when the interaction generation strategy is identified, the determination of the interaction mode based on the environment light and the environment sound includes: when the environment light is in the standard environment light range and the environment sound is not in the standard environment sound range, the interaction analysis module determines the interaction mode to be gesture interaction; when the environment light is not in the standard environment light range and the environment sound is in the standard environment sound range, the interaction analysis module determines the interaction mode to be voice interaction; and when the environment light is not in the standard environment light range and the environment sound is not in the standard environment sound range, the interaction analysis module determines the interaction mode to be touch screen interaction.

[0068] Specifically, when the interaction analysis module identifies the interaction generation strategy, two parameters of the current environment (ambient light and ambient sound) are detected, and the interaction analysis module is connected with the microphone array, image acquisition device, light sensor, etc. to determine whether the environmental factors are within the standard environment range corresponding to these devices. The interaction analysis module first calls the standard ambient light range and the standard ambient sound range as the judgment basis. The standard ambient light range is determined according to the requirements of the image acquisition device for image capture, and the standard ambient sound range is determined according to the requirements of the microphone array and other sound acquisition devices for sound capture. Then, it is compared one by one whether the current ambient light conforms to the standard ambient light range and whether the ambient sound conforms to the standard ambient sound range. If the detected ambient light is within the standard ambient light range, it means that the camera can capture clear images when capturing gestures or taking pictures, and the ambient sound is not within the standard ambient sound range, indicating that the background noise is high, and the voice is easily disturbed by noise, leading to recognition errors. The interaction analysis module excludes voice interaction prone to errors and determines the interaction mode as gesture interaction. If it is detected that the ambient light is not within the standard ambient light range, such as too dark or too bright, leading to blurred image details, gesture recognition is prone to misjudgment, and the ambient sound is within the standard ambient sound range, indicating that the background noise is low, and the voice can be accurately captured and analyzed. The interaction analysis module excludes gesture interaction prone to errors and determines the interaction mode as voice interaction. If it is detected that the ambient light is not within the standard ambient light range and the ambient sound is also not within the standard ambient sound range, gesture interaction is ineffective due to poor image quality, and voice interaction is ineffective due to noise interference. The interaction analysis module then selects touch screen interaction as the interaction mode, which is least affected by environmental factors.

[0069] It can be understood that different interaction modes have different environmental conditions. Gesture interaction relies on clear visual images, voice interaction relies on a sound environment with as little interference as possible, and touch screen interaction has lower requirements for light and sound and is almost unaffected by extreme conditions. Based on the judgment of the standard range of ambient light and ambient sound, the short board of each interaction mode in dealing with the environment is avoided, and the failure of interaction due to environmental mismatch is avoided, ensuring that the interaction mode always matches the current environmental conditions.

[0070] In some embodiments of the present application, when the historical interaction strategy is identified, the judgment of whether to re-determine the interaction mode based on the ambient light and the ambient sound includes: the interaction analysis module obtains and logs the historical interaction mode corresponding to the same historical login information as the login information, and records it as the to-be-determined interaction mode. When the to-be-determined interaction mode does not meet the corresponding standard environmental conditions, the to-be-determined interaction mode is cancelled, and the interaction mode is re-determined. When the to-be-determined interaction mode meets the corresponding standard environmental conditions, the interaction mode is determined according to the to-be-determined interaction mode. The standard environmental conditions are the standard ambient light range and the standard ambient sound range.

[0071] Specifically, if the to-be-determined interaction mode does not meet the standard environment condition, for example, the to-be-determined interaction mode is gesture interaction, but the current environment light is not within the standard environment light range, the accuracy of gesture recognition cannot be guaranteed, the interaction analysis module cancels the to-be-determined interaction mode, and determines the adaptive interaction mode again according to the current environment light and environment sound. When the to-be-determined interaction mode is voice interaction, the same applies. The historical interaction mode is the adaptive selection of the user in a specific environment in the past. In the case that the current use environment of the user may change, if the historical interaction mode that does not match the current environment is directly used, there is a risk of interaction failure. Through the judgment of the standard environment condition, the familiar interaction mode of the user can be retained (when the environment is adaptive), and the interaction problem caused by the change of the environment can be avoided (when the environment is not adaptive, the interaction mode is determined again), effectively balancing the user's use habit and the adaptability of the environment, while reducing unnecessary adjustment of the interaction mode, and improving the response efficiency of the system.

[0072] In some embodiments of the present application, when it is determined to be touch screen interaction or gesture interaction, the environment water mist concentration and the environment particle concentration are obtained, and it is determined whether to adjust the shooting parameter based on the environment water mist concentration and the environment particle concentration. When determining the real-time facial data, the interaction processing module obtains the environment water mist concentration and the environment particle concentration, and performs data standardization on the environment water mist concentration and the environment particle concentration. The target environment water mist concentration and the target environment particle concentration are determined, and the environment visibility is determined according to the target environment water mist concentration and the target environment particle concentration. When the environment visibility is greater than or equal to the environment visibility threshold, it is determined that the shooting parameter is not adjusted, and the current shooting parameter is used to determine the real-time facial data. When the environment visibility is less than the environment visibility threshold, it is determined that the shooting parameter is adjusted, and the exposure compensation is increased / the aperture is increased. The real-time facial data is determined according to the adjusted shooting parameter.

[0073] Specifically, since the collection of real-time facial data needs to capture facial features, the required environmental conditions are more demanding. When it is determined to be a touch screen interaction or a gesture interaction, the interaction processing module obtains the current environmental water mist concentration and the current environmental particle concentration, such as the water mist formed by the air conditioner condensate in the showcase of the shopping mall and the dust concentration generated by the flow of personnel, from devices such as water mist detection sensors and laser particle sensors. Since the original data units and numerical ranges of the two types of concentrations may differ in different scenarios, the data is standardized to eliminate the differences in dimensions, thereby determining the target environmental water mist concentration and the target environmental particle concentration. The environmental visibility is 1 minus the weighted sum of the target environmental water mist concentration and the target environmental particle concentration. The higher the target environmental water mist concentration and the target environmental particle concentration, the lower the environmental visibility. Conversely, the lower the target environmental water mist concentration and the target environmental particle concentration, the higher the environmental visibility. The environmental visibility directly reflects the clarity of the camera's shot picture in the current environment. The environmental visibility threshold is dynamically determined according to the actual environment in which the system is located (the size of the area, the number of people, etc.). The embodiment preferably is 0.5. If the environmental visibility is greater than or equal to the environmental visibility threshold, it means that the current environment has less impact on shooting, and the camera can capture a clear facial picture. Then it is determined not to adjust the shooting parameters (such as aperture size, exposure compensation, focus mode, etc.), and the current shooting parameters are used to collect and determine the real-time facial data. If the environmental visibility is less than the environmental visibility threshold, it means that water mist or particles cause the picture to be blurred and the number of noise points to increase, which will affect the quality of the real-time facial data. Then it is determined to adjust the shooting parameters. The specific adjustment is to increase the exposure compensation to improve the overall brightness of the picture to offset the picture darkness caused by water mist or to increase the aperture to increase the amount of light entering the lens to reduce the noise points caused by particles. After the shooting parameters are adjusted, the real-time facial data is collected according to the adjusted shooting parameters. By adjusting the exposure compensation or the aperture, the accuracy and reliability of the collection of real-time facial data in touch screen interaction or gesture interaction are ensured, thereby ensuring the stability of the recommended results.

[0074] In some embodiments of the present application, when determining the emotion type based on the real-time facial data and the convolutional neural network model, the interaction processing module acquires a facial sample set, samples the facial sample set to determine a training set and a test set, pre-selects an initial convolutional neural network model, iteratively trains the initial convolutional neural network model according to the training set, verifies the initial convolutional neural network model after iterative training according to the test set, if the ternary loss of the initial convolutional neural network model after the current iterative training is greater than the ternary loss of the initial convolutional neural network model after the previous iterative training, gradient clipping is performed on the initial convolutional neural network model after the current iterative training, and iterative training is continued, if the ternary loss of the initial convolutional neural network model after the current iterative training is less than or equal to the ternary loss of the initial convolutional neural network model after the previous iterative training, the iterative training is stopped, the initial convolutional neural network model after the current iterative training is determined as the convolutional neural network model, and the real-time facial data is substituted into the convolutional neural network model to determine the emotion type.

[0075] Specifically, the face sample set includes face data of different age groups such as children, young people, middle-aged and old people, different genders, different skin colors, different facial features such as wearing glasses and growing a beard, different facial poses such as front face, side face, slightly low head, slightly high head, and rich expression detail samples under different emotion types (such as joy, irritability, calmness, seriousness, etc.), for example, the emotion of "joy" includes different degrees of joyful expressions such as laughing, smiling, and slightly raising the corners of the mouth, and the emotion of "irritability" includes subtle expressions such as frowning, turning up the nose, and tight eyes, to ensure that the model can learn the differentiated facial features of different emotions. The sample is divided into a training set and a test set, the training set is used for model learning of emotion types, and the test set is used for verification of model recognition effect, and the sampling ratio is usually 4:1, then an initial convolutional neural network model is preselected, the initial convolutional neural network model is iteratively trained according to the training set, so that the model gradually learns the corresponding relationship between facial features and emotion types, the model performance is verified by using the test set after each round of iterative training, the optimization direction of the model is judged by comparing the ternary loss, the ternary loss is used to measure the accuracy of the model in distinguishing emotion types, if the ternary loss of the initial convolutional neural network model after the current iterative training is greater than the ternary loss of the initial convolutional neural network model after the previous iterative training, it indicates that the model has not reached a stable performance level, then the gradient clipping is performed on the current model to limit the gradient value range, to avoid out-of-control training, and then the iterative training is continued until the ternary loss of the initial convolutional neural network model after the iterative training is less than or equal to the ternary loss of the initial convolutional neural network model after the previous iterative training. If the ternary loss of the initial convolutional neural network model after the current iterative training is less than or equal to the ternary loss of the initial convolutional neural network model after the previous iterative training, it indicates that the performance of the model is rising or has stabilized, then the iterative training can be stopped, the initial convolutional neural network model after the current iterative training is determined as the convolutional neural network model, and the real-time face data is substituted into the convolutional neural network model to determine the emotion type, to ensure that the recommended content can fit the user's mood, thereby improving the automation and intelligence level of the system.

[0076] In some embodiments of the present application, when it is determined to be a voice interaction, the audio signal is acquired and converted into emotional text, including: the interaction processing module acquires an initial audio signal and performs noise reduction to determine an audio signal, the interaction processing module converts the audio signal into emotional text and determines the emotional category of each noun and the emotional category of each adjective in the emotional text, the emotional category includes positive emotion, negative emotion and neutral emotion, the positive number ratio of the total positive emotion quantity to the total emotion quantity is determined, the negative number ratio of the total negative emotion quantity to the total emotion quantity is determined, the neutral number ratio of the total neutral emotion quantity to the total emotion quantity is determined, the positive number ratio, the negative number ratio and the neutral number ratio are arranged in descending order, and the emotional orientation of the emotional text is determined according to the arrangement result.

[0077] Specifically, when it is determined to be a voice interaction, the interaction processing module first acquires an initial audio signal containing the user's voice and environmental noise such as a shopping mall background sound, filters the noise of the non-voice segment through noise reduction processing to eliminate invalid interference, and obtains a pure audio signal to avoid subsequent conversion distortion caused by noise. The audio signal is converted into an emotional text through audio-to-text technology, and all nouns and adjectives in the emotional text are extracted. Nouns and adjectives are the core carriers of emotional information. They are compared with an emotional library to identify the corresponding emotional categories. The emotional library is updated online at regular intervals, and the frequency mechanism of regular updates is set to ensure the timeliness and accuracy of the content of the emotional library. The interaction processing module can accurately identify and analyze the emotional categories in the emotional text, thereby capturing the latest emotional expression methods and vocabulary changes. In determining the emotional category of each noun and the emotional category of each adjective (positive emotion, negative emotion, and neutral emotion), the sum of the total number of positive emotions, the total number of negative emotions, and the total number of neutral emotions is determined as the total number of emotions. The positive number ratio, the negative number ratio, and the neutral number ratio are calculated respectively. Finally, they are sorted in descending order. If a certain number ratio is in the first place in the sorting, the emotional orientation of the emotional text is consistent with the emotional category corresponding to the number ratio (such as the maximum positive number ratio, the emotional orientation is positive). If the number ratios of two or three are in the first place in the sorting at the same time, it indicates that the emotional orientation of the emotional text is multi-emotion, indicating that the user's mood is not in a stable state, and therefore no video is recommended for playing.

[0078] In some embodiments of the present application, when determining the emotional audio fluctuation based on the audio signal and the deep learning model, and determining the emotional type according to the emotional text content and the emotional audio fluctuation, it includes: the interaction processing module determines the time domain image based on the audio signal and the deep learning model, acquires all adjacent amplitude intervals in the time domain image, counts the first amplitude interval number whose amplitude interval is greater than the amplitude interval threshold, and counts the second amplitude interval number whose amplitude interval is less than or equal to the amplitude interval threshold. When the first amplitude interval number is greater than the second amplitude interval number, it is determined that the emotional audio fluctuation is negative fluctuation. When the first amplitude interval number is less than the second amplitude interval number, it is determined that the emotional audio fluctuation is positive fluctuation. When the first amplitude interval number is equal to the second amplitude interval number, it is determined that the emotional audio fluctuation is neutral fluctuation. The amplitude interval threshold is determined according to the speaking speed of the audio signal. The emotional type is determined based on the emotional orientation and the emotional audio fluctuation.

[0079] Specifically, the interaction processing module inputs the acquired audio signal into the deep learning model, extracts the time domain features of the audio such as the strength variation of the sound, the signal fluctuation in the time dimension, and generates a time domain image from the deep learning model, which intuitively presents the amplitude state of the audio signal changing with time. The training process of the deep learning model is consistent with that of the initial convolutional neural network model, which will not be repeated here. Then, all adjacent amplitude intervals, i.e. the intervals of the audio signals at adjacent time points, are extracted from the time domain image, which reflects the degree of audio fluctuation. The amplitude interval threshold is determined according to the speech rate of the audio signal. Speech rate variation is an important way of expressing emotions. Accordingly, the emotional state conveyed by the user is determined. The slower the speech rate, the more likely the user's emotional state is sad. The longer the amplitude interval, the more the first amplitude interval. The faster the speech rate, the more likely the user's emotional state is happy. The more cheerful the conversation, the less the pause time, resulting in shorter amplitude intervals. The more the second amplitude interval, the emotional audio fluctuation is dynamically determined according to the comparison between the first amplitude interval and the second amplitude interval. The final emotional type is determined by combining the emotional orientation (e.g. positive, negative, neutral) determined by the emotional text and the current determined emotional audio fluctuation. For example, if the emotional orientation is positive and the audio fluctuation is positive, the emotional type is positive and happy. When there is a difference between the emotional orientation and the emotional audio fluctuation, such as a negative emotional orientation and a positive audio fluctuation, it also indicates that the user's mood is not stable at this time, so no video is recommended for playing. By combining the emotional orientation of the emotional text and the emotional audio fluctuation, the risk of relying solely on the emotional text and ignoring the emotional audio fluctuation is effectively avoided. The dual-dimension combination of the emotional orientation and the emotional audio fluctuation improves the accuracy of emotional type judgment, making the emotional recognition in the interactive scene more efficient and stable, thereby ensuring that the recommended content can match the user's real emotions. At the same time, the whole process does not require human intervention, improving the automation level of the system.

[0080] In some embodiments of the present application, when determining whether there is a difference in matching between the playing type and the emotional type, and recommending playing a video according to the determination result, it includes: when there is a difference in matching between the playing type and the emotional type, a video is recommended for playing according to the emotional type; and when there is no difference in matching between the playing type and the emotional type, a video is recommended for playing according to the playing type.

[0081] Specifically, the interactive recommendation module first calls the determined playing type of the real-time playing video, such as education, current affairs, dance, etc., and the interactive processing module outputs the emotional type, such as joy, irritability, neutrality, etc. Whether the content atmosphere of the playing type and the user state of the emotional type match is used to judge the matching difference. For example, the playing type is a serious education video, and the emotional type is "irritability", which indicates that the current state of the user is not suitable for serious content, so it is determined that there is a matching difference. Then, according to the emotional type, a light and entertaining playing video is recommended instead of continuing to recommend serious content. If the playing type is a light and entertaining video, and the emotional type is "joy", which indicates that the state of the user matches the content atmosphere, it is determined that there is no matching difference, and the playing video is recommended according to the playing type, that is, the current playing type is continued, and the same style and same field of entertainment video content is recommended. Through the matching of the form of video analysis and the interactive process, the emotional type is dynamically identified, the matching difference is judged, the "content type" and "real-time emotion" are effectively considered, so as to ensure that the user can find the content meeting the demand in a large number of videos, and the intelligent level and flexibility of the recommendation result are ensured.

[0082] In summary, the beneficial effects of the present application are that: the video analysis module matches the historical interaction strategy for old users and generates the interaction strategy for new users by comparing the login information with the historical identity library, ensuring the pertinence of the determined interaction mode, the interactive analysis module dynamically determines or adjusts the interaction mode based on the environmental light and environmental sound, switches to voice interaction in dim light, and switches to gesture interaction in noisy scenes, ensuring the adaptability of the interaction mode to the real-time playing video, thereby improving the compatibility of the system to the environment, the interactive processing module accurately perceives the emotional type for different interaction modes, in touch screen interaction or gesture interaction, the shooting parameters under the influence of the environmental mist concentration and the environmental particle concentration are adjusted to ensure the accuracy of the real-time facial data, and in voice interaction, the emotional type is determined through the dual dimensions of emotional text and emotional audio fluctuation, realizing the dynamic capture of the real-time state of different interaction modes, laying a data foundation for recommending playing videos, thereby avoiding the risk of feedback lag, recommending playing videos according to the matching difference between the emotional type and the playing type, ensuring the matching degree of the form of video analysis and the interactive process, and ensuring the intelligent level and flexibility of the recommendation result.

[0083] In another preferred mode based on the above embodiment, referring to Figure 2 The present embodiment provides an interactive method based on video content analysis, which is used for the above-mentioned interactive system based on video content analysis, comprising:

[0084] S100: Obtain the playing type of the real-time playing video, obtain the login information and compare the login information in the historical identity library, and determine the historical interaction strategy or the interaction generation strategy according to the comparison result.

[0085] S200: when the interaction generation strategy is identified, determining the interaction mode based on the ambient light and the ambient sound, and when the historical interaction strategy is identified, determining whether to re-determine the interaction mode based on the ambient light and the ambient sound, the interaction mode including a touch screen interaction, a gesture interaction, and a voice interaction.

[0086] S300: when it is determined to be the touch screen interaction or the gesture interaction, acquiring an ambient water mist concentration and an ambient particle concentration, determining whether to adjust the shooting parameter based on the ambient water mist concentration and the ambient particle concentration, determining the real-time facial data, and determining the emotion type based on the real-time facial data and a convolutional neural network model, and when it is determined to be the voice interaction, acquiring an audio signal and converting the audio signal into an emotional text, determining an emotional audio fluctuation based on the audio signal and a deep learning model, and determining the emotion type according to the emotional text and the emotional audio fluctuation.

[0087] S400: determining whether there is a matching difference according to the playing type and the emotion type, and recommending a playing video according to the determination result.

[0088] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer usable program code.

[0089] The present application is described with reference to flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The devices that implement the functions specified in one or more flows and / or blocks.

[0090] These computer program instructions can also be stored in a computer readable storage medium that can guide the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable storage medium produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1the function specified in the one or more blocks.

[0091] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, so that the instructions executed on the computer or other programmable devices provide processes for implementing the flow Figure 1 one or more flows and / or blocks Figure 1 the steps of the function specified in the one or more blocks.

[0092] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit it, although the above embodiments of the present application have been described in detail, those skilled in the art should understand: the specific embodiments of the present application can be modified or replaced by the same, without departing from the spirit and scope of the present application, any modification or equivalent replacement, which should be covered in the protection scope of the claims of the present application.

Claims

1. An interactive system based on video content analysis, characterized in that, include: The video parsing module is configured to obtain the playback type of the real-time video, obtain login information and compare the login information with the historical identity database, and determine the historical interaction strategy or interaction generation strategy based on the comparison result. The interaction analysis module is configured to determine the interaction method based on ambient light and ambient sound when the interaction generation strategy is identified, and to determine whether to re-determine the interaction method based on ambient light and ambient sound when the historical interaction strategy is identified. The interaction method includes touch screen interaction, gesture interaction and voice interaction. The interaction processing module is configured to, when determined to be a touch screen interaction or gesture interaction, acquire the ambient water mist concentration and ambient particle concentration, determine whether to adjust the shooting parameters based on the ambient water mist concentration and ambient particle concentration, determine real-time facial data, and determine the emotion type based on the real-time facial data and a convolutional neural network model; when determined to be a voice interaction, acquire the audio signal and convert it into emotional text, determine the emotional audio fluctuation based on the audio signal and a deep learning model, and determine the emotion type based on the emotional text and emotional audio fluctuation. The interactive recommendation module is configured to determine whether there is a matching difference based on the playback type and emotion type, and recommend videos to play based on the determination result; When the interaction is identified as a voice interaction, acquiring the audio signal and converting it into emotional text includes: The interactive processing module acquires the initial audio signal and performs noise reduction to determine the audio signal; The interactive processing module converts the audio signal into emotional text and determines the emotional category of each noun and each adjective in the emotional text. The emotional categories include positive emotions, negative emotions, and neutral emotions. Determine the ratio of the total number of positive emotions to the total number of positive emotions, the ratio of the total number of negative emotions to the total number of negative emotions, and the ratio of the total number of neutral emotions to the total number of neutral emotions. The positive, negative, and neutral quantity ratios are arranged in descending order, and the emotional orientation of the emotional text is determined based on the arrangement result. When determining emotional audio fluctuations based on the audio signal and a deep learning model, and determining the emotion type based on the emotional text and emotional audio fluctuations, the process includes: The interactive processing module determines the time-domain image based on the audio signal and the deep learning model, and obtains all adjacent amplitude intervals in the time-domain image; The number of first amplitude intervals with amplitude intervals greater than the amplitude interval threshold is counted, and the number of second amplitude intervals with amplitude intervals less than or equal to the amplitude interval threshold is counted. When the number of the first amplitude intervals is greater than the number of the second amplitude intervals, the emotional audio fluctuation is determined to be a negative fluctuation. When the number of the first amplitude intervals is less than the number of the second amplitude intervals, the emotional audio fluctuation is determined to be a positive fluctuation. When the number of the first amplitude intervals is equal to the number of the second amplitude intervals, the emotional audio fluctuation is determined to be a neutral fluctuation. The amplitude interval threshold is determined based on the speech rate of the audio signal, and the emotion type is determined based on the emotional orientation and the emotional audio fluctuation.

2. The interactive system based on video content analysis according to claim 1, characterized in that, When acquiring login information and comparing it with a historical identity database, and determining a historical interaction strategy or interaction generation strategy based on the comparison result, the process includes: The historical identity database includes several historical login records and several historical interaction methods, and each historical login record corresponds to a historical interaction method. When the historical identity database contains historical login information that is identical to the login information, the video parsing module determines it as the historical interaction strategy. When there is no historical login information in the historical identity database that is the same as the login information, the video parsing module determines the interaction generation strategy.

3. The interactive system based on video content analysis according to claim 2, characterized in that, When the interaction generation strategy is identified, determining the interaction method based on ambient lighting and ambient sound includes: When the ambient light is within the standard ambient light range and the ambient sound is outside the standard ambient sound range, the interaction analysis module determines the interaction method as the gesture interaction. When the ambient light is outside the standard ambient light range and the ambient sound is within the standard ambient sound range, the interaction analysis module determines the interaction method as the voice interaction. When the ambient light is outside the standard ambient light range and the ambient sound is outside the standard ambient sound range, the interaction analysis module determines the interaction method as the touch screen interaction.

4. The interactive system based on video content analysis according to claim 3, characterized in that, When the historical interaction strategy is identified, determining whether to re-determine the interaction method based on the ambient lighting and ambient sound includes: The interaction analysis module obtains the historical interaction methods corresponding to the same historical login information as the login information, and records them as the interaction methods to be determined. If the interaction method to be determined does not meet the corresponding standard environmental conditions, then the interaction method to be determined is cancelled and the interaction method is re-determined. When the interaction method to be determined meets the corresponding standard environmental conditions, the interaction method is determined according to the interaction method to be determined. The standard environmental conditions are the standard ambient light range and the standard ambient sound range.

5. The interactive system based on video content analysis according to claim 4, characterized in that, When the interaction is determined to be a touchscreen interaction or gesture interaction, the ambient water mist concentration and ambient particle concentration are acquired. Based on the ambient water mist concentration and ambient particle concentration, it is determined whether to adjust the shooting parameters. When determining real-time facial data, the following is included: The interactive processing module acquires the environmental water mist concentration and environmental particle concentration, performs data standardization on the environmental water mist concentration and environmental particle concentration, determines the target environmental water mist concentration and target environmental particle concentration, and determines the environmental visibility based on the target environmental water mist concentration and target environmental particle concentration. When the environmental visibility is greater than or equal to the environmental visibility threshold, it is determined that the shooting parameters will not be adjusted, and the real-time facial data will be determined based on the current shooting parameters. When the ambient visibility is less than the ambient visibility threshold, it is determined that the shooting parameters should be adjusted. The adjustment is to increase exposure compensation / increase the aperture, and the real-time facial data is determined based on the adjusted shooting parameters.

6. The interactive system based on video content analysis according to claim 5, characterized in that, When determining the emotion type based on the real-time facial data and the convolutional neural network model, the following are included: The interactive processing module acquires a facial sample set and samples the facial sample set to determine the training set and test set. An initial convolutional neural network model is pre-selected, and the initial convolutional neural network model is iteratively trained based on the training set. The initial convolutional neural network model after iterative training is then validated based on the test set. If the ternary loss of the initial convolutional neural network model after the current iteration is greater than the ternary loss of the initial convolutional neural network model after the previous iteration, then gradient pruning is performed on the initial convolutional neural network model after the current iteration, and iterative training continues. If the ternary loss of the initial convolutional neural network model after the current iteration is less than or equal to the ternary loss of the initial convolutional neural network model after the previous iteration, then the iterative training is stopped, and the initial convolutional neural network model after the current iteration is determined as the convolutional neural network model. The real-time facial data is then substituted into the convolutional neural network model to determine the emotion type.

7. The interactive system based on video content analysis according to claim 6, characterized in that, When determining whether there is a matching difference based on the playback type and emotion type, and recommending videos to play based on the determination result, the process includes: When there is a discrepancy between the matching between the playback type and the emotion type, a video is recommended to be played based on the emotion type. When there is no difference in the matching between the playback type and the emotion type, the video to be played is recommended based on the playback type.

8. An interactive method based on video content analysis, used to apply the interactive system based on video content analysis as described in any one of claims 1-7, characterized in that, include: Obtain the playback type of the real-time video, obtain login information and compare the login information with the historical identity database, and determine the historical interaction strategy or interaction generation strategy based on the comparison result; When the interaction generation strategy is identified, the interaction method is determined based on ambient light and ambient sound. When the historical interaction strategy is identified, it is determined whether to re-determine the interaction method based on ambient light and ambient sound. The interaction method includes touch screen interaction, gesture interaction and voice interaction. When the interaction is determined to be a touch screen interaction or a gesture interaction, the ambient water mist concentration and ambient particle concentration are acquired. Based on the ambient water mist concentration and ambient particle concentration, it is determined whether to adjust the shooting parameters, real-time facial data is determined, and the emotion type is determined based on the real-time facial data and a convolutional neural network model. When the interaction is determined to be a voice interaction, the audio signal is acquired and converted into emotional text. Based on the audio signal and a deep learning model, the emotional audio fluctuation is determined, and the emotion type is determined based on the emotional text and the emotional audio fluctuation. Based on the playback type and emotion type, determine whether there is a matching difference, and recommend videos to play based on the determination result.

Citation Information

Patent Citations

  • Multi-mode interactive intelligent control system

    CN118226967A

  • Intelligent video content analysis and interaction system

    CN119342296A