Intelligent optimization system and method integrating video resource scheduling and voice emotion recognition
Through an intelligent optimization system that integrates video resource scheduling and voice emotion recognition, users' emotions are perceived in real time and video resources are dynamically adjusted, the problem that existing systems cannot perceive user emotions is solved, and a high-quality and personalized multimedia service experience is achieved.
Patent Information
- Application Number
- CN202510373564.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing video service system cannot effectively perceive and respond to the emotional state of users, resulting in the inability to personalize resource allocation and video content adjustments according to user needs, affecting the user experience.
Design an intelligent optimization system that integrates video resource scheduling and voice emotion recognition. Through the voice emotion recognition module, users' voice streams are collected in real time, acoustic features are extracted and emotional classification is performed, and emotional labels are output; the video resource scheduling module dynamically adjusts video encoding parameters, transmission protocols and rendering priority based on emotional labels; the multi-modal decision-making engine integrates voice emotion labels, network status data and user historical behavior data to generate real-time resource allocation strategies.
It realizes personalized resource allocation and content adjustment, improves the personalization and intelligence of user experience, ensures smooth video playback, good picture quality and sound quality, optimizes resource utilization, improves emotional matching and QoS scores, and enhances system adaptability.
Smart Images

Figure CN120151548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of artificial intelligence and multimedia communication, and more specifically to an intelligent optimization system and method that integrates video resource scheduling and speech emotion recognition. Background Art
[0002] With the rapid development of Internet technology and multimedia applications, users have higher and higher requirements for the quality of video services. During video playback, the instability of the network state and the limitations of device resources often lead to problems such as video stuttering and blurred picture quality, seriously affecting the user experience. At the same time, existing video service systems often lack the perception and response to the user's emotional state, and cannot perform personalized resource allocation and video content adjustment according to the user's emotional needs.
[0003] Traditional video resource scheduling mainly makes static or simple dynamic adjustments based on network parameters and device performance, and fails to fully consider the user's subjective feelings. Although speech emotion recognition technology has made certain progress, there are still deficiencies in its deep integration and practical application in multimedia application scenarios. For example, in scenarios such as online education, video conferencing, and film and television entertainment, if video resources can be optimized in real - time according to the user's speech emotion, the interaction effect and user satisfaction will be greatly improved. Therefore, developing an intelligent optimization system and method that can integrate video resource scheduling and speech emotion recognition has important practical significance and application value. Summary of the Invention
[0004] In order to overcome the above - mentioned defects of the prior art, the present invention provides an intelligent optimization system and method that integrates video resource scheduling and speech emotion recognition to solve the problems existing in the above - mentioned background art.
[0005] The present invention provides the following technical solutions: An intelligent optimization system that integrates video resource scheduling and speech emotion recognition, comprising:
[0006] A speech emotion recognition module: used to collect the user's speech stream in real - time, extract acoustic features through a deep - learning model and perform emotion classification, and output emotion labels;
[0007] A video resource scheduling module: used to monitor the network state and device resources, and dynamically adjust video encoding parameters, transmission protocols, and rendering priorities according to the emotion labels;
[0008] A multi - modal decision - making engine: based on a reinforcement - learning model, integrating speech emotion labels, network state data, and user historical behavior data to generate real - time resource allocation strategies.
[0009] Further, the speech emotion recognition module includes:
[0010] The acoustic feature extraction unit adopts a hybrid architecture of Transformer and CNN to perform spatio-temporal feature fusion on the MFCC features and spectrograms of speech signals;
[0011] The emotion classification unit constructs a classifier based on a pre-trained emotion database (such as RAVDESS, IEMOCAP) and outputs the emotion category and confidence.
[0012] Furthermore, the speech emotion recognition module integrates lightweight model compression technologies, including at least one of knowledge distillation and quantization-aware training, and the model inference latency is less than 50ms.
[0013] Furthermore, the dynamic adjustment of the video resource scheduling module includes:
[0014] Real-time monitoring of network bandwidth fluctuations, device GPU load, and video resolution requirements;
[0015] Establish an emotion-resource mapping strategy table to define the correspondence between emotion categories and video parameters, including:
[0016] Angry emotion: Trigger the high frame rate mode (≥60fps) and the low latency transmission protocol (UDP priority);
[0017] Sad emotion: Trigger the high color saturation mode and enhanced audio bandwidth allocation.
[0018] Furthermore, the emotion-resource mapping strategy table further includes an emotion urgency level classification. High-urgency emotions (such as anxiety, anger) trigger preemptive resource allocation and suspend the GPU occupancy of low-priority tasks.
[0019] Furthermore, the multi-modal decision engine includes:
[0020] Input layer: Receive speech emotion labels, network throughput, packet loss rate, and user historical interaction behaviors (such as the frequency of manual resolution adjustment);
[0021] Output layer: Generate video encoding parameters (H.265 / AV1 bitrate), transmission protocol weights (UDP / TCP mixing ratio), and rendering resource allocation ratio;
[0022] Decision model: Trained based on the Proximal Policy Optimization (PPO) algorithm, and the reward function includes emotion matching degree, QoS score, and resource utilization rate.
[0023] Furthermore, the training process of the decision model includes:
[0024] Construct a simulation environment to simulate different network states and emotion scenarios;
[0025] Iteratively optimize the policy through offline reinforcement learning. The model update period is 24 hours, and incremental learning is used to adapt to new user data.
[0026] Provide an intelligent optimization method that integrates video resource scheduling and speech emotion recognition, including the following steps:
[0027] S1: Real-time collect the user's speech stream, perform emotion classification through a lightweight deep learning model, and generate emotion labels and confidence levels.
[0028] S2: Query the emotion-resource mapping policy table according to the emotion label to determine the priority of video parameter adjustment.
[0029] S3: Combine the current network bandwidth, device load, and priority list to generate resource allocation instructions through a multimodal decision-making engine.
[0030] S4: Dynamically adjust the video encoder bitrate, transmission protocol, and GPU rendering resources, and monitor the user's feedback behavior to update the policy.
[0031] Furthermore, the dynamic adjustment described in step S4 includes:
[0032] When the user's emotion is "confused", increase the video source resolution to 4K and reduce the background special effect rendering accuracy.
[0033] When the user's emotion is "calm", enable the bandwidth-saving mode, limit the peak bitrate, and delay the transmission of non-critical frames.
[0034] The technical effects and advantages of the present invention:
[0035] Through speech emotion recognition, the system of the present invention realizes personalized resource allocation and content adjustment, greatly improving the personalization and intelligence of the user experience; the video resource scheduling module dynamically adjusts parameters according to the network, device, and emotion to ensure smooth video playback with excellent picture quality and sound quality; the multimodal decision-making engine integrates multiple data to generate strategies, optimizes resource utilization, improves the emotion matching degree and QoS score; the intelligent optimization method has a clear process, accurately optimizes video resources, continuously updates the policy through user feedback, enhances the system's self-adaptability, and provides users with a high-quality, personalized multimedia service experience in all aspects.
[0036] Specifically: Enhancement of personalized service experience: The present invention can perceive the user's emotion in real time through the voice emotion recognition module, enabling the video service system to break through the traditional single mode and perform personalized resource allocation and video content adjustment according to the user's emotional state. For example, in the online education scenario, when it is recognized that the student is in a confused emotion, the system automatically increases the video source resolution to 4K and reduces the background special effect rendering accuracy, allowing the student to see the key details of the teaching content more clearly, thereby greatly enhancing the personalization and intelligence of the user experience and increasing the user's satisfaction and dependence on the video service.
[0037] Optimization of video playback quality: The video resource scheduling module closely combines the network state and device resources and dynamically adjusts video parameters based on emotion tags. In the case of network bandwidth fluctuations or high device GPU loads, appropriate responses are made according to different emotional needs. For example, when it is detected that the user is in an angry mood, the high frame rate mode (≥60fps) and low latency transmission protocol (UDP priority) are triggered to ensure the smoothness and timeliness of video playback, avoid stuttering, and effectively improve the smoothness, picture quality, and sound quality of video playback, ensuring a good viewing experience for users under different network environments and device conditions.
[0038] Enhancement of resource utilization efficiency and system performance: The multi-modal decision engine integrates multiple data and generates real-time resource allocation strategies through reinforcement learning. By comprehensively analyzing voice emotion tags, network state data, and user historical behavior data, the resource utilization efficiency is continuously optimized. Taking the video conferencing scenario as an example, the decision model is trained based on the proximal policy optimization (PPO) algorithm, and the reward function includes emotion matching degree, QoS score, and resource utilization rate, which can reasonably allocate video encoding parameters (H.265 / AV1 bit rate), transmission protocol weights (UDP / TCP mixing ratio), and rendering resource allocation ratio, improve the emotion matching degree and service quality score (QoS), and thus enhance the overall performance and adaptability of the system, better meeting diverse application requirements.
[0039] System adaptability and continuous optimization: The intelligent optimization method has a clear process and each step is closely coordinated, enabling real-time and accurate optimization of video resources. And by listening to the user's feedback behavior, the strategy is continuously updated. For example, in the film and television entertainment scenario, if the system finds that the user frequently manually adjusts the resolution, it indicates that the current resource allocation strategy may not be accurate enough. At this time, the system will retrain and optimize the decision model based on the user behavior data to further enhance the system's adaptability, enabling it to continuously adapt to new user needs and changing network environments and other factors, and always providing high-quality video services for users. Description of the Drawings
[0040] Figure 1 It is the system architecture diagram in the present invention;
[0041] Figure 2 This is the flowchart of the method in the present invention. Specific embodiments
[0042] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.
[0044] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0045] Embodiment 1: Online education scenario
[0046] System construction
[0047] Construction of the voice emotion recognition module
[0048] Acoustic feature extraction unit: Select a hybrid architecture model of Transformer and CNN pre-trained on a voice dataset related to the education field. According to the voice characteristics in the online education scenario, the model parameters are fine-tuned. For example, the intermediate dimension of the feed-forward neural network layer in the Transformer is adjusted to 2048 to better capture the complex features of the voice in the education scenario. The model is trained using the massive teacher-student interaction voice data accumulated by the online education platform. After 50 epochs of training, the fusion accuracy of the MFCC features and the spatio-temporal features of the spectrogram of the voice signal in the education scenario reaches 93%.
[0049] Emotion classification unit: Based on a dedicated emotion database in the education field (such as a database containing the emotional expressions of students when learning different knowledge points), select a deep neural network (DNN) as the classifier model. During the training process, the Adagrad optimizer with a learning rate of 0.0005 is used. After 30 rounds of training, the emotion category recognition accuracy of the model on the education scenario test set reaches 90%. Through quantization-aware training technology, the weights and activation values of the model are quantized to 8-bit integers, and the inference latency of the model is reduced to 40 ms, and the accuracy only drops by 1%.
[0050] Video Resource Scheduling Module Setup
[0051] Network Status and Device Resource Monitoring Section: In an online education platform where 5,000 students are studying online simultaneously, network monitoring tools are used to obtain real-time network bandwidth fluctuation data. During the peak period of course live streaming, the network bandwidth fluctuation range is between 3Mbps and 18Mbps. The API provided by the device driver is used to monitor the GPU load of students' devices. It is found that when students use high-definition tablet devices to watch video courses and perform operations such as taking notes, the GPU load varies between 20% and 70%. According to the characteristics of the education scenario, an emotion-resource mapping strategy table is established. For example, when it is recognized that students are in the "strong thirst for knowledge" emotional state, the video resolution is increased to 1080p, and the audio bandwidth allocation for the key knowledge points of the video is increased by 20% to highlight the key content. For the "attention distraction" emotional state with a relatively high degree of urgency, the GPU occupancy of low-priority tasks such as automatic background synchronization of learning materials is paused to ensure smooth video playback.
[0052] Multi-modal Decision Engine Setup
[0053] Input Layer Configuration: Obtain students' historical learning behavior data from the database of the online education platform, such as the time students stay at different knowledge points and the frequency of asking questions. Integrate the emotion labels output by the speech emotion recognition module and the real-time network status data. For example, it is found that during the explanation of a certain complex knowledge point, the speech emotion recognition is "confused", and the historical data shows that the student stays at similar knowledge points for a long time. At the same time, the current network bandwidth is 8Mbps and the packet loss rate is 1%.
[0054] Output Layer Settings: According to the requirements of the online education scenario, set the video encoding parameter to H.265 with a bit rate of 3Mbps, the transmission protocol weight is 60% for UDP and 40% for TCP, and the rendering resource allocation ratio is to allocate 70% of the rendering resources to the course video content and 30% to interactive elements (such as bullet screens and note areas).
[0055] Decision Model Training: Build a simulation environment to simulate the learning scenarios of students in various emotional states under different network conditions. Use a network simulator to simulate various combinations of network bandwidth from 2Mbps to 20Mbps, latency from 10ms to 80ms, and packet loss rate from 0% to 8%, and train the decision model in combination with the emotional data of the education scenario. After 800 iterations of training, the emotion matching degree of the model in the education scenario test reaches 88%, the QoS score is improved from the original 3.5 points (out of 5 points) to 4.2 points, and the resource utilization rate is increased from 65% to 78%. The model update cycle is set to 12 hours to quickly adapt to the changes in students' learning behaviors and network status.
[0056] Method implementation
[0057] S1: Speech emotion classification
[0058] In an online education live classroom, students use learning devices with built-in microphones to collect voice streams in real time. The collected voice data is transmitted to the speech emotion recognition module through a stable network. The lightweight deep learning model in the module extracts MFCC features from the voice signal, generates a 13-dimensional MFCC feature vector, and simultaneously generates a spectrogram. After feature fusion by the hybrid Transformer and CNN architecture of the acoustic feature extraction unit, the DNN classifier of the emotion classification unit outputs the emotion label and confidence. For example, in a live class, the voices of 100 students are analyzed, and the number of accurately recognized emotion labels is 92, with an average confidence of 0.88.
[0059] S2: Determine the priority of video parameter adjustment
[0060] Input the emotion label generated in step S1 into the video resource scheduling module to query the emotion-resource mapping strategy table dedicated to the education scenario. If the emotion label is "strong thirst for knowledge", according to the strategy table, it is determined that the priority of increasing the video resolution and audio bandwidth allocation is higher, because in this emotional state, students are eager to acquire more knowledge, and clearer video and more prominent audio are helpful for learning.
[0061] S3: Generate resource allocation instructions
[0062] The multimodal decision-making engine receives data such as the speech emotion label, the current network bandwidth, the device load, and the priority list determined in step S2. For example, when the network bandwidth is 10 Mbps, the device GPU load is 50%, and the emotion label is "confused", the decision-making engine, according to the trained model, combines factors such as emotion matching degree and resource utilization rate, and determines that the video encoding parameter is H.265 with a bit rate of 4 Mbps, the transmission protocol weight is 70% for UDP and 30% for TCP, and the rendering resource allocation ratio is to allocate 80% of the rendering resources to the main part of the course video to highlight the knowledge point content.
[0063] S4: Dynamically adjust video resources and update the strategy
[0064] The video resource scheduling module dynamically adjusts the video encoder bitrate, transmission protocol, and GPU rendering resources according to the resource allocation instructions generated by the multi-modal decision engine. When the student's emotion is "confused", the video source resolution is increased to 1080p and the volume of the knowledge point explanation audio is increased by 20%. At the same time, the system listens to the user feedback behavior through the feedback mechanism of the online education platform, such as the student's after-class evaluation, learning effect test scores, etc. If it is found that the student's learning effect test scores do not improve significantly after adjusting the resources for a certain knowledge point explanation video, it indicates that the current resource allocation strategy may need to be optimized. The system will update the emotion-resource mapping strategy table and the model parameters of the multi-modal decision engine based on this feedback data. After one update, in the subsequent teaching of the same knowledge point, the average learning effect test scores of students increased by 10 points.
[0065] Embodiment 2: Video conferencing scenario
[0066] System construction
[0067] Construction of the voice emotion recognition module
[0068] Acoustic feature extraction unit: Adopt a hybrid architecture model of Transformer and CNN jointly pre-trained on a general speech dataset and a video conferencing scenario speech dataset. In view of the characteristics of multi-person interaction and environmental noise in the voice of the video conferencing scenario, the model structure is optimized. For example, a convolutional layer for noise suppression is added to the CNN part, and the convolutional kernel size is 5x5. Train with a large amount of video conferencing voice data. After 70 epochs of training, the model's accuracy of fusing the MFCC features and spatio-temporal features of the spectrogram of the voice signal reaches 91% in a complex video conferencing environment.
[0069] Emotion classification unit: Based on a database containing various emotional expressions in the video conferencing scenario, select the support vector machine (SVM) as the classifier model. During the training process, by adjusting the kernel function parameters and penalty factors, after 40 rounds of training, the emotion category recognition accuracy of the model on the video conferencing scenario test set reaches 87%. Using the knowledge distillation technology, a larger teacher model (such as a neural network with more hidden layers) is used to guide the SVM student model to learn, and the model inference latency is reduced to 45ms, and the accuracy remains stable.
[0070] Construction of the video resource scheduling module
[0071] Network status and device resource monitoring section: In a large-scale video conference with 100 participants, a network monitoring tool is used to monitor network bandwidth fluctuations in real time. It is found that the network bandwidth fluctuates in the range of 4 Mbps - 20 Mbps during the conference. The GPU load of the participants' devices is monitored through the API provided by the device driver. When the device is running video conferencing software, document sharing software, etc. simultaneously, the GPU load varies between 30% - 80%. An emotion-resource mapping strategy table for the video conferencing scenario is established. For example, when it is recognized that a participant is in an "excited" emotional state, a high frame rate mode (≥50 fps) and a low-latency transmission protocol (UDP priority) are triggered to better display the interactive atmosphere in the conference. For high-urgency emotions such as "irritable", the GPU occupancy of low-priority tasks such as automatically downloading meeting minutes is paused to ensure the smoothness of the video conference.
[0072] Multi-modal decision engine construction
[0073] Input layer configuration: Historical interaction behavior data of the participants, such as speaking duration, camera switching frequency, etc., is obtained from the logs of the video conferencing system. It is integrated with the emotion labels output by the speech emotion recognition module and the real-time network status data. For example, in a video conference, a participant's speech emotion is recognized as "excited", the historical data shows that they are actively speaking, and the current network bandwidth is 12 Mbps with a packet loss rate of 1.5%.
[0074] Output layer settings: According to the requirements of the video conferencing scenario, the video encoding parameters are set as H.265 with a bitrate of 5 Mbps, the transmission protocol weights are UDP at 70% and TCP at 30%, and the rendering resource allocation ratio is to allocate 75% of the rendering resources to the participants' video images and 25% to auxiliary content such as shared documents.
[0075] Decision model training: A simulation environment is constructed to simulate video conferencing scenarios of participants in various emotional states under different network conditions. A network simulator is used to simulate various combinations of network bandwidth from 3 Mbps - 25 Mbps, latency from 10 ms - 100 ms, and packet loss rate from 0% - 10%. The decision model is trained in combination with the emotion data of the video conferencing scenario. After 900 iterations of training, the emotion matching degree of the model in the video conferencing scenario test reaches 86%, the QoS score is improved from the original 3 points (out of 5) to 4 points, and the resource utilization rate is increased from 60% to 75%. The model update cycle is set to 24 hours, and an incremental learning method is adopted to incorporate new video conferencing data into the training process.
[0076] Method implementation
[0077] S1: Speech emotion classification
[0078] During the video conference, participants collect the voice stream in real time through the device microphone, and the collected voice data is transmitted over the network to the voice emotion recognition module. The lightweight deep learning model in the module extracts MFCC features and generates spectrograms for the voice signal. After feature fusion through the hybrid Transformer and CNN architecture of the acoustic feature extraction unit, the SVM classifier of the emotion classification unit outputs the emotion label and confidence level. For example, in a 2-hour video conference, when analyzing the voices of all participants, the proportion of accurately recognized emotion labels reaches 88%, and the average confidence level reaches 0.86.
[0079] S2: Determine the priority of video parameter adjustment
[0080] Input the emotion label generated in step S1 into the video resource scheduling module to query the emotion-resource mapping policy table for the video conference scenario. If the emotion label is "excited", according to the policy table, it is determined that the adjustment priorities for the high frame rate mode and low-latency transmission protocol are relatively high, because in this emotional state, participants are more concerned about the interaction effect and real-time nature of the conference.
[0081] S3: Generate resource allocation instructions
[0082] The multimodal decision-making engine receives data such as the voice emotion label, current network bandwidth, device load, and the priority list determined in step S2. For example, when the network bandwidth is 15 Mbps, the device GPU load is 60%, and the emotion label is "excited", the decision-making engine, based on the trained model, combines factors such as emotion matching degree and resource utilization rate to determine that the video encoding parameter is H.265 with a bitrate of 6 Mbps, the transmission protocol weight is 80% UDP and 20% TCP, and the rendering resource allocation ratio is to allocate 80% of the rendering resources to the video images of the participants to highlight the expressions and actions of the participants.
[0083] S4: Dynamically adjust video resources and update the policy
[0084] The video resource scheduling module dynamically adjusts the video encoder bitrate, transmission protocol, and GPU rendering resources according to the resource allocation instructions generated by the multimodal decision-making engine. When the emotion of the participants is "excited", the video frame rate is increased to 50 fps and the UDP transmission protocol is preferentially used. At the same time, the system listens to the user feedback behavior through the feedback function of the video conference system, such as the evaluation of the conference experience by the participants and the number of disconnections. If it is found that the number of disconnections increases during the conference, it indicates that there may be problems with the current resource allocation strategy in terms of network stability. The system will update the emotion-resource mapping policy table and the model parameters of the multimodal decision-making engine based on this feedback data. After one update, during subsequent video conferences of the same scale, the average number of disconnections decreases by 3 times.
[0085] Example 3: Film and TV entertainment scene
[0086] System Construction
[0087] Speech emotion recognition module construction
[0088] Acoustic feature extraction unit: A Transformer and CNN hybrid architecture model pre-trained on film and television related speech datasets was selected. In view of the rich emotional expressions and diverse timbre characteristics of speech in film and television entertainment scenes, the attention mechanism of the model was improved, and an adaptive attention module was added to better focus on speech features of different emotions. Using speech data from massive film and television clips for training, after 60 epochs of training, the model achieved an accuracy rate of 94% in fusing the MFCC features of the speech signal of the film and television scene and the spatiotemporal features of the spectrogram.
[0089] Emotion classification unit: Based on a database containing various film and television emotional expressions, a deep neural network (DNN) was selected as the classifier model. During the training process, the RMSProp optimizer with a learning rate of 0.0008 was used. After 35 rounds of training, the model's emotion category recognition accuracy on the film and television entertainment scene test set reached 91%. By combining quantitative perception training and pruning technology, the number of model parameters was reduced by 30%, the model inference delay was reduced to 35ms, and the accuracy rate only dropped by 0.5%.
[0090] Video resource scheduling module construction
[0091] Network status and device resource monitoring section: In a platform with 100,000 users watching movies and TV online at the same time, use network monitoring tools to obtain real-time network bandwidth fluctuation data. During the prime viewing time in the evening, the network bandwidth fluctuates between 2Mbps and 15Mbps. Use the API provided by the device driver to monitor the GPU load of the user-end device. When the user uses a smart TV to watch high-definition movies and performs operations such as fast forward and pause, the GPU load varies between 15% and 60%. Establish an emotion-resource mapping strategy table for film and television entertainment scenes. For example, when it is identified that the user is in a "tense" emotional state, increase the contrast of the video screen by 15% to enhance the creation of a tense atmosphere. For the emotional state of "boredom" that has a lower degree of urgency but affects the user experience, appropriately reduce the video encoding bit rate to save bandwidth resources.
[0092] Building a multimodal decision engine
[0093] Input layer configuration: Obtain the user's historical viewing behavior data from the user behavior database of the film and television entertainment platform, such as the viewing duration of different types of films, the preference for favorite films, etc. Integrate the emotion tags output by the speech emotion recognition module and the real-time network status data. For example, when a user is watching a suspense movie, the speech emotion recognition is "nervous", the historical data shows that the user prefers suspense movies, and the current network bandwidth is 7 Mbps with a packet loss rate of 2%.
[0094] Output layer settings: According to the requirements of the film and television entertainment scenario, set the video encoding parameter to H.265 with a bit rate of 4 Mbps, the transmission protocol weight is 60% for UDP and 40% for TCP, and the rendering resource allocation ratio is to allocate 80% of the rendering resources to the film and television video screen and 20% to auxiliary elements such as subtitles.
[0095] Decision model training: Build a simulation environment to simulate the film and television viewing scenarios of users in various emotional states under different network states. Use a network simulator to simulate various combinations of network bandwidth from 1 Mbps to 20 Mbps, latency from 10 ms to 90 ms, and packet loss rate from 0% to 9%, and train the decision model in combination with the emotional data of the film and television entertainment scenario. After 1000 iterations of training, the emotion matching degree of the model in the film and television entertainment scenario test reaches 89%, the QoS score is improved from the original 3.2 points (out of 5 points) to 4.3 points, and the resource utilization rate is increased from 63% to 77%. The model update cycle is set to 18 hours to adapt to the changes in film and television content and user preferences.
[0096] Method implementation
[0097] S1: Speech emotion classification
[0098] During the user's viewing of the film and television, the speech stream is collected in real time through the microphone of the smart TV or mobile device, and the collected speech data is transmitted to the speech emotion recognition module through the network. The lightweight deep learning model in the module extracts MFCC features and generates spectrograms for the speech signal. After feature fusion by the hybrid architecture of Transformer and CNN in the acoustic feature extraction unit, the DNN classifier in the emotion classification unit outputs the emotion tags and confidence levels. For example, in the speech analysis of 1000 users watching film and television, the number of accurately identified emotion tags is 910, and the average confidence level reaches 0.89.
[0099] S2: Determine the priority of video parameter adjustment
[0100] Input the sentiment tags generated in step S1 into the video resource scheduling module to query the sentiment-resource mapping policy table for the film and television entertainment scenario. If the sentiment tag is "tense", according to the policy table, it is determined that the priority of improving the video picture contrast is relatively high, because in this emotional state, enhancing the picture effect can better meet the emotional needs of users.
[0101] S3: Generate a resource allocation instruction
[0102] The multimodal decision-making engine receives data such as voice sentiment tags, the current network bandwidth, device load, and the priority list determined in step S2. For example, when the network bandwidth is 8 Mbps, the device GPU load is 40%, and the sentiment tag is "tense", the decision-making engine, based on the trained model, combines factors such as sentiment matching degree and resource utilization rate, and determines that the video encoding parameter is H.265 with a bitrate of 5 Mbps, the transmission protocol weight is 70% for UDP and 30% for TCP, and the rendering resource allocation ratio is to allocate 85% of the rendering resources to the main picture of the film and television video to highlight the key plot.
[0103] S4: Dynamically adjust video resources and update the policy
[0104] The video resource scheduling module dynamically adjusts the video encoder bitrate, transmission protocol, and GPU rendering resources according to the resource allocation instruction generated by the multimodal decision-making engine. When the user's emotion is "tense", the video picture contrast is increased by 15%. At the same time, the system listens to the user's feedback behavior through the user feedback channels of the film and television entertainment platform, such as the user's film reviews and ratings. If it is found that the user is not satisfied with the resource adjustment effect of a certain type of film and television in a specific emotional state, it means that the current resource allocation policy may need to be optimized. The system will update the sentiment-resource mapping policy table and the model parameters of the multimodal decision-making engine according to these feedback data. After one update, the average score of the user's ratings has increased by 0.5 points in the subsequent play of the same type of film and television.
[0105] The above embodiments cover multiple typical scenarios such as online education, video conferencing, and film and television entertainment, fully demonstrating the effectiveness and practicality of the intelligent optimization system and method of integrating video resource scheduling and voice emotion recognition of the present invention in different multimedia application scenarios. Through the actual application and optimization in each scenario, it can effectively improve the quality of the user experience in the multimedia interaction process and meet the diverse needs of users.
[0106] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0107] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
[0108] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent optimization system integrating video resource scheduling and speech emotion recognition, characterized in that: include: Speech emotion recognition module: used to collect user speech streams in real time, extract acoustic features through deep learning models, perform emotion classification, and output emotion labels; Video resource scheduling module: used to monitor network status and device resources, and dynamically adjust video encoding parameters, transmission protocols, and rendering priorities according to emotion tags; Multimodal decision engine: Based on the reinforcement learning model, it integrates voice emotion labels, network status data and user historical behavior data to generate real-time resource allocation strategies.
2. According to claim 1, an intelligent optimization system and method integrating video resource scheduling and speech emotion recognition is characterized in that: The speech emotion recognition module comprises: The acoustic feature extraction unit uses a hybrid architecture of Transformer and CNN to perform spatiotemporal feature fusion on the MFCC features and spectrograms of speech signals. The sentiment classification unit builds a classifier based on the pre-trained sentiment database (such as RAVDESS and IEMOCAP) and outputs the sentiment category and confidence.
3. An intelligent optimization system integrating video resource scheduling and speech emotion recognition according to claim 1 or 2, characterized in that: The speech emotion recognition module integrates lightweight model compression technology, including at least one of knowledge distillation and quantization perception training, and the model inference delay is less than 50ms.
4. The intelligent optimization system integrating video resource scheduling and speech emotion recognition according to claim 1 is characterized in that: The dynamic adjustment of the video resource scheduling module includes: Real-time monitoring of network bandwidth fluctuations, device GPU load, and video resolution requirements; Establish an emotion-resource mapping strategy table to define the correspondence between emotion categories and video parameters, including: Angry emotion: trigger high frame rate mode (≥60fps) and low latency transmission protocol (UDP preferred); Sad mood: triggers high color saturation mode and enhanced audio bandwidth allocation.
5. The intelligent optimization system integrating video resource scheduling and speech emotion recognition according to claim 4 is characterized in that: The emotion-resource mapping strategy table further includes emotion urgency classification, where high-urgency emotions (such as anxiety and anger) trigger preemptive resource allocation and suspend GPU occupancy of low-priority tasks.
6. The intelligent optimization system integrating video resource scheduling and speech emotion recognition according to claim 1 is characterized in that: The multimodal decision engine comprises: Input layer: receives speech emotion labels, network throughput, packet loss rate, and user historical interaction behaviors (such as the frequency of manual resolution adjustment); Output layer: generates video encoding parameters (H.265 / AV1 bit rate), transmission protocol weight (UDP / TCP mixed ratio), and rendering resource allocation ratio; Decision model: Based on the proximal policy optimization (PPO) algorithm training, the reward function includes sentiment matching, QoS score and resource utilization.
7. The intelligent optimization system integrating video resource scheduling and speech emotion recognition according to claim 6 is characterized in that: The training process of the decision model includes: Build a simulation environment to simulate different network states and emotional scenarios; Through offline reinforcement learning iterative optimization strategy, the model update cycle is 24 hours, and incremental learning is used to adapt to new user data.
8. An intelligent optimization method integrating video resource scheduling and speech emotion recognition, characterized in that: The following steps are involved: S1: Collect user voice streams in real time, perform sentiment classification through a lightweight deep learning model, and generate sentiment labels and confidence levels; S2: query the emotion-resource mapping strategy table according to the emotion label to determine the priority of video parameter adjustment; S3: Generate resource allocation instructions through a multimodal decision engine based on the current network bandwidth, device load, and priority list; S4: Dynamically adjust the video encoder bit rate, transmission protocol, and GPU rendering resources, and monitor user feedback to update the strategy.
9. The intelligent optimization method integrating video resource scheduling and speech emotion recognition according to claim 8 is characterized in that: The dynamic adjustment in step S4 includes: When the user's emotion is "confused", the video source resolution is increased to 4K and the background special effects rendering accuracy is reduced; when the user's emotion is "calm", the bandwidth saving mode is enabled to limit the bit rate peak and delay the transmission of non-key frames.
Citation Information
Cited By
Application program management system and method driven by virtual engine
CN120635279A
A virtual engine driven application management system and method
CN120635279B
Audio and video processing system and method supporting AI intelligent repair technology
CN121126017A