Intelligent processing and analyzing method and system for audios and videos in video communication

By introducing intelligent processing analysis methods and modules into the video communication system, identifying scenarios and switching processing strategies, and optimizing resource usage with user feedback, the problem of processing data and resource consumption in the existing technology is solved, and an efficient and low-latency video communication experience is achieved.

CN120111037APending Publication Date: 2025-06-06ZHEJIANG DAXIN TECH CO LTD

Patent Information

Application Number
CN202510117791.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06

Smart Images

  • Figure CN120111037A_ABST
    Figure CN120111037A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent processing analysis method and system for audios and videos in video communication, and relates to the technical field of information retrieval, and the system comprises an audio and video collection module which is used for collecting and transmitting real-time video stream data in the video communication; the scene recognition module comprises a scene classification unit and a self-adaptive adjustment unit, and the scene classification unit is used for receiving the real-time video stream data of the audio and video acquisition module, recognizing a scene of current video communication according to the real-time video stream data, and performing classification according to the scene to obtain a classification result of the current video communication scene; audio and video data are processed in real time through the real-time analysis and processing module, it is ensured that quick response and low-delay communication experience can be provided in scenes such as emergency response and video conferences with high requirements for real-time performance, and the module optimizes use of computing resources through intelligent scheduling and resource management so as to meet the requirements for real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information retrieval, and more specifically to a method and system for intelligent processing and analysis of audio and video in video communication. Background Art

[0002] It should be noted that in today's digital age, video communication has become a core tool in remote collaboration, social interaction, and online education. With the advancement of technology, the demand for video communication is growing, and the requirements for communication quality are also getting higher and higher. In order to improve the experience of video communication, intelligent processing and analysis technology of audio and video has emerged;

[0003] However, in actual use, video communication has different usage scenarios, such as video conferencing, online education, social video calls, and emergency response situations. In different usage scenarios, the analysis and processing methods of audio and video are also different. Existing processing and analysis solutions are difficult to process and integrate data from different modalities and understand and associate information between different modalities.

[0004] Moreover, video communication requires low latency and high real-time performance, especially in scenarios such as emergency response and video conferencing, which consumes a lot of computing resources. It is difficult to optimize resource consumption while ensuring real-time performance. Summary of the invention

[0005] To solve the above problems, the present invention provides a method and system for intelligent processing and analysis of audio and video in video communication.

[0006] The present invention provides an intelligent processing and analysis method and system for audio and video in video communication, comprising an audio and video acquisition module, wherein the audio and video acquisition module is used to acquire real-time video stream data in video communication and transmit it;

[0007] A scene recognition module, the scene recognition module includes a scene classification unit and an adaptive adjustment unit, the scene classification unit is used to receive the real-time video stream data of the audio and video acquisition module, and identify the scene of the current video communication according to the real-time video stream data, classify according to the scene, obtain the classification result of the current video communication scene, and transmit it to the adaptive adjustment unit;

[0008] The adaptive adjustment unit is used to receive the communication scenario classification result, switch the corresponding audio and video processing strategy according to the communication scenario classification result, and transmit the processing strategy signal;

[0009] An interactive feedback module, which is used to collect and transmit user feedback information on video communication;

[0010] The processing and analysis module is used to receive the real-time video stream data of the audio and video acquisition module, the processing strategy signal of the adaptive adjustment module and the feedback information of the interactive feedback module to generate real-time processing and analysis results.

[0011] Preferably, the specific working steps of the audio and video acquisition module are as follows:

[0012] Detect available audio and video input devices;

[0013] Synchronously capture the video and audio streams in the video communication through the audio and video input device to obtain real-time video stream data;

[0014] Compress and encode the captured real-time video stream data;

[0015] The compressed and encoded real-time video stream data is transmitted to the scene recognition module.

[0016] Preferably, the specific working steps of the scene classification unit are as follows:

[0017] The classification results are set in advance, including video conferencing, online education, social video calls, and emergency response, and the known category features of each classification result are obtained;

[0018] Receive real-time video stream data from the audio and video acquisition module, and extract key features from the real-time video stream data, including color features, texture features, shape features, motion features, and space features;

[0019] Calculate the similarity between the extracted key features and the known category features of each classification result;

[0020] Construct a similarity matrix to record the similarity between all key features and known category features of each classification result;

[0021] A classification threshold is set for each classification result, and the calculated similarity is compared with the classification threshold to obtain a classification result of the current video communication scene, and the result is transmitted to the adaptive adjustment unit.

[0022] Preferably, the specific steps of extracting key features from real-time video stream data are as follows:

[0023] Parse the real-time video stream into a continuous sequence of image frames, extract individual image frames from the video stream at fixed time intervals or based on specific event trigger sampling;

[0024] Convert a color image frame to a grayscale image;

[0025] Apply SIFT algorithm to the grayscale image to detect key points and extract corresponding feature descriptors;

[0026] Calculate the gradient magnitude and direction of each pixel in the image frame and generate a directional gradient histogram;

[0027] The SIFT feature descriptor and the HOG feature vector are merged to form a comprehensive feature vector.

[0028] Preferably, the specific working mode of the adaptive adjustment unit is as follows:

[0029] Presetting the parameter data of the video communication, specifically including video parameters: resolution, frame rate, bit rate, and adaptability preference;

[0030] Resolution, frame rate, and bit rate include three levels: high, medium, and low;

[0031] Audio parameters: sampling rate, encoding bit rate and audio processing;

[0032] The sampling rate and encoding bit rate include three levels: high, medium and low;

[0033] Network parameters: bandwidth estimation and packet loss compensation;

[0034] Obtain the classification result of the current video communication scene. If the classification result is a video conference, adjust the video parameters to: resolution: high; frame rate: high; bit rate: high; and the adaptive preference is to maintain the frame rate;

[0035] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: high; audio processing: turn on echo cancellation and automatic gain control;

[0036] The network parameters are adjusted as follows: Bandwidth estimation: real-time monitoring of network conditions and dynamic adjustment of bit rate; Packet loss compensation: enabling packet loss compensation mechanism;

[0037] If the classification result is online education, the video parameters are adjusted to: resolution: medium; frame rate: low; bit rate: medium; adaptive preference is adjusted to ensure video quality;

[0038] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control;

[0039] The network parameters are adjusted as follows: Bandwidth estimation: monitor network conditions in real time and adjust the bit rate dynamically; Congestion control: enable congestion control;

[0040] If the classification result is a social video call, the video parameters are adjusted as follows: resolution: medium; frame rate: medium; bit rate: medium; adaptive preference: maintain balance;

[0041] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control;

[0042] The network parameters are adjusted as follows: Packet loss compensation: Enable packet loss compensation mechanism; Congestion control: Enable congestion control;

[0043] If the classification result is emergency response, the video parameters are adjusted to: resolution: low; frame rate: low; bit rate: low; adaptive preference: prioritize frame rate;

[0044] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control;

[0045] The network parameters are adjusted as follows: Bandwidth estimation: real-time monitoring of network conditions and dynamic adjustment of bit rate; Packet loss compensation: enabling packet loss compensation mechanism;

[0046] A processing strategy signal is generated according to the parameter adjustment and transmitted to the processing and analysis module.

[0047] Preferably, when the adaptive adjustment unit is working, the parameter data of the video communication needs to be dynamically adjusted in real time, and the specific steps are as follows:

[0048] For bit rate adjustment, according to the formula Calculate the bit rate B after real-time dynamic adjustment, where B 0 is the basic bit rate, PL is the packet loss rate, EWMA bw is the exponentially weighted moving average bandwidth estimate, BW max is the maximum bandwidth observed historically;

[0049] For frame rate adjustment, according to the formula The frame rate FR after real-time dynamic adjustment is calculated, where FR is the adjusted frame rate, FR 0 Is the base frame rate, EWMA la is the exponentially weighted moving average delay, T is the target delay threshold;

[0050] According to the calculated bit rate B after real-time dynamic adjustment and the frame rate FR after real-time dynamic adjustment, the parameter data of the video communication needs to be dynamically adjusted in real time.

[0051] Preferably, the specific working steps of the interactive feedback module are as follows:

[0052] Design a user-friendly feedback interface that allows users to provide feedback during or after the video call and obtain user feedback scores;

[0053] Monitor user interaction behavior in real time;

[0054] The user's feedback score and the user's interactive behavior are used as feedback information and transmitted to the processing and analysis module.

[0055] Preferably, the specific working steps of the processing and analysis module are as follows:

[0056] Collect user rating data and count the number of user interactions during video communication;

[0057] Evaluate and classify user rating data and the number of user interactions during video communication into three levels: high, medium, and low;

[0058] If the user rating is "high" and the number of negative interactions is "low", the feedback information is marked as good;

[0059] If the user rating is "Medium" and the number of negative interactions is "Low" or "Medium", mark the feedback as good;

[0060] If the user rating is "Medium or Low" and the number of negative interactions is "Medium" or "High", the feedback information is marked as Medium;

[0061] If the user rating is "low" or the number of negative interactions is "high", the feedback information is marked as poor;

[0062] If the feedback information is excellent, the parameter setting of the adaptive adjustment unit can already meet the user's needs and no adjustment is required;

[0063] If the feedback information is good, the parameter setting of the adaptive adjustment unit is adjusted by increasing the bit rate and frame rate by 5% based on the parameter setting of the adaptive adjustment unit;

[0064] If the feedback information is medium, the parameter setting of the adaptive adjustment unit is adjusted, and the bit rate and frame rate are increased by 10% based on the parameter setting of the adaptive adjustment unit;

[0065] If the feedback information is poor, the parameter setting of the adaptive adjustment unit is adjusted, then the bit rate and frame rate are increased by 20% based on the parameter setting of the adaptive adjustment unit.

[0066] Preferably, the specific working steps of evaluating and classifying the user rating data and the number of user interactions during the video communication process include:

[0067] Perform statistical analysis on user rating data and the number of user interactions during video communications, including calculation of historical averages and standard deviations;

[0068] If the user rating data and the number of user interactions during video communication are less than one standard deviation below the mean, it is a low grade;

[0069] If the user rating data and the number of user interactions during video communication are within one standard deviation of the mean, it is considered medium;

[0070] If the user rating data and the number of user interactions during video communication are abnormal by one standard deviation above the mean, it is considered a high level.

[0071] The present invention also proposes a method for intelligent processing and analysis of audio and video in video communication, comprising the following steps:

[0072] Step 1: Collect real-time video stream data in video communication, identify the current video communication scene according to the real-time video stream data, classify according to the scene, and obtain the current video communication scene classification result;

[0073] Step 2: Switch the corresponding audio and video processing strategy according to the communication scenario classification result, and transmit the processing strategy signal;

[0074] Step 3: Generate real-time processing and analysis results based on real-time video stream data, processing strategy signals and feedback information.

[0075] Beneficial effects: The multimodal data fusion module integrates multiple data such as video, audio and text, and uses advanced algorithms to understand and associate information between different modalities. This enables the system to provide richer and more accurate analysis results, such as evaluating learning effects by analyzing students' facial expressions and voice feedback in online education scenarios, or quickly identifying and responding to emergencies through video analysis in emergency response scenarios;

[0076] The system processes audio and video data in real time through the real-time analysis and processing module, ensuring rapid response and providing low-latency communication experience in scenarios with high real-time requirements such as emergency response and video conferencing. The module optimizes the use of computing resources through intelligent scheduling and resource management to meet real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a flow chart of the management system of the present invention. DETAILED DESCRIPTION

[0078] Application scenarios: In actual use, video communication has different usage scenarios, such as video conferencing, online education, social video calls, and emergency response situations. In different usage scenarios, the analysis and processing methods of audio and video are also different. Existing processing and analysis solutions are difficult to process and integrate data from different modalities and understand and associate information between different modalities.

[0079] Moreover, video communication requires low latency and high real-time performance, especially in scenarios such as emergency response and video conferencing, which consumes a lot of computing resources. It is difficult to optimize resource consumption while ensuring real-time performance.

[0080] like Figure 1As shown: an intelligent processing and analysis system for audio and video in video communication, including an audio and video acquisition module, the audio and video acquisition module is used to collect real-time video stream data in video communication and transmit;

[0081] A scene recognition module, the scene recognition module includes a scene classification unit and an adaptive adjustment unit, the scene classification unit is used to receive the real-time video stream data of the audio and video acquisition module, and identify the scene of the current video communication according to the real-time video stream data, classify according to the scene, obtain the classification result of the current video communication scene, and transmit it to the adaptive adjustment unit;

[0082] The adaptive adjustment unit is used to receive the communication scenario classification result, switch the corresponding audio and video processing strategy according to the communication scenario classification result, and transmit the processing strategy signal;

[0083] An interactive feedback module, which is used to collect and transmit user feedback information on video communication;

[0084] The processing and analysis module is used to receive the real-time video stream data of the audio and video acquisition module, the processing strategy signal of the adaptive adjustment module and the feedback information of the interactive feedback module to generate real-time processing and analysis results.

[0085] It should be noted that the audio and video acquisition module transmits the collected real-time video stream data to the scene recognition module to provide input for scene recognition, and adjusts the acquisition parameters, such as resolution and frame rate, according to the feedback from the adaptive adjustment unit to adapt to different processing requirements;

[0086] The scene recognition module receives data from the audio and video acquisition module and outputs the scene classification result to the adaptive adjustment unit;

[0087] The interactive feedback module transmits the user feedback information to the processing and analysis module for personalized processing, and adjusts the collection method and content of the user feedback according to the scene classification result of the scene recognition module;

[0088] The processing and analysis module receives the real-time video stream data from the audio and video acquisition module and the processing strategy signal from the adaptive adjustment module, and adjusts the processing strategy according to the user feedback information from the interactive feedback module to meet the user's needs;

[0089] The audio and video acquisition module is the starting point of the data flow. Its data is used by the scene recognition module and the processing and analysis module. The output of the scene recognition module (scene classification result) is used by the adaptive adjustment unit and the processing and analysis module. The output of the interactive feedback module (user feedback information) is used by the processing and analysis module.

[0090] The adaptive adjustment unit adjusts the processing strategy according to the output of the scene recognition module and transmits it to the processing analysis module, and the processing analysis module adjusts the processing strategy according to the output of the interactive feedback module;

[0091] The output of the processing and analysis module can be fed back to the audio and video acquisition module to optimize the data acquisition quality; at the same time, the processing and analysis results can also be fed back to the interactive feedback module for further collecting user feedback;

[0092] It should also be noted that the system automatically identifies the current communication scenario (such as video conferencing, online education, etc.) through the scene recognition module, and according to the recognition result, the adaptive adjustment unit dynamically switches to the audio and video processing strategy that best suits the current scenario. This adaptive mechanism enables the system to flexibly respond to the needs of different scenarios, optimize resource allocation, and improve processing efficiency;

[0093] The multimodal data fusion module integrates multiple data such as video, audio and text, and uses advanced algorithms to understand and associate information between different modalities. This enables the system to provide richer and more accurate analysis results, such as evaluating learning effects by analyzing students' facial expressions and voice feedback in online education scenarios, or quickly identifying and responding to emergencies through video analysis in emergency response scenarios;

[0094] The system processes audio and video data in real time through the real-time analysis and processing module, ensuring fast response and low-latency communication experience in scenarios with high real-time requirements such as emergency response and video conferencing. This module optimizes the use of computing resources through intelligent scheduling and resource management to meet real-time requirements;

[0095] The system dynamically adjusts the priority and resource allocation of analysis tasks according to the current network conditions and computing resources. This dynamic resource management strategy helps to optimize resource consumption while ensuring real-time performance. Especially on resource-constrained mobile devices, it can balance battery usage and processing performance to extend the device's usage time.

[0096] As a further embodiment, the specific working steps of the audio and video acquisition module are as follows:

[0097] Detect available audio and video input devices; such as cameras and microphones;

[0098] Synchronously capture the video and audio streams in the video communication through the audio and video input device to obtain real-time video stream data;

[0099] Compress and encode the captured real-time video stream data to reduce the data volume and adapt to network transmission;

[0100] The compressed and encoded real-time video stream data is transmitted to the scene recognition module.

[0101] It should be noted that,input is provided for scene recognition;

[0102] As a further embodiment, the specific working steps of the scene classification unit are as follows:

[0103] Classification results are set in advance, specifically including video conferencing, online education, social video calls, and emergency response, and known category features of each classification result are obtained: It should be noted that the known category features of each classification result include color features, texture features, shape features, motion features, and space features;

[0104] For example, if the classification result is a video conference, the characteristics are that the background color is usually relatively simple, such as the neutral-toned walls or curtains in a conference room, and the background texture may be relatively smooth, such as the wall of an office or the table in a conference room. There may be frontal images of multiple people and the rectangular shape of a shared screen; there is less movement and the movements of the participants are relatively stable; spatial characteristics: the scene is relatively closed, and the standard layout of a conference room may appear in the background;

[0105] If the classification result is online education, the color feature is that the background color may be more diverse, but usually contains the color of teaching materials, such as whiteboards or books; the texture feature is that the background may contain the texture of educational materials such as books and paper;

[0106] Shape features are the possible appearance of a teacher’s frontal image and specific shapes of teaching materials, such as an open book;

[0107] The movement characteristics are that teachers may produce certain movements when writing or manipulating teaching materials;

[0108] The spatial characteristics are that the scene may be relatively open, and teaching-related layouts may appear in the background, such as classrooms;

[0109] If the classification result is a social video call, the color feature is that the background color may be more diverse, reflecting the characteristics of the personal space;

[0110] Texture features are textures of household items that may be included in the background, such as sofas, carpets, etc.

[0111] The shape feature is the shape in which various personal items may appear, such as photo frames, plants, etc.;

[0112] Motion features mean that participants may have more dynamic behaviors in the video, such as gestures, movements, etc.

[0113] The spatial characteristics are that the scene may be more casual, and the background may contain daily life environments;

[0114] If the classification result is emergency response, the color feature is that the background color may be relatively simple and bright, such as the red or orange of the emergency vehicle;

[0115] Texture features are textures that may be contained in the background that are unique to emergency scenes, such as the rough texture of firefighter uniforms;

[0116] Shape features are those of possible emergency equipment, such as a fire hose reel or an ambulance;

[0117] The motion feature is that there may be fast motion in the scene, such as people running or vehicles moving;

[0118] The spatial feature is that the scene may be relatively open, and outdoor environments or emergency scenes may appear in the background;

[0119] Receive real-time video stream data from the audio and video acquisition module, and extract key features from the real-time video stream data, including color features, texture features, shape features, motion features, and space features;

[0120] It should be noted that the color features are obtained by calculating the average color value of all pixels in the video frame;

[0121] Texture features: describe the statistical characteristics of image texture, such as contrast, energy, uniformity, etc.;

[0122] Shape features: Identify edge information in video frames and describe the outline of objects; such as bounding rectangles, minimum covering circles, etc., to describe the shape features of objects;

[0123] Motion features: Estimate pixel motion between video frames to describe dynamic changes in the scene; Motion vectors calculated by block matching algorithms to describe the direction and speed of object movement;

[0124] Spatial features: key points detected by algorithms such as SIFT and SURF, and feature descriptors generated by algorithms such as SIFT and ORB;

[0125] The extracted key features and the known category features of each classification result are similarly calculated; it should be noted that the similarity between the features is calculated using mathematical formulas such as Euclidean distance or cosine similarity;

[0126] For example, if the known categories are "video conference", "online education", "social video call" and "emergency response", the similarity between the features of the current video frame and the features of these four categories is calculated;

[0127] The specific calculation method can be calculated by Euclidean distance or cosine similarity;

[0128] Construct a similarity matrix to record the similarities between all key features and known category features of each classification result; it should be noted that for N features, the similarity matrix is ​​an NxN matrix;

[0129] A classification threshold is set for each classification result, and the calculated similarity is compared with the classification threshold to obtain a classification result of the current video communication scene, and the result is transmitted to the adaptive adjustment unit.

[0130] It should be noted that the classification threshold of each classification result can be determined by maximizing the F1 score to determine the optimal threshold. For each possible threshold, the corresponding precision and recall rate are calculated; each possible threshold is obtained by giving a batch of possible numbers by the staff;

[0131] The F1 score of each threshold is calculated by multiplying 2 by precision and recall, and then divided by precision plus recall. The threshold with the largest F1 score is selected as the optimal threshold.

[0132] Precision refers to the proportion of samples that are actually positive among all samples predicted to be positive;

[0133] Recall refers to the proportion of samples predicted to be positive among all samples that are actually positive.

[0134] For example, if the features of a video frame have the highest similarity with the features of the “video conference” category and exceed a preset threshold, then the video frame is classified as “video conference”;

[0135] As a further embodiment, the specific steps of extracting key features from real-time video stream data are as follows:

[0136] Parse the real-time video stream into a continuous sequence of image frames, extract individual image frames from the video stream at fixed time intervals or based on specific event trigger sampling;

[0137] Converts color image frames to grayscale images; simplifies the data and highlights the texture and contour information of the image.

[0138] Apply SIFT algorithm to the grayscale image to detect key points and extract corresponding feature descriptors;

[0139] Calculate the gradient magnitude and direction of each pixel in the image frame and generate a directional gradient histogram;

[0140] The SIFT feature descriptor and the HOG feature vector are merged to form a comprehensive feature vector.

[0141] As a further embodiment, the specific working mode of the adaptive adjustment unit is as follows:

[0142] Presetting the parameter data of the video communication, specifically including video parameters: resolution, frame rate, bit rate, and adaptability preference;

[0143] Resolution, frame rate, and bit rate include three levels: high, medium, and low. It should be noted that in this embodiment, the three levels include resolution: high (1080p), medium (720p), and low (480p).

[0144] Frame rate: High (30fps), Medium (20fps), Low (15fps);

[0145] Bit rate: high (1000kbps), medium (500-1000kbps), low (<500kbps);

[0146] Audio parameters: sampling rate, encoding bit rate and audio processing;

[0147] The sampling rate and encoding bit rate include three levels: high, medium and low. It should be noted that the sampling rate is high (48kHz), medium (44.1kHz) and low (32kHz).

[0148] Encoding bitrate: high (128kbps), medium (96kbps), low (64kbps);

[0149] Network parameters: bandwidth estimation and packet loss compensation;

[0150] Obtain the classification result of the current video communication scene. If the classification result is a video conference, adjust the video parameters to: resolution: high; frame rate: high; bit rate: high; and the adaptive preference is to maintain the frame rate;

[0151] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: high; audio processing: turn on echo cancellation and automatic gain control;

[0152] The network parameters are adjusted as follows: Bandwidth estimation: real-time monitoring of network conditions and dynamic adjustment of bit rate; Packet loss compensation: enabling packet loss compensation mechanism;

[0153] If the classification result is online education, the video parameters are adjusted to: resolution: medium; frame rate: low; bit rate: medium; adaptive preference is adjusted to ensure video quality;

[0154] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control

[0155] The network parameters are adjusted as follows: Bandwidth estimation: monitor network conditions in real time and adjust the bit rate dynamically; Congestion control: enable congestion control;

[0156] If the classification result is a social video call, the video parameters are adjusted as follows: resolution: medium; frame rate: medium; bit rate: medium; adaptive preference: maintain balance;

[0157] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control;

[0158] The network parameters are adjusted as follows: Packet loss compensation: Enable packet loss compensation mechanism; Congestion control: Enable congestion control;

[0159] If the classification result is emergency response, the video parameters are adjusted to: resolution: low; frame rate: low; bit rate: low; adaptive preference: prioritize frame rate;

[0160] The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control;

[0161] The network parameters are adjusted as follows: Bandwidth estimation: real-time monitoring of network conditions and dynamic adjustment of bit rate; Packet loss compensation: enabling packet loss compensation mechanism;

[0162] A processing strategy signal is generated according to the parameter adjustment and transmitted to the processing and analysis module.

[0163] It should be noted that for video conferencing:

[0164] High resolution and frame rate: Video conferencing requires clear facial expressions and body language for better communication and understanding;

[0165] High bit rate: ensures that the video quality is as high as possible under network conditions;

[0166] Prioritize frame rate: When network bandwidth is limited, maintaining a smooth video experience is more important than sacrificing frame rate to maintain resolution;

[0167] Clear video and smooth dynamic images can improve participation and interactivity in remote meetings, and high-quality audio can ensure clear communication of voice information and reduce misunderstandings;

[0168] For online education:

[0169] Adapt to different network conditions: Online education needs to adapt to the network environment of different students to ensure the accessibility of content;

[0170] Prioritize video quality: Educational content often requires clear presentation of details, such as graphics and text;

[0171] Appropriate video quality can improve students' learning experience and help them absorb information better. Reasonable bit rate adjustment can ensure that educational resources are most effectively distributed under limited network conditions.

[0172] For social video calls:

[0173] Balance video quality and network bandwidth: Social calls usually focus more on real-time interaction, and need to find a balance between video quality and network bandwidth;

[0174] Maintain audio quality: Clear audio is essential to maintaining social interactions;

[0175] Appropriate video and audio quality can enhance social interactions between users, provide a more natural communication experience, and reduce the impact of network problems: By dynamically adjusting parameters, the impact of network fluctuations on call quality can be reduced;

[0176] For emergency response:

[0177] Prioritize frame rate: In an emergency, real-time performance is more important than video quality, and on-site information needs to be delivered quickly.

[0178] Adapt to possible low-bandwidth environments: Emergency response may occur in areas with poor network conditions;

[0179] High frame rate and low bit rate ensure that key information can be delivered in real time in emergency situations. Clear audio and key video images can help emergency personnel make decisions quickly.

[0180] By adjusting the above parameters, applications in different video communication scenarios can better adapt to specific needs and restrictions, thereby providing a better user experience and more efficient communication effects;

[0181] As a further embodiment, when the adaptive adjustment unit is working, the parameter data of the video communication needs to be dynamically adjusted in real time, and the specific steps are as follows:

[0182] For bit rate adjustment, according to the formula Calculate the bit rate B after real-time dynamic adjustment, where B 0 is the basic bit rate, PL is the packet loss rate, EWMA bw is the exponentially weighted moving average bandwidth estimate, BW max is the maximum bandwidth observed in history; it should be noted that the basic bit rate is obtained by adjusting the parameters of the processing strategy signal, the packet loss rate refers to the proportion of data packets lost during data transmission, which is obtained by dividing the number of lost data packets by the total number of sent data packets, and the exponentially weighted moving average bandwidth is estimated;

[0183] It should also be noted that the maximum bandwidth BW observed historically max Data updated in real time;

[0184] Exponentially Weighted Moving Average Bandwidth Estimation EWMA bw The specific way to obtain is as follows:

[0185] Select a smoothing factor b, which is a value between 0 and 1, usually in the range of 0.1 to 0.3, to determine the weight of new observations relative to old observations;

[0186] At each time point t, obtain the bandwidth observation value bw at the current moment;

[0187] It is obtained by multiplying the smoothing factor b by the bandwidth observation value bw at the current time point, and adding the bandwidth observation value bw at the previous time point multiplied by (1-b);

[0188] For frame rate adjustment, according to the formula The frame rate FR after real-time dynamic adjustment is calculated, where FR is the adjusted frame rate, FR 0 Is the base frame rate, EWMA la is the exponentially weighted moving average delay, and T is the target delay threshold. It should be noted that the target delay threshold is a preset delay value used to determine the delay range that the system should respond to. This value can be set according to the specific needs of the application scenario. For example, in a video conference, you may want the delay to be no more than 200 milliseconds, so TT can be set to 200 milliseconds.

[0189] It should be noted that the exponentially weighted moving average lag EWMA la The specific way to obtain is as follows:

[0190] Select a smoothing factor α, which is a value between 0 and 1, usually in the range of 0.1 to 0.3, to determine the weight of new observations relative to old observations;

[0191] At each time point t, obtain the current network delay observation value la;

[0192] It is obtained by multiplying the network delay observation value la at the current time point by the smoothing factor α, and adding the network delay observation value la at the previous time point multiplied by (1-α);

[0193] According to the calculated bit rate B after real-time dynamic adjustment and the frame rate FR after real-time dynamic adjustment, the parameter data of the video communication needs to be dynamically adjusted in real time.

[0194] It should be noted that not only can the parameters be adjusted according to the needs of different scenarios, but they can also be dynamically optimized according to the real-time network conditions, which improves the creativity and applicability of the solution. These formulas can help the system respond to network changes more intelligently, thereby providing a more stable and high-quality video communication experience;

[0195] As a further embodiment, the specific working steps of the interactive feedback module are as follows:

[0196] Design a user-friendly feedback interface to allow users to provide feedback during or after the video communication and obtain the user's feedback score; it should be noted that the user's feedback is provided in the form of a score, which can be an integer from 1 to 5 or 1 to 10;

[0197] Real-time monitoring of user interaction behaviors, such as muting and turning off the camera. In video communications, user interaction behaviors, such as muting and turning off the camera, are usually considered negative behaviors because they may mean a decline in user experience or an increase in communication barriers.

[0198] The user's feedback score and the user's interactive behavior are used as feedback information and transmitted to the processing and analysis module.

[0199] As a further embodiment, the specific working steps of the processing and analysis module are as follows:

[0200] Collect user rating data and count the number of user interactions during video communication;

[0201] Evaluate and classify user rating data and the number of user interactions during video communication into three levels: high, medium, and low;

[0202] If the user rating is "high" and the number of negative interactions is "low", the feedback information is marked as good;

[0203] If the user rating is "Medium" and the number of negative interactions is "Low" or "Medium", mark the feedback as good;

[0204] If the user rating is "Medium or Low" and the number of negative interactions is "Medium" or "High", the feedback information is marked as Medium;

[0205] If the user rating is "low" or the number of negative interactions is "high", the feedback information is marked as poor;

[0206] If the feedback information is excellent, the parameter setting of the adaptive adjustment unit can already meet the user's needs and no adjustment is required;

[0207] If the feedback information is good, the parameter setting of the adaptive adjustment unit is adjusted by increasing the bit rate and frame rate by 5% based on the parameter setting of the adaptive adjustment unit;

[0208] If the feedback information is medium, the parameter setting of the adaptive adjustment unit is adjusted, and the bit rate and frame rate are increased by 10% based on the parameter setting of the adaptive adjustment unit;

[0209] If the feedback information is poor, the parameter setting of the adaptive adjustment unit is adjusted, then the bit rate and frame rate are increased by 20% based on the parameter setting of the adaptive adjustment unit.

[0210] It should be noted that by adjusting parameters according to user feedback, user needs can be met more accurately and the overall user experience and satisfaction can be improved. Increasing the bit rate and frame rate can improve the clarity and fluency of video communication, especially when network conditions permit, and can better utilize the available bandwidth. For feedback of "good" and "medium", moderately increasing the bit rate and frame rate can improve the communication quality without sacrificing too much bandwidth; for feedback of "poor", significantly increasing the bit rate and frame rate can be used as a quick response measure to improve the user experience as soon as possible. Only when the user feedback is "excellent" will no adjustment be made, which helps to avoid unnecessary waste of resources and ensures that the system maintains the economy of resource use while meeting user needs;

[0211] As a further embodiment, the specific working steps of evaluating and classifying the user rating data and the number of user interactions during the video communication process include:

[0212] Perform statistical analysis on user rating data and the number of user interactions during video communications, including calculation of historical averages and standard deviations;

[0213] If the user rating data and the number of user interactions during video communication are less than one standard deviation below the mean, it is a low grade;

[0214] If the user rating data and the number of user interactions during video communication are within one standard deviation of the mean, it is considered medium;

[0215] If the user rating data and the number of user interactions during video communication are abnormal by one standard deviation above the average, it is considered a high level. Setting the threshold by the historical average and standard deviation makes the classification more objective and reduces the influence of subjective judgment;

[0216] The present invention also proposes a method for intelligent processing and analysis of audio and video in video communication, comprising the following steps:

[0217] Step 1: Collect real-time video stream data in video communication, identify the current video communication scene according to the real-time video stream data, classify according to the scene, and obtain the current video communication scene classification result;

[0218] Step 2: Switch the corresponding audio and video processing strategy according to the communication scenario classification result, and transmit the processing strategy signal;

[0219] Step 3: Generate real-time processing and analysis results based on real-time video stream data, processing strategy signals and feedback information.

[0220] Working principle:

[0221] The audio and video acquisition module transmits the collected real-time video stream data to the scene recognition module to provide input for scene recognition, and adjusts the acquisition parameters, such as resolution and frame rate, according to the feedback from the adaptive adjustment unit to adapt to different processing requirements;

[0222] The scene recognition module receives data from the audio and video acquisition module and outputs the scene classification result to the adaptive adjustment unit;

[0223] The interactive feedback module transmits the user feedback information to the processing and analysis module for personalized processing, and adjusts the collection method and content of the user feedback according to the scene classification result of the scene recognition module;

[0224] The processing and analysis module receives the real-time video stream data from the audio and video acquisition module and the processing strategy signal from the adaptive adjustment module, and adjusts the processing strategy according to the user feedback information from the interactive feedback module to meet the user's needs;

[0225] The audio and video acquisition module is the starting point of the data flow. Its data is used by the scene recognition module and the processing and analysis module. The output of the scene recognition module (scene classification result) is used by the adaptive adjustment unit and the processing and analysis module. The output of the interactive feedback module (user feedback information) is used by the processing and analysis module.

[0226] The adaptive adjustment unit adjusts the processing strategy according to the output of the scene recognition module and transmits it to the processing analysis module, and the processing analysis module adjusts the processing strategy according to the output of the interactive feedback module;

[0227] The output of the processing and analysis module can be fed back to the audio and video acquisition module to optimize the data acquisition quality; at the same time, the processing and analysis results can also be fed back to the interactive feedback module for further collecting user feedback;

[0228] It should also be noted that the system automatically identifies the current communication scenario (such as video conferencing, online education, etc.) through the scene recognition module, and according to the recognition result, the adaptive adjustment unit dynamically switches to the audio and video processing strategy that best suits the current scenario. This adaptive mechanism enables the system to flexibly respond to the needs of different scenarios, optimize resource allocation, and improve processing efficiency;

[0229] The multimodal data fusion module integrates multiple data such as video, audio and text, and uses advanced algorithms to understand and associate information between different modalities. This enables the system to provide richer and more accurate analysis results, such as evaluating learning effects by analyzing students' facial expressions and voice feedback in online education scenarios, or quickly identifying and responding to emergencies through video analysis in emergency response scenarios;

[0230] The system processes audio and video data in real time through the real-time analysis and processing module, ensuring fast response and low-latency communication experience in scenarios with high real-time requirements such as emergency response and video conferencing. This module optimizes the use of computing resources through intelligent scheduling and resource management to meet real-time requirements;

[0231] The system dynamically adjusts the priority and resource allocation of analysis tasks based on the current network conditions and computing resources. This dynamic resource management strategy helps to optimize resource consumption while ensuring real-time performance, especially on resource-constrained mobile devices. It can balance battery usage and processing performance and extend the device's usage time.

[0232] The above are only preferred implementations of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical staff in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of this template.

Claims

1. An intelligent processing and analysis system for audio and video in video communication, characterized in that: It includes an audio and video acquisition module, which is used to collect real-time video stream data in video communication and transmit it; A scene recognition module, the scene recognition module includes a scene classification unit and an adaptive adjustment unit, the scene classification unit is used to receive the real-time video stream data of the audio and video acquisition module, and identify the scene of the current video communication according to the real-time video stream data, classify according to the scene, obtain the classification result of the current video communication scene, and transmit it to the adaptive adjustment unit; The adaptive adjustment unit is used to receive the communication scenario classification result, switch the corresponding audio and video processing strategy according to the communication scenario classification result, and transmit the processing strategy signal; An interactive feedback module, which is used to collect and transmit user feedback information on video communication; A processing and analysis module, the processing and analysis module is used to receive the real-time video stream data of the audio and video acquisition module, the processing strategy signal of the adaptive adjustment module and the feedback information of the interactive feedback module to generate a real-time processing and analysis result; The specific working steps of the scene classification unit are as follows: The classification results are set in advance, including video conferencing, online education, social video calls, and emergency response, and the known category features of each classification result are obtained; Receive real-time video stream data from the audio and video acquisition module, and extract key features from the real-time video stream data, including color features, texture features, shape features, motion features, and space features; Calculate the similarity between the extracted key features and the known category features of each classification result; Construct a similarity matrix to record the similarity between all key features and known category features of each classification result; Setting a classification threshold for each classification result, comparing the calculated similarity with the classification threshold, obtaining a classification result of the current video communication scene, and transmitting the result to the adaptive adjustment unit; Parse the real-time video stream into a continuous sequence of image frames, extract individual image frames from the video stream at fixed time intervals or based on specific event trigger sampling; Convert a color image frame to a grayscale image; Apply SIFT algorithm to grayscale images to detect key points and extract corresponding feature descriptors; Calculate the gradient magnitude and direction of each pixel in the image frame and generate a directional gradient histogram; The SIFT feature descriptor and the HOG feature vector are combined to form a comprehensive feature vector.

2. The intelligent processing and analysis system for audio and video in video communication according to claim 1, characterized in that: The specific working steps of the audio and video acquisition module are as follows: Detect available audio and video input devices; Synchronously capture the video and audio streams in the video communication through the audio and video input device to obtain real-time video stream data; Compress and encode the captured real-time video stream data; The compressed and encoded real-time video stream data is transmitted to the scene recognition module.

3. The intelligent processing and analysis system for audio and video in video communication according to claim 1, characterized in that: The specific working mode of the adaptive adjustment unit is as follows: Presetting the parameter data of the video communication, specifically including video parameters: resolution, frame rate, bit rate, and adaptability preference; Resolution, frame rate, and bit rate include three levels: high, medium, and low; Audio parameters: sampling rate, encoding bit rate and audio processing; The sampling rate and encoding bit rate include three levels: high, medium and low; Network parameters: bandwidth estimation and packet loss compensation; Obtain the classification result of the current video communication scene, and if the classification result is a video conference, adjust the video parameters to: resolution: high; Frame rate: high; Bit rate: high; Adaptive preference is to maintain frame rate; The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: high; audio processing: turn on echo cancellation and automatic gain control; The network parameters are adjusted as follows: Bandwidth estimation: real-time monitoring of network conditions and dynamic adjustment of bit rate; Packet loss compensation: enabling packet loss compensation mechanism; If the classification result is online education, the video parameters are adjusted as follows: Resolution: Medium; Frame rate: low; Bit rate: medium; Adaptive preference is adjusted to ensure video quality; The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control; The network parameters are adjusted as follows: Bandwidth estimation: monitor network conditions in real time and adjust the bit rate dynamically; Congestion control: enable congestion control; If the classification result is a social video call, the video parameters are adjusted as follows: Resolution: Medium; Frame rate: Medium; Bit rate: Medium; Adaptive preference: Keep balanced; The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control; The network parameters are adjusted as follows: Packet loss compensation: Enable packet loss compensation mechanism; Congestion control: Enable congestion control; If the classification result is emergency response, the video parameters are adjusted to: resolution: low; Frame rate: low; Bit rate: low; Adaptive preference: prioritize frame rate; The audio parameters are adjusted as follows: sampling rate: high; encoding bit rate: medium; audio processing: turn on echo cancellation and automatic gain control; The network parameters are adjusted as follows: Bandwidth estimation: real-time monitoring of network conditions and dynamic adjustment of bit rate; Packet loss compensation: enabling packet loss compensation mechanism; A processing strategy signal is generated according to the parameter adjustment and transmitted to the processing and analysis module.

4. The intelligent processing and analysis system for audio and video in video communication according to claim 3, characterized in that: When the adaptive adjustment unit is working, the parameter data of the video communication needs to be dynamically adjusted in real time, and the specific steps are as follows: For bit rate adjustment, according to the formula The bit rate B after real-time dynamic adjustment is calculated, where B0 is the basic bit rate, PL is the packet loss rate, and EWMA bw is the exponentially weighted moving average bandwidth estimate, BW max is the maximum bandwidth observed historically; For frame rate adjustment, according to the formula Calculate the real-time dynamically adjusted frame rate FR, where FR is the adjusted frame rate, FR0 is the base frame rate, and EWMA la is the exponentially weighted moving average delay, T is the target delay threshold; According to the calculated bit rate B after real-time dynamic adjustment and the frame rate FR after real-time dynamic adjustment, the parameter data of the video communication needs to be dynamically adjusted in real time.

5. The intelligent processing and analysis system for audio and video in video communication according to claim 1, characterized in that: The specific working steps of the interactive feedback module are as follows: Design a user-friendly feedback interface that allows users to provide feedback during or after the video call and obtain user feedback scores; Monitor user interaction behavior in real time; The user's feedback score and the user's interactive behavior are used as feedback information and transmitted to the processing and analysis module.

6. The intelligent processing and analysis system for audio and video in video communication according to claim 1, characterized in that: The specific working steps of the processing and analysis module are as follows: Collect user rating data and count the number of user interactions during video communication; Evaluate and classify user rating data and the number of user interactions during video communication into three levels: high, medium, and low; If the user rating is "high" and the number of negative interactions is "low", mark the feedback information as good; If the user rating is "Medium" and the number of negative interactions is "Low" or "Medium", mark the feedback as good; If the user rating is "Medium or Low" and the number of negative interactions is "Medium" or "High", mark the feedback as Medium; If the user rating is "low" or the number of negative interactions is "high", mark the feedback information as poor; If the feedback information is excellent, the parameter setting of the adaptive adjustment unit can already meet the user's needs and no adjustment is required; If the feedback information is good, the parameter setting of the adaptive adjustment unit is adjusted by increasing the bit rate and frame rate by 5% based on the parameter setting of the adaptive adjustment unit; If the feedback information is medium, the parameter setting of the adaptive adjustment unit is adjusted, and the bit rate and frame rate are increased by 10% based on the parameter setting of the adaptive adjustment unit; If the feedback information is poor, the parameter setting of the adaptive adjustment unit is adjusted, then the bit rate and frame rate are increased by 20% based on the parameter setting of the adaptive adjustment unit.

7. The intelligent processing and analysis system for audio and video in video communication according to claim 6, characterized in that: The specific working steps of evaluating and classifying the user rating data and the number of user interactions during the video communication process include: Perform statistical analysis on user rating data and the number of user interactions during video communications, including calculation of historical averages and standard deviations; If the user rating data and the number of user interactions during video communication are less than one standard deviation below the mean, it is a low grade; If the user rating data and the number of user interactions during video communication are within one standard deviation of the mean, it is considered medium; If the user rating data and the number of user interactions during video communication are abnormal by one standard deviation above the mean, it is considered a high level.

8. A method for intelligent processing and analysis of audio and video in video communication, applicable to a system for intelligent processing and analysis of audio and video in video communication according to claims 1 to 7, characterized in that: The following steps are involved: Step 1: Collect real-time video stream data in video communication, identify the current video communication scene according to the real-time video stream data, classify according to the scene, and obtain the current video communication scene classification result; Step 2: Switch the corresponding audio and video processing strategy according to the communication scenario classification result, and transmit the processing strategy signal; Step 3: Generate real-time processing and analysis results based on real-time video stream data, processing strategy signals and feedback information.

Citation Information

Patent Citations

  • Mobile streaming media QoE (Quality of Experience) correction method and server

    CN104796443A

  • Audio and video communication method and device, computer device and readable storage medium

    CN110225285A

  • Acoustic anomaly detection method for end-to-end unsupervised deep support network

    CN110706720A

  • Threshold determination method and system, storage medium and terminal

    CN111814990A

  • Image retrieval method and device, electronic equipment and readable storage medium

    CN111930985A

Cited By

  • Audio and video processing system and method supporting AI intelligent repair technology

    CN121126017A