Intelligent audio-video hybrid transmission strategy and system

CN119420989BActive Publication Date: 2026-08-11SHENZHEN SHOWTOP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,由于硬件故障、软件冲突或网络延迟等原因,视频画面异常与音频数据异常事件频发,影响用户体验和业务运营

Benefits of technology

[0055] This invention discloses an intelligent audio-video hybrid transmission strategy and system. During a monitoring period, the system captures user interactions and audio-video data in real time, separating the video and audio streams. The video stream undergoes frame extraction and parameter comparison to generate a first anomaly assessment. Interaction data is used to locate the scene, perform feature and semantic analysis, and compare with database information to obtain a second anomaly assessment. Combining the two assessment results, audio data anomalies are located and optimized according to a preset strategy to generate optimized audio. This invention effectively performs anomaly analysis and reasonable hybrid transmission of audio and video data, effectively optimizing the transmission of audio and video data across multiple terminals and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119420989B_ABST
    Figure CN119420989B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent audio-video hybrid transmission strategy and system. During a monitoring period, the system captures user interactions and audio-video data in real time, separating the video and audio streams. The video stream undergoes frame extraction and parameter comparison to generate a first anomaly assessment. Interaction data is used to locate the scene, perform feature and semantic analysis, and compare with database information to obtain a second anomaly assessment. Combining the two assessment results, audio data anomalies are located and optimized according to a preset strategy to generate optimized audio. This invention effectively performs anomaly analysis and reasonable hybrid transmission of audio and video data, effectively optimizing the transmission of audio and video data across multiple terminals and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video data analysis, and more specifically, to intelligent audio and video hybrid transmission strategies and systems. Background Technology

[0002] Multi-screen system platforms are becoming increasingly prevalent in business, education, and entertainment. However, due to hardware failures, software conflicts, or network latency, video and audio data anomalies occur frequently, impacting user experience and business operations. Existing audio and video anomaly detection methods often rely on manual monitoring or single-device feedback, resulting in drawbacks such as slow response, high false alarm rates, and inability to adapt to dynamic changes. Furthermore, they cannot provide rapid, real-time, and accurate anomaly assessment and adjustment, leading to a poor user experience. Therefore, developing an efficient intelligent audio and video hybrid data analysis strategy and related methods is a crucial issue that urgently needs to be addressed. Summary of the Invention

[0003] This invention overcomes the shortcomings of the prior art and proposes an intelligent audio and video hybrid transmission strategy and system.

[0004] The first aspect of this invention provides an intelligent audio and video hybrid transmission strategy, comprising:

[0005] Within a monitoring cycle, user interaction data and real-time audio and video data are acquired in real time. The real-time audio and video data are then separated to obtain real-time video data and real-time audio data.

[0006] Video frames are extracted from real-time video data to obtain image data. The basic parameters of the image data are analyzed and compared with the preset image parameters to generate the first anomaly assessment result.

[0007] Based on user interaction data, the interaction time nodes and screen interaction type information are obtained. Based on the interaction time nodes, the corresponding screens are extracted from the screen data to form detection screen data. Feature extraction, image recognition and text recognition are performed on the detection screen data to generate semantic recognition results.

[0008] By filtering real-time comparative semantic information of the screen from the screen semantic database using user interaction data, calculating the similarity between the semantic recognition result and the real-time comparative semantic information, and performing screen anomaly analysis through threshold comparison, a second anomaly assessment result based on semantic analysis is generated.

[0009] Based on the first and second anomaly assessment results, obtain the anomaly time nodes and anomaly status information. Based on the anomaly time nodes, extract the corresponding audio segment data from the real-time audio and video data. Through a preset audio processing strategy, optimize the audio segment data based on the anomaly status information to generate optimized audio segment data.

[0010] The first and second anomaly assessment results are sent to the user terminal device. The user terminal device adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with the optimized audio segment data to obtain the second audio and video data.

[0011] In this solution, within a monitoring cycle, user interaction data and real-time audio / video data are acquired in real time, and the real-time audio / video data is separated to obtain real-time video data and real-time audio data. Specifically:

[0012] The system platform is used to count the number of currently connected user terminals and analyze the current audio and video data transmission volume.

[0013] The monitoring cycle is set based on the number of user terminals and the amount of audio and video data transmitted.

[0014] Within a monitoring period, user interaction data and real-time audio and video data are acquired in real time through user terminals.

[0015] The real-time audio and video data is parsed and separated into audio and video data to obtain real-time video data and real-time audio data.

[0016] In this solution, the process of extracting video frames from real-time video data to obtain image data, analyzing and comparing the basic parameters of the image data with preset image parameters, and generating a first anomaly assessment result, specifically involves:

[0017] Keyframes are extracted from real-time video data to obtain an image set;

[0018] The image set is preprocessed with image enhancement, smoothing, and normalization to obtain image data;

[0019] Based on the image data, statistical image parameter information is obtained, including resolution, aspect ratio, and frame rate, and basic parameters are obtained.

[0020] The basic parameters are compared with the preset screen parameters. If there is a difference, a first anomaly assessment result is generated based on the parameter difference. At the same time, the anomaly time node is recorded in the first anomaly assessment result.

[0021] In this solution, the steps of obtaining interaction time nodes and screen interaction type information based on user interaction data, extracting corresponding screen segments from the screen data based on the interaction time nodes to form detection screen data, and performing feature extraction, image recognition, and text recognition on the detection screen data to generate semantic recognition results are as follows:

[0022] Obtain interaction commands and interaction time points through user interaction data;

[0023] Set the time nodes before and after the interaction based on the interaction time nodes;

[0024] By extracting corresponding image information from the image data at the time nodes before and after the interaction, the first image data and the second image data are obtained.

[0025] A CNN-based image recognition model is constructed to perform image segmentation and image recognition on the first and second image data respectively, and the image recognition results are obtained.

[0026] The image recognition results are converted into first semantic information;

[0027] Based on OCR technology, the first image data and the second image data are preprocessed by grayscale conversion and noise reduction, and the characters are recognized and extracted by the OCR system to obtain the second semantic information.

[0028] The first semantic information and the second semantic information are integrated to form the semantic recognition result.

[0029] In this solution, the steps of filtering real-time comparative semantic information from the image semantic database using user interaction data, calculating the similarity between the semantic recognition result and the real-time comparative semantic information, performing image anomaly analysis through threshold comparison, and generating a second anomaly assessment result based on semantic analysis are as follows:

[0030] Get interaction commands based on user interaction data;

[0031] Based on interactive commands, the corresponding real-time comparison semantic information of the screen is filtered and obtained from the system database;

[0032] The first and second semantic information in the semantic recognition results are compared with the real-time semantic information to calculate the similarity. During the similarity calculation process, the semantic information similarity between the texts is calculated by the bag-of-words model and cosine similarity, and the semantic similarity is finally obtained.

[0033] The semantic similarity and the resemblance threshold are compared to identify the anomalies. Anomaly assessment is performed through comparative analysis. At the same time, the time points of the anomalies are recorded and a second anomaly assessment result is generated.

[0034] In this solution, the steps of obtaining abnormal time nodes and abnormal situation information based on the first and second abnormal assessment results, extracting corresponding audio segment data from real-time audio and video data based on the abnormal time nodes, optimizing the audio segment data based on the abnormal situation information using a preset audio processing strategy, and generating optimized audio segment data are as follows:

[0035] Information on abnormal time points and abnormal conditions is obtained by combining the results of the first and second anomaly assessments.

[0036] Data segmentation and extraction are performed on real-time audio and video data based on abnormal time nodes to obtain corresponding audio segment data;

[0037] By using a preset audio processing strategy, the system matches processing strategies based on abnormal situation information, and optimizes audio segment data based on the matching results to generate optimized audio segment data.

[0038] In this solution, the first anomaly assessment result and the second anomaly assessment result are sent to the user terminal device. The user terminal device adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with optimized audio segment data to obtain the second audio-visual data. Specifically:

[0039] The first anomaly assessment result and the second anomaly assessment result are sent to the user terminal device;

[0040] By analyzing abnormal situations through user terminal devices, and based on preset adjustment instructions, corresponding adjustment instructions are selected for real-time adjustment of different abnormal screen states.

[0041] Record the adjusted video data to obtain optimized video data. Mix the optimized video data with the optimized audio segment data to obtain the second audio and video data.

[0042] A second aspect of the present invention also provides an intelligent audio and video hybrid transmission system, the system comprising: a memory and a processor, wherein the memory includes an intelligent audio and video hybrid transmission program, and the intelligent audio and video hybrid transmission program, when executed by the processor, performs the following steps:

[0043] Within a monitoring cycle, user interaction data and real-time audio and video data are acquired in real time. The real-time audio and video data are then separated to obtain real-time video data and real-time audio data.

[0044] Video frames are extracted from real-time video data to obtain image data. The basic parameters of the image data are analyzed and compared with the preset image parameters to generate the first anomaly assessment result.

[0045] Based on user interaction data, the interaction time nodes and screen interaction type information are obtained. Based on the interaction time nodes, the corresponding screens are extracted from the screen data to form detection screen data. Feature extraction, image recognition and text recognition are performed on the detection screen data to generate semantic recognition results.

[0046] By filtering real-time comparative semantic information of the screen from the screen semantic database using user interaction data, calculating the similarity between the semantic recognition result and the real-time comparative semantic information, and performing screen anomaly analysis through threshold comparison, a second anomaly assessment result based on semantic analysis is generated.

[0047] Based on the first and second anomaly assessment results, obtain the anomaly time nodes and anomaly status information. Based on the anomaly time nodes, extract the corresponding audio segment data from the real-time audio and video data. Through a preset audio processing strategy, optimize the audio segment data based on the anomaly status information to generate optimized audio segment data.

[0048] The first and second anomaly assessment results are sent to the user terminal device. The user terminal device adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with the optimized audio segment data to obtain the second audio and video data.

[0049] In this solution, within a monitoring cycle, user interaction data and real-time audio / video data are acquired in real time, and the real-time audio / video data is separated to obtain real-time video data and real-time audio data. Specifically:

[0050] The system platform is used to count the number of currently connected user terminals and analyze the current audio and video data transmission volume.

[0051] The monitoring cycle is set based on the number of user terminals and the amount of audio and video data transmitted.

[0052] Within a monitoring period, user interaction data and real-time audio and video data are acquired in real time through user terminals.

[0053] The real-time audio and video data is parsed and separated into audio and video data to obtain real-time video data and real-time audio data.

[0054] A third aspect of the present invention also provides a computer-readable storage medium including an intelligent audio-video hybrid transmission program, wherein when the intelligent audio-video hybrid transmission program is executed by a processor, it implements the steps of the intelligent audio-video hybrid transmission strategy as described in any of the preceding claims.

[0055] This invention discloses an intelligent audio-video hybrid transmission strategy and system. During a monitoring period, the system captures user interactions and audio-video data in real time, separating the video and audio streams. The video stream undergoes frame extraction and parameter comparison to generate a first anomaly assessment. Interaction data is used to locate the scene, perform feature and semantic analysis, and compare with database information to obtain a second anomaly assessment. Combining the two assessment results, audio data anomalies are located and optimized according to a preset strategy to generate optimized audio. This invention effectively performs anomaly analysis and reasonable hybrid transmission of audio and video data, effectively optimizing the transmission of audio and video data across multiple terminals and improving user experience. Attached Figure Description

[0056] Figure 1A flowchart of an intelligent audio-video hybrid transmission strategy according to the present invention is shown;

[0057] Figure 2 A block diagram of an intelligent audio and video hybrid transmission system according to the present invention is shown. Detailed Implementation

[0058] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0059] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0060] Figure 1 A flowchart of an intelligent audio and video hybrid transmission strategy according to the present invention is shown.

[0061] like Figure 1 As shown, the first aspect of the present invention provides an intelligent audio and video hybrid transmission strategy, comprising:

[0062] S102: Within a monitoring cycle, real-time user interaction data and real-time audio and video data are acquired, and the real-time audio and video data are separated to obtain real-time video data and real-time audio data.

[0063] S104: Extract video frames from real-time video data to obtain image data, analyze and compare whether the basic parameters of the image data are consistent with the preset image parameters, and generate the first anomaly assessment result.

[0064] S106: Based on user interaction data, obtain interaction time nodes and screen interaction type information; based on interaction time nodes, extract corresponding screens from screen data to form detection screen data; perform feature extraction, image recognition and text recognition on the detection screen data to generate semantic recognition results.

[0065] S108: Select real-time comparison semantic information of the screen from the screen semantic database through user interaction data, calculate the similarity between the semantic recognition result and the real-time comparison semantic information, perform screen anomaly analysis through threshold comparison, and generate a second anomaly evaluation result based on semantic analysis.

[0066] S110: Based on the first and second anomaly assessment results, obtain the anomaly time node and anomaly status information; based on the anomaly time node, extract the corresponding audio segment data from the real-time audio and video data; through a preset audio processing strategy, optimize the audio segment data based on the anomaly status information to generate optimized audio segment data.

[0067] S112, the first anomaly assessment result and the second anomaly assessment result are sent to the user terminal device, the user terminal device adjusts the abnormal state of the screen in real time, and generates optimized real-time video data. The real-time video data is mixed with the optimized audio segment data to obtain the second audio and video data.

[0068] This invention overcomes the limitations of existing audio and video devices in content transmission and playback, enabling flexible transmission and strategic playback of video and audio content through anomaly analysis and response strategies.

[0069] This system includes an audio / video transmission module and an audio / video policy playback module. The audio / video transmission module is responsible for transmitting the audio / video content of one device to multiple user devices (or terminal devices on various platforms) in real time and stably. This module utilizes preset audio encoding technologies and network transmission protocols to ensure high-quality audio transmission and low latency. Simultaneously, the module also has automatic audio signal recognition and classification functions to distinguish between the device's own audio and externally transmitted audio. The audio / video policy playback module is responsible for recording abnormal situations, policy applications, and historical processing policies.

[0070] The mixed audio and video data can be recorded and used to replace the original distorted and abnormal video data and distorted audio data. In real-time audio and video terminal applications, such as surveillance applications, when abnormal video occurs, the audio and video optimization transmission strategy of this invention can optimize the distorted and abnormal video data and store it in the surveillance database, replacing the original lossy audio and video data and improving the adaptability to audio and video data transmission optimization.

[0071] In addition, real-time analysis of anomalies enables real-time recovery of distorted audio and video data in the later stages. Timely anomaly analysis and data optimization and recovery can improve the authenticity of audio and video data and effectively optimize subsequent data maintenance work, especially in applications based on audio and video monitoring needs.

[0072] According to an embodiment of the present invention, the step of acquiring user interaction data and real-time audio / video data in real time within a monitoring period, and separating the real-time audio / video data to obtain real-time video data and real-time audio data specifically involves:

[0073] The system platform is used to count the number of currently connected user terminals and analyze the current audio and video data transmission volume.

[0074] The monitoring cycle is set based on the number of user terminals and the amount of audio and video data transmitted.

[0075] Within a monitoring period, user interaction data and real-time audio and video data are acquired in real time through user terminals.

[0076] The real-time audio and video data is parsed and separated into audio and video data to obtain real-time video data and real-time audio data.

[0077] It should be noted that in this invention, the system platform is the central platform for image monitoring, responsible for tasks such as image capture and analysis of different user terminals. The monitoring cycle is dynamically set based on the current audio and video data transmission volume of the user terminals. Specifically, it is based on the analysis of the number of user terminals and the audio and video data transmission volume. The larger the volume of data transmission, the shorter the monitoring cycle can be set to reduce the system's computing power consumption. Alternatively, the user can set the cycle according to their actual needs.

[0078] User terminals include terminal interaction devices and display devices, used to display visual and audio content to users, such as mobile terminals, tablets, and computers. The system platform connects to the user terminals via the Internet and is responsible for real-time visual analysis, real-time audio analysis, and anomaly monitoring and control of different user terminals.

[0079] According to an embodiment of the present invention, the step of extracting video frames from real-time video data to obtain image data, analyzing and comparing whether the basic parameters of the image data are consistent with preset image parameters, and generating a first anomaly assessment result specifically includes:

[0080] Keyframes are extracted from real-time video data to obtain an image set;

[0081] The image set is preprocessed with image enhancement, smoothing, and normalization to obtain image data;

[0082] Based on the image data, statistical image parameter information is obtained, including resolution, aspect ratio, and frame rate, and basic parameters are obtained.

[0083] The basic parameters are compared with the preset screen parameters. If there is a difference, a first anomaly assessment result is generated based on the parameter difference. At the same time, the anomaly time node is recorded in the first anomaly assessment result.

[0084] It should be noted that the first anomaly assessment result is based on image transmission parameters and is used to assess transmission loss, effectively reflecting the impact of image anomalies during transmission. The first anomaly assessment result includes anomaly information and the corresponding time point of the anomaly occurrence.

[0085] According to an embodiment of the present invention, the step of obtaining interaction time nodes and screen interaction type information based on user interaction data, extracting corresponding screen segments from screen data based on interaction time nodes to form detection screen data, and performing feature extraction, image recognition, and text recognition on the detection screen data to generate semantic recognition results specifically includes:

[0086] Obtain interaction commands and interaction time points through user interaction data;

[0087] Set the time nodes before and after the interaction based on the interaction time nodes;

[0088] By extracting corresponding image information from the image data at the time nodes before and after the interaction, the first image data and the second image data are obtained.

[0089] A CNN-based image recognition model is constructed to perform image segmentation and image recognition on the first and second image data respectively, and the image recognition results are obtained.

[0090] The image recognition results are converted into first semantic information;

[0091] Based on OCR technology, the first image data and the second image data are preprocessed by grayscale conversion and noise reduction, and the characters are recognized and extracted by the OCR system to obtain the second semantic information.

[0092] The first semantic information and the second semantic information are integrated to form the semantic recognition result.

[0093] It should be noted that the pre-interaction and post-interaction time points are used to analyze whether the changes in the screen before and after the interaction conform to the corresponding interaction procedure. For example, whether the screen changes accordingly when the user performs operations such as pausing, zooming, entering a certain interface, or exiting a certain interface. The image recognition model improves the recognition rate by training on historical screen information. The screen recognition results correspond to different results based on different scenarios. For example, in applications such as commercial tablets and multi-monitor commercial conferences, the recognition typically identifies the specific operation interface of the software screen, such as the primary interface and secondary interfaces, to determine whether there are any screen anomalies during the interaction process. OCR technology is a text recognition technology that performs text recognition and extraction on images. The OCR system is based on a CNN network for character feature extraction and recognition.

[0094] According to an embodiment of the present invention, the step of filtering real-time comparison semantic information of the screen from the screen semantic database through user interaction data, calculating the similarity between the semantic recognition result and the real-time comparison semantic information, performing screen anomaly analysis through threshold comparison, and generating a second anomaly evaluation result based on semantic analysis specifically includes:

[0095] Get interaction commands based on user interaction data;

[0096] Based on interactive commands, the corresponding real-time comparison semantic information of the screen is filtered and obtained from the system database;

[0097] The first and second semantic information in the semantic recognition results are compared with the real-time semantic information to calculate the similarity. During the similarity calculation process, the semantic information similarity between the texts is calculated by the bag-of-words model and cosine similarity, and the semantic similarity is finally obtained.

[0098] The semantic similarity and the resemblance threshold are compared to identify the anomalies. Anomaly assessment is performed through comparative analysis. At the same time, the time points of the anomalies are recorded and a second anomaly assessment result is generated.

[0099] It should be noted that the real-time image comparison semantic information includes first and second comparison semantic information, which are compared with the first semantic information and the second semantic information, respectively. In this invention, the first anomaly assessment result analyzes image anomalies from the perspective of transmitted data, while the second anomaly assessment result analyzes image anomalies from the perspective of interaction anomalies; the two reflect different anomalies. Anomalies include black screen, white screen, distorted screen, green screen, freeze, third-party application dialog box lag, page loading lag, etc. The second anomaly assessment result includes anomaly information and the corresponding time point of the anomaly occurrence.

[0100] According to an embodiment of the present invention, the step of obtaining abnormal time nodes and abnormal situation information based on the first abnormality assessment result and the second abnormality assessment result, extracting corresponding audio segment data from real-time audio and video data based on the abnormal time nodes, optimizing the audio segment data based on the abnormal situation information through a preset audio processing strategy, and generating optimized audio segment data, specifically includes:

[0101] Information on abnormal time points and abnormal conditions is obtained by combining the results of the first and second anomaly assessments.

[0102] Data segmentation and extraction are performed on real-time audio and video data based on abnormal time nodes to obtain corresponding audio segment data;

[0103] By using a preset audio processing strategy, the system matches processing strategies based on abnormal situation information, and optimizes audio segment data based on the matching results to generate optimized audio segment data.

[0104] It is worth mentioning that during audio and video data transmission, audio data anomalies are generally difficult to assess, analyze, and recover in real time. Therefore, this invention performs real-time assessment and analysis of video anomalies. When video anomalies occur, the corresponding audio data often exhibits audio distortion, excessive noise, or repetitive playback of certain audio segments, all caused by data anomalies. Therefore, this invention records video anomalies at specific time points, extracts and analyzes the corresponding audio data, and performs appropriate audio masking and noise reduction to reduce the impact of audio anomalies. The adjustment process is based on a preset audio processing strategy; different anomalies correspond to different optimization processes. For example, when the video freezes or stutters, audio can be masked; when the video is blurry or distorted, the audio data often exhibits distortion and excessive noise, in which case the corresponding optimization strategy could be audio noise reduction. The preset audio processing strategy can be set by the user. Optimized audio segments are used to replace the original audio segments.

[0105] According to an embodiment of the present invention, the step of sending the first anomaly assessment result and the second anomaly assessment result to the user terminal device, adjusting the abnormal state of the screen in real time through the user terminal device, generating optimized real-time video data, and mixing the real-time video data with optimized audio segment data to obtain the second audio-video data specifically involves:

[0106] The first anomaly assessment result and the second anomaly assessment result are sent to the user terminal device;

[0107] By analyzing abnormal situations through user terminal devices, and based on preset adjustment instructions, corresponding adjustment instructions are selected for real-time adjustment of different abnormal screen states.

[0108] Record the adjusted video data to obtain optimized video data. Mix the optimized video data with the optimized audio segment data to obtain the second audio and video data.

[0109] It should be noted that the preset adjustment command is an application adjustment strategy pre-set by the user. For example, if the screen abnormality is a decrease in screen resolution or a blurry screen, the screen can be adjusted by commands such as adjusting the transmission line or adjusting the network resource allocation. If the screen abnormality is a situation where the operation interface freezes, the screen can be adjusted by sending a refresh command.

[0110] Audio and video data mixing includes steps such as audio and video data encoding, time calibration, and encapsulation of synthesized data.

[0111] During audio and video processing, after adjustments, the optimized audio and video data is recorded, replacing the original audio and video data, and stored in the system database.

[0112] In audio and video image control, there are two methods: real-time control and delayed control. In delayed data control, after optimizing and analyzing to obtain the second audio and video data, the original abnormal image audio and video data is replaced, thereby achieving effective data maintenance of audio and video data and reducing unnecessary data storage.

[0113] According to an embodiment of the present invention, it further includes:

[0114] Within N monitoring periods, the system platform acquires N audio transmission data from the user.

[0115] Using a user as the unit of analysis, a waveform comparison method based on audio is used to compare the waveform similarity of N audio transmission data. If the similarity is greater than a preset similarity threshold, the user is marked as the user to be analyzed.

[0116] Analyze all users and record all users to be analyzed to form a user group to be analyzed;

[0117] In the user group to be analyzed, feature extraction based on frequency, amplitude and audio waveform is performed on all audio transmission data of each user to obtain user audio features;

[0118] Using user audio features as clustering sample data, and based on the DBSCAN clustering model, the user group to be analyzed is clustered and grouped into multiple user groups;

[0119] Within N monitoring periods, monitor and record the audio anomalies and corresponding audio anomaly handling strategies for each user group to obtain audio anomaly handling record information.

[0120] Each user group corresponds to an audio anomaly handling record;

[0121] Using a user group as the analysis unit, high-frequency anomalies are extracted from the audio anomaly processing record information to obtain high-frequency anomaly information. The corresponding high-frequency anomaly information is then correlated with the corresponding processing strategies to obtain general audio anomaly strategy information.

[0122] Analyze all user groups to obtain multiple common audio anomaly policy information, and store the multiple common audio anomaly policy information in the system database;

[0123] In the next monitoring cycle, determine whether the real-time user belongs to the user group to be analyzed. If so, based on the user group to which the real-time user belongs, obtain the general audio anomaly strategy information from the system database to judge and process the audio anomaly.

[0124] The general response strategy (i.e., general audio anomaly strategy information) generated by this invention can be applied to large-scale audio transmission and processing. When the system data processing pressure exceeds the warning threshold, the strategy can be adopted to reduce the system's data analysis pressure while ensuring the audio anomaly handling capability. In particular, it can effectively reduce the system server pressure in large-scale user platforms.

[0125] In the extraction of high-frequency anomalies, specifically, the frequency of audio anomalies is counted, and audio anomalies exceeding a preset frequency are recorded, along with their corresponding processing strategies, to obtain general audio anomaly strategy information. The total audio transmission data for each user includes N audio transmissions from that user.

[0126] The waveform similarity comparison of N audio transmission data involves comparing the audio waveforms of one transmission with the next. This waveform comparison reflects the similarity of audio features, and the comparison process generates multiple similarity values. In the user group analyzed by this invention, the N audio transmission data exhibit a certain degree of similarity. Therefore, each user in this group has certain audio demand habits and regularities in their audio transmission data. For example, in video-on-demand or audio-on-demand platforms, users often request similar audio and video data based on a particular interest when using their devices. When corresponding high-frequency anomalies are identified, the corresponding strategies are imported into the database as cached data. In subsequent real-time audio transmissions, anomaly handling and judgment can be performed using general strategies, reducing the system's data analysis pressure and improving system stability.

[0127] Figure 2 A block diagram of an intelligent audio and video hybrid transmission system according to the present invention is shown.

[0128] A second aspect of the present invention also provides an intelligent audio and video hybrid transmission system 2, the system comprising: a memory 21 and a processor 22, wherein the memory includes an intelligent audio and video hybrid transmission program, and the intelligent audio and video hybrid transmission program, when executed by the processor, performs the following steps:

[0129] Within a monitoring cycle, user interaction data and real-time audio and video data are acquired in real time. The real-time audio and video data are then separated to obtain real-time video data and real-time audio data.

[0130] Video frames are extracted from real-time video data to obtain image data. The basic parameters of the image data are analyzed and compared with the preset image parameters to generate the first anomaly assessment result.

[0131] Based on user interaction data, the interaction time nodes and screen interaction type information are obtained. Based on the interaction time nodes, the corresponding screens are extracted from the screen data to form detection screen data. Feature extraction, image recognition and text recognition are performed on the detection screen data to generate semantic recognition results.

[0132] By filtering real-time comparative semantic information of the screen from the screen semantic database using user interaction data, calculating the similarity between the semantic recognition result and the real-time comparative semantic information, and performing screen anomaly analysis through threshold comparison, a second anomaly assessment result based on semantic analysis is generated.

[0133] Based on the first and second anomaly assessment results, obtain the anomaly time nodes and anomaly status information. Based on the anomaly time nodes, extract the corresponding audio segment data from the real-time audio and video data. Through a preset audio processing strategy, optimize the audio segment data based on the anomaly status information to generate optimized audio segment data.

[0134] The first and second anomaly assessment results are sent to the user terminal device. The user terminal device adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with the optimized audio segment data to obtain the second audio and video data.

[0135] This invention overcomes the limitations of existing audio and video devices in content transmission and playback, enabling flexible transmission and strategic playback of video and audio content through anomaly analysis and response strategies.

[0136] This system includes an audio / video transmission module and an audio / video policy playback module. The audio / video transmission module is responsible for transmitting the audio / video content of one device to multiple user devices (or terminal devices on various platforms) in real time and stably. This module utilizes preset audio encoding technologies and network transmission protocols to ensure high-quality audio transmission and low latency. Simultaneously, the module also has automatic audio signal recognition and classification functions to distinguish between the device's own audio and externally transmitted audio. The audio / video policy playback module is responsible for recording abnormal situations, policy applications, and historical processing policies.

[0137] The mixed audio and video data can be recorded and used to replace the original distorted and abnormal video data and distorted audio data. In real-time audio and video terminal applications, such as surveillance applications, when abnormal video occurs, the audio and video optimization transmission strategy of this invention can optimize the distorted and abnormal video data and store it in the surveillance database, replacing the original lossy audio and video data and improving the adaptability to audio and video data transmission optimization.

[0138] In addition, real-time analysis of anomalies enables real-time recovery of distorted audio and video data in the later stages. Timely anomaly analysis and data optimization and recovery can improve the authenticity of audio and video data and effectively optimize subsequent data maintenance work, especially in applications based on audio and video monitoring needs.

[0139] According to an embodiment of the present invention, the step of acquiring user interaction data and real-time audio / video data in real time within a monitoring period, and separating the real-time audio / video data to obtain real-time video data and real-time audio data specifically involves:

[0140] The system platform is used to count the number of currently connected user terminals and analyze the current audio and video data transmission volume.

[0141] The monitoring cycle is set based on the number of user terminals and the amount of audio and video data transmitted.

[0142] Within a monitoring period, user interaction data and real-time audio and video data are acquired in real time through user terminals.

[0143] The real-time audio and video data is parsed and separated into audio and video data to obtain real-time video data and real-time audio data.

[0144] It should be noted that in this invention, the system platform is the central platform for image monitoring, responsible for tasks such as image capture and analysis of different user terminals. The monitoring cycle is dynamically set based on the current audio and video data transmission volume of the user terminals. Specifically, it is based on the analysis of the number of user terminals and the audio and video data transmission volume. The larger the volume of data transmission, the shorter the monitoring cycle can be set to reduce the system's computing power consumption. Alternatively, the user can set the cycle according to their actual needs.

[0145] User terminals include terminal interaction devices and display devices, used to display visual and audio content to users, such as mobile terminals, tablets, and computers. The system platform connects to the user terminals via the Internet and is responsible for real-time visual analysis, real-time audio analysis, and anomaly monitoring and control of different user terminals.

[0146] According to an embodiment of the present invention, the step of extracting video frames from real-time video data to obtain image data, analyzing and comparing whether the basic parameters of the image data are consistent with preset image parameters, and generating a first anomaly assessment result specifically includes:

[0147] Keyframes are extracted from real-time video data to obtain an image set;

[0148] The image set is preprocessed with image enhancement, smoothing, and normalization to obtain image data;

[0149] Based on the image data, statistical image parameter information is obtained, including resolution, aspect ratio, and frame rate, and basic parameters are obtained.

[0150] The basic parameters are compared with the preset screen parameters. If there is a difference, a first anomaly assessment result is generated based on the parameter difference. At the same time, the anomaly time node is recorded in the first anomaly assessment result.

[0151] It should be noted that the first anomaly assessment result is based on image transmission parameters and is used to assess transmission loss, effectively reflecting the impact of image anomalies during transmission. The first anomaly assessment result includes anomaly information and the corresponding time point of the anomaly occurrence.

[0152] According to an embodiment of the present invention, the step of obtaining interaction time nodes and screen interaction type information based on user interaction data, extracting corresponding screen segments from screen data based on interaction time nodes to form detection screen data, and performing feature extraction, image recognition, and text recognition on the detection screen data to generate semantic recognition results specifically includes:

[0153] Obtain interaction commands and interaction time points through user interaction data;

[0154] Set the time nodes before and after the interaction based on the interaction time nodes;

[0155] By extracting corresponding image information from the image data at the time nodes before and after the interaction, the first image data and the second image data are obtained.

[0156] A CNN-based image recognition model is constructed to perform image segmentation and image recognition on the first and second image data respectively, and the image recognition results are obtained.

[0157] The image recognition results are converted into first semantic information;

[0158] Based on OCR technology, the first image data and the second image data are preprocessed by grayscale conversion and noise reduction, and the characters are recognized and extracted by the OCR system to obtain the second semantic information.

[0159] The first semantic information and the second semantic information are integrated to form the semantic recognition result.

[0160] It should be noted that the pre-interaction and post-interaction time points are used to analyze whether the changes in the screen before and after the interaction conform to the corresponding interaction procedure. For example, whether the screen changes accordingly when the user performs operations such as pausing, zooming, entering a certain interface, or exiting a certain interface. The image recognition model improves the recognition rate by training on historical screen information. The screen recognition results correspond to different results based on different scenarios. For example, in applications such as commercial tablets and multi-monitor commercial conferences, the recognition typically identifies the specific operation interface of the software screen, such as the primary interface and secondary interfaces, to determine whether there are any screen anomalies during the interaction process. OCR technology is a text recognition technology that performs text recognition and extraction on images. The OCR system is based on a CNN network for character feature extraction and recognition.

[0161] According to an embodiment of the present invention, the step of filtering real-time comparison semantic information of the screen from the screen semantic database through user interaction data, calculating the similarity between the semantic recognition result and the real-time comparison semantic information, performing screen anomaly analysis through threshold comparison, and generating a second anomaly evaluation result based on semantic analysis specifically includes:

[0162] Get interaction commands based on user interaction data;

[0163] Based on interactive commands, the corresponding real-time comparison semantic information of the screen is filtered and obtained from the system database;

[0164] The first and second semantic information in the semantic recognition results are compared with the real-time semantic information to calculate the similarity. During the similarity calculation process, the semantic information similarity between the texts is calculated by the bag-of-words model and cosine similarity, and the semantic similarity is finally obtained.

[0165] The semantic similarity and the resemblance threshold are compared to identify the anomalies. Anomaly assessment is performed through comparative analysis. At the same time, the time points of the anomalies are recorded and a second anomaly assessment result is generated.

[0166] It should be noted that the real-time image comparison semantic information includes first and second comparison semantic information, which are compared with the first semantic information and the second semantic information, respectively. In this invention, the first anomaly assessment result analyzes image anomalies from the perspective of transmitted data, while the second anomaly assessment result analyzes image anomalies from the perspective of interaction anomalies; the two reflect different anomalies. Anomalies include black screen, white screen, distorted screen, green screen, freeze, third-party application dialog box lag, page loading lag, etc. The second anomaly assessment result includes anomaly information and the corresponding time point of the anomaly occurrence.

[0167] According to an embodiment of the present invention, the step of obtaining abnormal time nodes and abnormal situation information based on the first abnormality assessment result and the second abnormality assessment result, extracting corresponding audio segment data from real-time audio and video data based on the abnormal time nodes, optimizing the audio segment data based on the abnormal situation information through a preset audio processing strategy, and generating optimized audio segment data, specifically includes:

[0168] Information on abnormal time points and abnormal conditions is obtained by combining the results of the first and second anomaly assessments.

[0169] Data segmentation and extraction are performed on real-time audio and video data based on abnormal time nodes to obtain corresponding audio segment data;

[0170] By using a preset audio processing strategy, the system matches processing strategies based on abnormal situation information, and optimizes audio segment data based on the matching results to generate optimized audio segment data.

[0171] It is worth mentioning that during audio and video data transmission, audio data anomalies are generally difficult to assess, analyze, and recover in real time. Therefore, this invention performs real-time assessment and analysis of video anomalies. When video anomalies occur, the corresponding audio data often exhibits audio distortion, excessive noise, or repetitive playback of certain audio segments, all caused by data anomalies. Therefore, this invention records video anomalies at specific time points, extracts and analyzes the corresponding audio data, and performs appropriate audio masking and noise reduction to reduce the impact of audio anomalies. The adjustment process is based on a preset audio processing strategy; different anomalies correspond to different optimization processes. For example, when the video freezes or stutters, audio can be masked; when the video is blurry or distorted, the audio data often exhibits distortion and excessive noise, in which case the corresponding optimization strategy could be audio noise reduction. The preset audio processing strategy can be set by the user. Optimized audio segments are used to replace the original audio segments.

[0172] According to an embodiment of the present invention, the step of sending the first anomaly assessment result and the second anomaly assessment result to the user terminal device, adjusting the abnormal state of the screen in real time through the user terminal device, generating optimized real-time video data, and mixing the real-time video data with optimized audio segment data to obtain the second audio-video data specifically involves:

[0173] The first anomaly assessment result and the second anomaly assessment result are sent to the user terminal device;

[0174] By analyzing abnormal situations through user terminal devices, and based on preset adjustment instructions, corresponding adjustment instructions are selected for real-time adjustment of different abnormal screen states.

[0175] Record the adjusted video data to obtain optimized video data. Mix the optimized video data with the optimized audio segment data to obtain the second audio and video data.

[0176] It should be noted that the preset adjustment command is an application adjustment strategy pre-set by the user. For example, if the screen abnormality is a decrease in screen resolution or a blurry screen, the screen can be adjusted by commands such as adjusting the transmission line or adjusting the network resource allocation. If the screen abnormality is a situation where the operation interface freezes, the screen can be adjusted by sending a refresh command.

[0177] Audio and video data mixing includes steps such as audio and video data encoding, time calibration, and encapsulation of synthesized data.

[0178] During audio and video processing, after adjustments, the optimized audio and video data is recorded, replacing the original audio and video data, and stored in the system database.

[0179] In audio and video image control, there are two methods: real-time control and delayed control. In delayed data control, after optimizing and analyzing to obtain the second audio and video data, the original abnormal image audio and video data is replaced, thereby achieving effective data maintenance of audio and video data and reducing unnecessary data storage.

[0180] A third aspect of the present invention also provides a computer-readable storage medium including an intelligent audio-video hybrid transmission program, wherein when the intelligent audio-video hybrid transmission program is executed by a processor, it implements the steps of the intelligent audio-video hybrid transmission strategy as described in any of the preceding claims.

[0181] This invention discloses an intelligent audio-video hybrid transmission strategy and system. During a monitoring period, the system captures user interactions and audio-video data in real time, separating the video and audio streams. The video stream undergoes frame extraction and parameter comparison to generate a first anomaly assessment. Interaction data is used to locate the scene, perform feature and semantic analysis, and compare with database information to obtain a second anomaly assessment. Combining the two assessment results, audio data anomalies are located and optimized according to a preset strategy to generate optimized audio. This invention effectively performs anomaly analysis and reasonable hybrid transmission of audio and video data, effectively optimizing the transmission of audio and video data across multiple terminals and improving user experience.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0183] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0184] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0185] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0186] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0187] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent audio-video hybrid transmission strategy, characterized in that, include: Within a monitoring cycle, user interaction data and real-time audio and video data are acquired in real time. The real-time audio and video data are then separated to obtain real-time video data and real-time audio data. Video frames are extracted from real-time video data to obtain image data. The basic parameters of the image data are analyzed and compared with the preset image parameters to generate the first anomaly assessment result. Based on user interaction data, the interaction time nodes and screen interaction type information are obtained. Based on the interaction time nodes, the corresponding screens are extracted from the screen data to form detection screen data. Feature extraction, image recognition and text recognition are performed on the detection screen data to generate semantic recognition results. By filtering real-time comparative semantic information of the screen from the screen semantic database using user interaction data, the semantic recognition results are compared with the real-time comparative semantic information to calculate similarity, and screen anomaly analysis is performed by threshold comparison to generate a second anomaly assessment result based on semantic analysis. Based on the first and second anomaly assessment results, obtain the anomaly time nodes and anomaly status information. Based on the anomaly time nodes, extract the corresponding audio segment data from the real-time audio and video data. Through a preset audio processing strategy, optimize the audio segment data based on the anomaly status information to generate optimized audio segment data. The first and second anomaly assessment results are sent to the user terminal device. The user terminal device adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with the optimized audio segment data to obtain the second audio and video data.

2. The intelligent audio and video hybrid transmission strategy according to claim 1, characterized in that, Within a monitoring cycle, user interaction data and real-time audio / video data are acquired in real time. The real-time audio / video data is then separated to obtain real-time video data and real-time audio data. Specifically: The system platform is used to count the number of currently connected user terminals and analyze the current audio and video data transmission volume. The monitoring cycle is set based on the number of user terminals and the amount of audio and video data transmitted. Within a monitoring period, user interaction data and real-time audio and video data are acquired in real time through user terminals. The real-time audio and video data is parsed and separated into audio and video data to obtain real-time video data and real-time audio data.

3. The intelligent audio and video hybrid transmission strategy according to claim 2, characterized in that, The process of extracting video frames from real-time video data to obtain image data, analyzing and comparing the basic parameters of the image data with preset image parameters, and generating a first anomaly assessment result, specifically involves: Keyframes are extracted from real-time video data to obtain an image set; The image set is preprocessed with image enhancement, smoothing, and normalization to obtain image data; Based on the image data, statistical image parameter information is obtained, including resolution, aspect ratio, and frame rate, and basic parameters are obtained. The basic parameters are compared with the preset screen parameters. If there is a difference, a first anomaly assessment result is generated based on the parameter difference. At the same time, the anomaly time node is recorded in the first anomaly assessment result.

4. The intelligent audio and video hybrid transmission strategy according to claim 3, characterized in that, The process involves obtaining interaction time points and screen interaction type information based on user interaction data, extracting corresponding screen segments from the screen data based on the interaction time points to form detection screen data, and performing feature extraction, image recognition, and text recognition on the detection screen data to generate semantic recognition results. Specifically: Obtain interaction commands and interaction time points through user interaction data; Set the time nodes before and after the interaction based on the interaction time nodes; By extracting corresponding image information from the image data at the time nodes before and after the interaction, the first image data and the second image data are obtained. A CNN-based image recognition model is constructed to perform image segmentation and image recognition on the first and second image data respectively, and the image recognition results are obtained. The image recognition results are converted into first semantic information; Based on OCR technology, the first image data and the second image data are preprocessed by grayscale conversion and noise reduction, and the characters are recognized and extracted by the OCR system to obtain the second semantic information. The first semantic information and the second semantic information are integrated to form the semantic recognition result.

5. The intelligent audio and video hybrid transmission strategy according to claim 4, characterized in that, The process involves filtering real-time comparative semantic information from the image semantic database using user interaction data, calculating the similarity between the semantic recognition result and the real-time comparative semantic information, performing image anomaly analysis through threshold comparison, and generating a second anomaly assessment result based on semantic analysis. Specifically: Get interaction commands based on user interaction data; Based on interactive commands, the corresponding real-time comparison semantic information of the screen is filtered and obtained from the system database; The first and second semantic information in the semantic recognition results are compared with the real-time semantic information to calculate the similarity. During the similarity calculation process, the semantic information similarity between the texts is calculated by the bag-of-words model and cosine similarity, and the semantic similarity is finally obtained. The semantic similarity and the similarity threshold are compared to identify the anomalies. Anomaly assessment is performed through comparative analysis. At the same time, the time points of the anomalies are recorded and a second anomaly assessment result is generated.

6. The intelligent audio and video hybrid transmission strategy according to claim 5, characterized in that, The process involves obtaining abnormal time nodes and abnormal situation information based on the first and second abnormality assessment results, extracting corresponding audio segment data from real-time audio and video data based on the abnormal time nodes, optimizing the audio segment data based on the abnormal situation information using a preset audio processing strategy, and generating optimized audio segment data. Information on abnormal time points and abnormal conditions is obtained by using the results of the first and second abnormal assessments. Data segmentation and extraction are performed on real-time audio and video data based on abnormal time nodes to obtain corresponding audio segment data; By using a preset audio processing strategy, the system matches processing strategies based on abnormal situation information, and optimizes audio segment data based on the matching results to generate optimized audio segment data.

7. The intelligent audio and video hybrid transmission strategy according to claim 6, characterized in that, The process involves sending the first and second anomaly assessment results to the user terminal device, which then adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with the optimized audio segment data to obtain the second audio-visual data. Specifically: The first anomaly assessment result and the second anomaly assessment result are sent to the user terminal device; By analyzing abnormal situations through user terminal devices, and based on preset adjustment instructions, corresponding adjustment instructions are selected for real-time adjustment of different abnormal screen states. Record the adjusted video data to obtain optimized video data. Mix the optimized video data with the optimized audio segment data to obtain the second audio and video data.

8. An intelligent audio and video hybrid transmission system, characterized in that, The system includes: a memory and a processor. The memory includes an intelligent audio and video mixing transmission program. When the intelligent audio and video mixing transmission program is executed by the processor, it performs the following steps: Within a monitoring cycle, user interaction data and real-time audio and video data are acquired in real time. The real-time audio and video data are then separated to obtain real-time video data and real-time audio data. Video frames are extracted from real-time video data to obtain image data. The basic parameters of the image data are analyzed and compared with the preset image parameters to generate the first anomaly assessment result. Based on user interaction data, the interaction time nodes and screen interaction type information are obtained. Based on the interaction time nodes, the corresponding screens are extracted from the screen data to form detection screen data. Feature extraction, image recognition and text recognition are performed on the detection screen data to generate semantic recognition results. By filtering real-time comparative semantic information of the screen from the screen semantic database using user interaction data, the semantic recognition results are compared with the real-time comparative semantic information to calculate similarity, and screen anomaly analysis is performed by threshold comparison to generate a second anomaly assessment result based on semantic analysis. Based on the first and second anomaly assessment results, obtain the anomaly time nodes and anomaly status information. Based on the anomaly time nodes, extract the corresponding audio segment data from the real-time audio and video data. Through a preset audio processing strategy, optimize the audio segment data based on the anomaly status information to generate optimized audio segment data. The first and second anomaly assessment results are sent to the user terminal device. The user terminal device adjusts the abnormal state of the screen in real time and generates optimized real-time video data. The real-time video data is then mixed with the optimized audio segment data to obtain the second audio and video data.

9. The intelligent audio and video hybrid transmission system according to claim 8, characterized in that, Within a monitoring cycle, user interaction data and real-time audio / video data are acquired in real time. The real-time audio / video data is then separated to obtain real-time video data and real-time audio data. Specifically: The system platform is used to count the number of currently connected user terminals and analyze the current audio and video data transmission volume. The monitoring cycle is set based on the number of user terminals and the amount of audio and video data transmitted. Within a monitoring period, user interaction data and real-time audio and video data are acquired in real time through user terminals. The real-time audio and video data is parsed and separated into audio and video data to obtain real-time video data and real-time audio data.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes an intelligent audio and video hybrid transmission program, which, when executed by a processor, implements the steps of the intelligent audio and video hybrid transmission strategy as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video matching method, device and equipment and storage medium

    CN112257595A

  • Audio data analysis method and system in complex scene and storage medium

    CN117116302A