Audio and video transmission quality monitoring method, system and device and storage medium

The quality evaluation and transmission method adjustment of audio and video data is performed through multimodal fusion deep learning algorithms, and combined with big data analysis technology to optimize the transmission strategy, the problem of lower audio and video transmission quality in complex network environments is solved, and efficient and accurate audio and video data transmission is achieved.

CN120050477APending Publication Date: 2025-05-27SHENZHEN XINGYIMEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197205.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has reduced audio and video transmission quality in complex network environments due to frequent switching, response delay, mismatch of content characteristics, etc.

Method used

The multimodal fusion deep learning algorithm is used to evaluate the quality of the original audio and video data, and the transmission method is adjusted based on the evaluation results, and the user feedback is collected in combination with big data analysis technology, factors affecting the transmission quality are determined, and the transmission method is further adjusted.

Benefits of technology

Real-time and accurate monitoring of audio and video transmission quality is achieved, the accuracy of detecting audio and video communication quality is improved, and the efficiency of audio and video data transmission is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050477A_ABST
    Figure CN120050477A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video transmission quality monitoring method, system and device and a storage medium, and the method comprises the steps: obtaining original audio and video data, carrying out the quality evaluation of the original audio and video data through employing a multi-modal fusion deep learning algorithm to obtain a quality evaluation result, and carrying out the monitoring of the audio and video transmission quality according to the quality evaluation result. And adjusting the transmission mode of the original audio and video data to obtain a first transmission mode. Obtaining feedback data of a user on the transmission quality of the audio and video data transmitted in the first transmission mode, analyzing the feedback data by using a big data analysis technology to determine factors influencing the transmission quality, and adjusting the first transmission mode according to the factors to obtain a second transmission mode, and transmitting the audio and video data in the second transmission mode. According to the embodiment of the invention, the audio and video transmission quality can be monitored in real time, and the accuracy of audio and video communication quality in the audio and video transmission process is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software communication, and in particular, to a method, system, device, and storage medium for monitoring the quality of audio and video transmission. Background Art

[0002] With the continuous development of network technology, audio and video transmission has become an indispensable part of people's daily lives. Quality monitoring of audio and video plays a crucial role in multiple fields such as multimedia communication, online education, remote work, and video entertainment.

[0003] In related technologies, the use of adaptive bitrate streaming based on network bandwidth and the dynamic adjustment of the video bitrate to adapt to different network conditions have improved the quality of audio and video data transmission.

[0004] However, the audio and video data transmission methods of the prior art may lead to a decrease in the quality of audio and video transmission due to factors such as frequent switching, response delay, and content characteristic mismatch in a complex network environment. Summary of the Invention

[0005] This application provides a method and system for monitoring the quality of audio and video transmission, which is used to accurately monitor the quality of audio and video transmission in real time and improve the accuracy of detecting the quality of audio and video communication during the audio and video transmission process.

[0006] In a first aspect, this application provides a method for monitoring the quality of audio and video transmission, and the method includes the following steps: Obtain the original audio and video data; Use a multi-modal fusion deep learning algorithm to perform quality assessment on the original audio and video data to obtain a quality assessment result; According to the quality assessment result, adjust the transmission method of the original audio and video data to obtain a first transmission method; Obtain feedback data on the transmission quality of the audio and video data transmitted in the first transmission method, use big data analysis technology to analyze the feedback data to determine the factors affecting the transmission quality, and adjust the first transmission method according to the factors to obtain a second transmission method; Transmit the audio and video data in the second transmission method.

[0007] Optionally, obtaining the original audio and video data specifically includes: Collect initial audio data and initial video data from an audio and video source; Convert the audio signal in the initial audio data into a time-frequency domain signal through short-time Fourier transform, perform Mel scale transformation on the time-frequency domain signal to obtain a Mel spectrum, and extract audio features from the Mel spectrum; Decompose the initial video data to form multiple image frames, adjust the multiple image frames to the same size, extract static image features from the image frames, and extract dynamic image features from the video frame sequence of the initial video data; Combine the audio features, the static image features, and the dynamic image features into the original audio-visual data.

[0008] Optionally, use a multi-modal fusion deep learning algorithm to perform quality assessment on the original audio-visual data to obtain a quality assessment result, specifically including: Construct a network structure model, which includes a convolutional neural network and a recurrent neural network. Through the network structure model, perform feature layer fusion, decision layer fusion, and intermediate layer fusion on the features of the original audio-visual data to obtain fused audio-visual data containing the audio features, the static image features, and the dynamic image features; Collect multiple pieces of the fused audio-visual data to form a fused audio-visual data set; Use the fused audio-visual data set to train the deep learning model, and adjust the model parameters through an optimization algorithm and a loss function; Through the deep learning model, process the fused audio-visual data and the original audio-visual data to obtain a quality assessment result.

[0009] Optionally, according to the quality assessment result, adjust the transmission method of the original audio-visual data to obtain a first transmission method, specifically including: Establish a long short-term memory network deep learning model according to the network condition and user device behavior. According to the quality assessment result, use the long short-term memory network deep learning model to obtain a prediction result of the transmission quality, and adjust the transmission method according to the prediction result. The transmission method includes an encoding method and a transmission rate.

[0010] Optionally, by obtaining feedback data on the transmission quality of the audio-visual data transmitted in the first transmission method, use big data analysis technology to analyze the feedback data to determine the factors affecting the transmission quality, and adjust the first transmission method according to the factors to obtain a second transmission method, specifically including: Obtain feedback data on the transmission quality of the audio-visual data transmitted in the adjusted transmission method; Use a data management service to analyze the feedback data and determine the factors affecting the transmission quality; Obtain the transmission requirements of the audio-visual data; Use an elastic computing service to select a computing instance according to the factors and the transmission requirements of the audio-visual data, and adjust the transmission method according to the factors.

[0011] Optionally, analyze the feedback data using the data management service to determine the factors affecting the transmission quality, specifically including: Upload the feedback data to an Amazon S3 bucket; Format the feedback data uploaded to the Amazon S3 bucket into CSV format and load it into a data processing pipeline; Process the feedback data in the data processing pipeline through the Apache Kafka messaging framework to determine the factors affecting the transmission quality.

[0012] Optionally, before adjusting the transmission method of the original audio-visual data according to the quality assessment result to obtain the first transmission method, specifically including: Perform the computing tasks of quality assessment and adjustment strategies at network edge nodes through edge computing technology; Deploy a deep learning model on the network edge node, adaptively adjust the deep learning model, and through the deep learning model, perform quality assessment on the collected original audio-visual data to obtain a quality assessment result.

[0013] In a second aspect, an embodiment of the present application provides a monitoring system for audio-visual transmission quality, including: A data acquisition module for acquiring original audio-visual data; A quality assessment module for performing quality assessment on the original audio-visual data using a multi-modal fusion deep learning algorithm to obtain a quality assessment result; An optimization transmission module for adjusting the transmission method of the original audio-visual data according to the quality assessment result; A data analysis module for acquiring feedback data on the transmission quality of the audio-visual data transmitted by the adjusted transmission method, analyzing the feedback data using big data analysis technology to determine the factors affecting the transmission quality, and adjusting the adjusted transmission method according to the factors; An execution module for transmitting the audio-visual data in the second transmission method.

[0014] In a third aspect, an embodiment of the present application provides a monitoring device for audio-visual transmission quality. The monitoring device for audio-visual transmission quality includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the monitoring device for audio-visual transmission quality to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0015] Fourthly, an embodiment of the present application provides a computer-readable storage medium, including instructions, which, when running on a monitoring device for audio and video transmission quality, cause the monitoring device for audio and video transmission quality to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0016] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Through a multi-modal fusion deep learning algorithm, multiple features of audio and video data are integrated to obtain a quality evaluation result. According to the quality evaluation result, the transmission method is adjusted, which can optimize data transmission in real time to adapt to different network conditions and user requirements. By collecting and analyzing user feedback, a large amount of user feedback data is stored, providing a rich data resource for subsequent quality optimization. Big data analysis technology assists in identifying factors affecting transmission quality, enabling the system to continuously learn and improve, and enhancing the accuracy of detecting audio and video communication quality during the audio and video transmission process.

[0017] 2. Preprocess the collected initial audio data and initial video data. The short-time Fourier transform helps capture the dynamic characteristics of the audio signal, providing a basis for subsequent feature extraction. The Mel-scale transform simulates the perception of the human auditory system for the audio signal, thereby extracting audio features more relevant to human hearing. By decomposing video data into image frames and extracting static and dynamic image features, the video content can be analyzed from different perspectives. Unifying the data format facilitates feature fusion and analysis.

[0018] 1. Integrate the audio features, static image features, and dynamic image features in the original audio and video data, and integrate multiple features of the audio and video data to obtain a quality evaluation result. Through a deep learning model, multi-modal fusion and quality evaluation of the original audio and video data are realized, improving the accuracy of the quality evaluation of the original audio and video data.

[0019] 2. Through a long short-term memory network deep learning model, accurate prediction of the transmission quality of the original audio and video data is realized, and accordingly, the transmission method of the original audio and video data is adjusted to improve the transmission efficiency of the audio and video data.

[0020] 5. Use the data management service to analyze user feedback data, quickly understand user needs and system performance, and provide data support for adjusting the transmission method. Use the elastic computing service to select a computing instance according to the factors and the transmission requirements of the audio and video data, adjust the transmission method according to the factors, dynamically select the computing instance, and realize the on-demand allocation and optimized use of resources. The data management service can process large-scale data streams and perform real-time analysis of user feedback, improving the efficiency and real-time performance of data processing.

[0021] 6. By completing the calculation tasks of quality assessment and adjustment strategies at the network edge nodes, the latency of data processing is significantly reduced. Processing the original audio-visual data at the edge nodes can reduce the amount of data that needs to be transmitted to the central server, relieve the pressure on the network bandwidth, and improve the efficiency of data processing at the same time. Description of the Drawings

[0022] Figure 1 is a schematic flowchart of the audio-visual transmission quality monitoring method in an embodiment of the present application; Figure 2 is a schematic structural diagram of the audio-visual transmission quality monitoring system in an embodiment of the present application; Figure 3 is a schematic structural diagram of the electronic device in an embodiment of the present application.

[0023] Reference numerals: 201, data acquisition module; 202, quality assessment module; 203, optimized transmission module; 204, data analysis module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Embodiments

[0024] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0025] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0026] In the description of the embodiments of the present application, the meaning of the term "plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0027] Figure 1It is a schematic flowchart of the audio - video transmission quality monitoring method in an embodiment of the present application.

[0028] Please refer to Figure 1 , the audio - video transmission quality monitoring method in an embodiment of the present application includes: S10. Obtain the original audio - video data; Specifically, collect the initial audio data and initial video data from the audio - video source. Convert the audio signal in the initial audio data into a time - frequency domain signal through the short - time Fourier transform, and perform the Mel - scale transform on the time - frequency domain signal to obtain the Mel - spectrum. Extract the audio features from the Mel - spectrum. Decompose the initial video data to form multiple image frames, adjust the multiple image frames to the same size, extract the static image features from the image frames, and extract the dynamic image features from the video frame sequence of the initial video data. Combine the audio features, static image features, and dynamic image features into the original audio - video data.

[0029] Among them, convert the audio signal into a time - frequency domain representation through the short - time Fourier transform (STFT): Use the discrete Fourier transform (DFT) and the fast Fourier transform (FFT) to calculate the frequency components in each short - time window. The DFT decomposes the short - time signal into a series of sine and cosine functions, and the fast Fourier transform (FFT) multiplies these basic frequency components by the time window and sums them to obtain a complex signal containing all frequency components. The short - time Fourier transform (STFT) helps to capture the dynamic characteristics of the audio signal and provides a basis for subsequent feature extraction. Perform an inverse operation on the result of the short - time Fourier transform (STFT) to obtain a set of Mel - frequency coefficients. Use the inverse Mel - frequency filter to convert the Mel - frequency coefficients back to the audio signal in the original time domain to obtain a feature vector with time and frequency information. The feature vector is the audio feature of the original audio data. The Mel - scale transform simulates the perception of the human auditory system for audio signals, thereby extracting audio features more relevant to human hearing.

[0030] Use an image - processing library such as OpenCV to calculate the scaling factor, adjust the size of the video frames in the original video data according to the scaling factor, segment the video frames in the original video data into multiple image frames, and adjust them to the same size. Start from the local features of static images, such as color, texture, shape, etc., extract the static image features from the image frames, and extract the dynamic image features from the video frame sequence of the initial video data by comparing the positional relationships between consecutive video frames; Convert the format of the adjusted audio - video data into a tensor for processing through a deep - learning model such as TensorFlow.

[0031] S11. Use a multi-modal fusion deep learning algorithm to perform quality assessment on the original audio-visual data to obtain a quality assessment result; Specifically, construct a network structure model. The network structure model includes a convolutional neural network and a recurrent neural network. Through the network structure model, perform feature layer fusion, decision layer fusion, and intermediate layer fusion on the features of the original audio-visual data to obtain fused audio-visual data containing the audio features, the static image features, and the dynamic image features. Collect multiple pieces of the fused audio-visual data to form a fused audio-visual data set, use the fused audio-visual data set to train the deep learning model, and adjust the model parameters through an optimization algorithm and a loss function. Through the deep learning model, process the fused audio-visual data and the original audio-visual data to obtain a quality assessment result.

[0032] Among them, use the tf.keras API to construct a model including a convolutional neural network and a recurrent neural network to fuse the audio features, the static image features, and the dynamic image features: The input layer is the starting point of the convolutional neural network, accepting the original data as input. The convolutional layer extracts local features from the global feature map. The activation layer performs a non-linear mapping on the output result of the convolutional layer. The pooling layer further screens the features, which can effectively reduce the number of parameters required for subsequent network layers. Then, through the recurrent neural network, perform feature layer fusion, decision layer fusion, and intermediate layer fusion on the audio features, the static image features, and the dynamic image features. Feature layer fusion immediately concatenates the audio and video features after feature extraction and sends them as a unified input to the subsequent network for processing. Decision layer fusion processes the audio and video features separately. Intermediate layer fusion is usually achieved through an attention mechanism or other feature selection methods. Obtain fused audio-visual data that simultaneously contains the audio features, the static image features, and the dynamic image features; Collect audio-visual data from public data platforms such as YouTube and Bilibili. Ensure that the collected data has high audio and video quality by setting filtering conditions (such as resolution and bit rate).

[0033] Preprocess the collected audio-visual data. The preprocessing steps include but are not limited to: Data deduplication: Remove duplicate or redundant data through a deduplication algorithm (such as hash value matching, deep learning model, etc.); Data augmentation: Increase the diversity of the data set by performing operations such as cropping, scaling, and rotating on the original data; Data annotation: Label each audio-visual frame to improve the generalization ability of the model. Multiple annotation methods can be used, such as manual annotation, semi-automatic annotation, full-automatic annotation, etc.

[0034] Fuse the audio-video dataset. The fusion methods include, but are not limited to: timestamp fusion and spatial location fusion. Timestamp fusion: Arrange data segments of different durations in chronological order according to the timestamp information of the audio-video data to ensure the coherence of the audio-video data; Spatial location fusion: Perform spatial normalization on the audio-video data by combining Geographic Information System (GIS) data or other spatial location information.

[0035] After completing the fusion of the audio-video dataset, evaluate the dataset to ensure its quality and usability. The evaluation metrics can include: dataset size, data category coverage, data quality, etc. According to the evaluation results, we can further adjust the data fusion strategy to improve the overall performance of the dataset.

[0036] Use the collected fused audio-video dataset to train the model through a model that includes a convolutional neural network and a recurrent neural network. Use optimization algorithms (such as Adam or RMSprop) and loss functions (such as mean squared error or cross entropy) to train the model. During the training process, callback functions such as early stopping (EarlyStopping) can be used to avoid overfitting. During the training process, the model will learn to extract effective information from the fused audio-video data and map it to the corresponding labels or categories. After training, use the deep learning model to process the fused audio-video data and the original audio-video data to obtain the F1 score of the prediction result of the deep learning model, and this F1 score is the quality evaluation result.

[0037] S12. According to the quality evaluation result, adjust the transmission method of the original audio-video data to obtain the first transmission method; Specifically, establish a long short-term memory network deep learning model according to the network condition and user device behavior. Use the long short-term memory network deep learning model to obtain the prediction result of the transmission quality according to the quality evaluation result, and adjust the transmission method according to the prediction result. The transmission method includes the encoding method and the transmission rate.

[0038] Among them, according to the quality evaluation result, use the long short-term memory (LSTM) network deep learning model to predict the transmission quality and automatically adjust the encoding method and the transmission rate to obtain the first transmission method.

[0039] Use tf.dataAPI to build an efficient data input pipeline to collect relevant network and device behavior data. Based on the collected data, you need to extract or build features that can represent network conditions and user behavior. Use tf.kerasAPI to define one or more LSTM layers in the long short delay memory network (LSTM) deep learning model to process sequence data, and add dense layers to make final predictions or classifications. Use optimization algorithms (such as Adam or RMSprop) and loss functions (such as mean square error or cross entropy) to train the model. During training, callback functions such as EarlyStopping can be used to avoid overfitting. After the model is trained, it is applied to real-time data to predict network conditions or user device behavior to support decision-making. Based on the model's prediction results, the encoding method and transmission rate can be automatically adjusted to optimize network performance and user experience.

[0040] In order to improve the transmission efficiency of the original audio and video data, edge computing technology is introduced to delegate some of the computing tasks of quality assessment and adjustment strategies to the edge nodes of the network to reduce data transmission delays and improve processing speed.

[0041] Optionally, before adjusting the transmission mode of the original audio and video data according to the quality evaluation result to obtain the first transmission mode, the method further includes: Perform quality assessment and strategy adjustment computing tasks at network edge nodes through edge computing technology; A deep learning model is deployed on the network edge node, and the deep learning model is adaptively adjusted. The quality of the collected original audio and video data is evaluated through the deep learning model to obtain a quality evaluation result, and the audio and video transmission method is adjusted according to the quality evaluation result.

[0042] Obtain audio and video data generated by users, including audio, video frames, video streams, etc. These data can be uploaded to network edge nodes in real time through mobile phones, cameras and other devices; Deploy a deep learning model on the edge node of the network and adaptively adjust the model so that it can quickly respond to user needs on the user side. The deep learning model directly analyzes and processes the audio and video data on the edge node without sending all the data to a remote server for processing. This significantly reduces data transmission delays, improves real-time performance, and saves bandwidth and computing resources. The deployment and training of the deep learning model can refer to step S11. According to the quality assessment results, the transmission strategy can be adjusted to adjust the transmission method of the original audio and video data. Please refer to step S12, which will not be repeated here.

[0043] S13. Obtain feedback data on the transmission quality of the audio - video data transmitted in the first transmission mode, analyze the feedback data using big data analysis technology to determine the factors affecting the transmission quality, and adjust the first transmission mode according to the factors to obtain a second transmission mode.

[0044] Specifically, collect the feedback of users on the transmission quality through the data management service in the cloud, and use big data analysis technology to mine and analyze the feedback data to obtain a second transmission mode and optimize the transmission quality in real - time. Analyze the feedback data using the data management service, upload the feedback data to an Amazon S3 bucket, format the feedback data uploaded to the Amazon S3 bucket into CSV format and load it into a data processing pipeline. Process the feedback data in the data processing pipeline through the Apache Kafka messaging framework to determine the factors affecting the transmission quality.

[0045] Among them, collect real - time user feedback data through methods such as HTTP request headers or Cookies. Upload the collected feedback data to an Amazon S3 bucket. This process can also be implemented using any appropriate method, for example, by transmitting data between the user and the service provider through an application. Once the data is collected in the S3 bucket, it can be further processed and analyzed; Convert the feedback data from the original format to CSV format and load it into a data processing pipeline. This step can be achieved through various programming languages and libraries. Load the data into the pipeline for data analysis; Use the Apache Kafka messaging framework to integrate real - time data streams into a data processing pipeline for real - time user feedback analysis. Kafka is a high - throughput, distributed, and reliable messaging system that can easily integrate a large amount of data streams into a system, analyze the feedback data of different users, and obtain analysis results: certain specific factors such as encoding methods and transmission rates have a certain impact on the transmission quality; Visualize the analysis results. You can use AWS data analysis and visualization tools such as Amazon QuickSight or Amazon Denali, etc. Convert the analysis results into easy - to - understand charts and dashboards. These tools can help present the data in an intuitive way, and determine the degree of association between certain specific factors such as encoding methods, transmission rates and transmission quality through the data presented in the charts and dashboards to determine the factors affecting the transmission quality.

[0046] Model on the cloud computing service EC2 and define performance and cost parameters for different workloads. Determine the tasks to be executed based on the factors affecting the transmission quality and the requirements of the workload. This model will determine the instances that can run under a specific workload. For example, if the task requires real-time data analysis, an instance type with higher computing power can be selected. If the task only needs to process a small amount of data, an instance type with lower computing power can be selected; Set the start, stop, and reservation policies of the instances according to the real-time nature and computing requirements of the tasks. For example, if the task is real-time, multiple instances can be selected to handle a large number of concurrent requests. If the task is non-real-time, a smaller number of instances can be selected to reduce costs; Monitor the real-time changes of the factors affecting the transmission quality and the requirements of the workload. Dynamically adjust the number and type of instances according to the execution of the tasks and the changes in the user feedback data. For example, when the user feedback increases, more instances can be automatically started to support real-time data analysis.

[0047] S14. Transmit the audio and video data in the second transmission mode.

[0048] The above describes the audio and video transmission quality monitoring method in the embodiments of the present application. Next, in combination with the above audio and video transmission quality monitoring method, the audio and video transmission quality monitoring system 200 in the embodiments of the application will be described in detail.

[0049] Please refer to Figure 2 , which is a schematic structural diagram of an audio and video transmission quality monitoring system provided by the embodiments of the present application. The audio and video transmission quality monitoring system 200 specifically includes: A data acquisition module 201 for acquiring original audio and video data; A quality evaluation module 202 for evaluating the quality of the original audio and video data to obtain a quality evaluation result; An optimized transmission module 203 for adjusting the transmission strategy according to the quality evaluation result to adjust the original audio and video data; A data analysis module 204, by collecting user feedback data on the transmission quality of the adjusted audio and video data, using big data analysis technology to mine and analyze the feedback data to determine the factors affecting the transmission quality, and adjusting the adjusted audio and video data according to the factors; An execution module 205 for transmitting the audio and video data in the second transmission mode.

[0050] Optionally, the data acquisition module 201 is used for: Collect initial audio data and initial video data from an audio and video source; Convert the audio signal in the initial audio data into a time-frequency domain signal through short-time Fourier transform, and perform Mel-scale transformation on the time-frequency domain signal to obtain a Mel spectrogram, and extract audio features from the Mel spectrogram; Decompose the initial video data to form multiple image frames, adjust the multiple image frames to the same size, extract static image features from the image frames, and extract dynamic image features from the video frame sequence of the initial video data; Combine the audio features, the static image features and the dynamic image features into the original audio-visual data.

[0051] Optionally, the quality assessment module 202 is used for: Construct a network structure model, which includes a convolutional neural network and a recurrent neural network. Through the network structure model, perform feature layer fusion, decision layer fusion and intermediate layer fusion on the features of the original audio-visual data to obtain the fused audio-visual data containing the audio features, the static image features and the dynamic image features; Collect multiple pieces of the fused audio-visual data to form a fused audio-visual data set; Use the fused audio-visual data set to train the deep learning model, and adjust the model parameters through an optimization algorithm and a loss function; Through the deep learning model, process the fused audio-visual data and the original audio-visual data to obtain a quality assessment result.

[0052] Optionally, the optimization transmission module 203 is used for: Establish a long short-term memory network deep learning model according to the network condition and the user device behavior. Use the long short-term memory network deep learning model to obtain a prediction result of the transmission quality according to the quality assessment result, and adjust the transmission method according to the prediction result. The transmission method includes an encoding method and a transmission rate.

[0053] Optionally, the data analysis module 204 is used for: Obtain feedback data on the transmission quality of the audio-visual data transmitted by the adjusted transmission method by the user; Use the data management service to analyze the feedback data and determine the factors affecting the transmission quality; Obtain the transmission requirements of the audio-visual data; Use the elastic computing service to select a computing instance according to the factors and the transmission requirements of the audio-visual data, and adjust the transmission method according to the factors.

[0054] Optionally, the data analysis module 204 is used for: Upload the feedback data to the Amazon S3 bucket; Format the feedback data uploaded to the Amazon S3 bucket into CSV format and load it into the data processing pipeline; Process the feedback data in the data processing pipeline through the Apache Kafka messaging framework to determine the factors affecting the transmission quality. Optionally, the system further includes an edge computing module, and the edge computing module is used for: Perform computing tasks for quality assessment and adjustment strategies at network edge nodes through edge computing technology; Deploy a deep learning model on the network edge node, adaptively adjust the deep learning model, and through the deep learning model, perform quality assessment on the collected original audio and video data to obtain a quality assessment result.

[0055] Optionally, the execution module 205 is used to transmit the audio and video data in the second transmission mode.

[0056] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be repeated here.

[0057] This embodiment also discloses an electronic device. Refer to Figure 3 , the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.

[0058] Among them, the communication bus 302 is used to realize the connection and communication between these components.

[0059] Among them, the user interface 303 may include a display screen (Display), a camera (Camera), a microphone (Microphone). Optionally, the user interface may further include a standard wired interface and a wireless interface.

[0060] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0061] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 505, it executes various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0062] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As Figure 3 shown, the memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for monitoring the audio and video transmission quality.

[0063] In Figure 3In the electronic device shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 301 can be used to call the application program for monitoring the audio and video transmission quality stored in the memory 305. When executed by one or more processors, the electronic device is enabled to execute the method of one or more of the above embodiments.

[0064] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0065] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0066] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0067] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0068] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0069] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 505 and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned memory 505 includes various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0070] The foregoing are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the disclosure of the specification, those skilled in the art will readily think of other implementation manners of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for monitoring audio and video transmission quality, characterized in that: include: Get original audio and video data; Performing quality assessment on the original audio and video data using a multimodal fusion deep learning algorithm to obtain a quality assessment result; According to the quality evaluation result, adjusting the transmission mode of the original audio and video data to obtain a first transmission mode; Obtaining user feedback data on the transmission quality of audio and video data transmitted in the first transmission mode, analyzing the feedback data using big data analysis technology to determine factors affecting the transmission quality, and adjusting the first transmission mode according to the factors to obtain a second transmission mode; The audio and video data are transmitted in the second transmission mode.

2. The method according to claim 1, characterized in that The obtaining of original audio and video data specifically includes: Collecting initial audio data and initial video data from an audio and video source; Converting the audio signal in the initial audio data into a time-frequency domain signal by short-time Fourier transform, performing Mel scale transformation on the time-frequency domain signal to obtain a Mel spectrum, and extracting audio features from the Mel spectrum; Decomposing the initial video data to form a plurality of image frames, adjusting the plurality of image frames to the same size, extracting static image features from the image frames, and extracting dynamic image features from a video frame sequence of the initial video data; The audio features, the static image features and the dynamic image features are combined into original audio and video data.

3. The method according to claim 1, characterized in that The method of using a multimodal fusion deep learning algorithm to perform quality assessment on the original audio and video data to obtain a quality assessment result specifically includes: Constructing a network structure model, wherein the network structure model includes a convolutional neural network and a recurrent neural network, and performing feature layer fusion, decision layer fusion and intermediate layer fusion on the features of the original audio and video data through the network structure model to obtain fused audio and video data containing the audio features, the static image features and the dynamic image features; Collecting a plurality of the fused audio and video data to form a fused audio and video data set; Using the fused audio and video data set to train the deep learning model, and adjusting the model parameters through an optimization algorithm and a loss function; The fused audio and video data and the original audio and video data are processed by the deep learning model to obtain a quality assessment result.

4. The method according to claim 1, characterized in that According to the quality evaluation result, the transmission mode of the original audio and video data is adjusted to obtain a first transmission mode, which specifically includes: A long-short delay memory network deep learning model is established according to network conditions and user equipment behavior, and a prediction result of transmission quality is obtained using the long-short delay memory network deep learning model according to quality assessment results. The transmission mode is adjusted according to the prediction result, and the transmission mode includes a coding mode and a transmission rate.

5. The method according to claim 1, characterized in that The obtaining of user feedback data on the transmission quality of the audio and video data transmitted in the first transmission mode, analyzing the feedback data using big data analysis technology to determine factors affecting the transmission quality, and adjusting the first transmission mode according to the factors to obtain the second transmission mode specifically includes: Obtaining user feedback data on the transmission quality of audio and video data transmitted by the adjusted transmission method; analyzing the feedback data using a data management service to determine factors affecting transmission quality; Obtaining transmission requirements for the audio and video data; Use elastic computing services, select computing instances based on the factors and the transmission requirements of the audio and video data, and adjust the transmission method based on the factors.

6. The method according to claim 5, characterized in that The using the data management service to analyze the feedback data to determine factors affecting the transmission quality specifically includes: Uploading the feedback data to an Amazon S3 bucket; Format the feedback data uploaded to the Amazon S3 bucket into CSV format and load it into the data processing pipeline; The feedback data in the data processing pipeline is processed through the Apache Kafka messaging framework to determine factors affecting the transmission quality.

7. The method according to claim 1, characterized in that Before adjusting the transmission mode of the original audio and video data according to the quality evaluation result to obtain the first transmission mode, the method further includes: Perform quality assessment and strategy adjustment computing tasks at network edge nodes through edge computing technology; A deep learning model is deployed on the network edge node, the deep learning model is adaptively adjusted, and the quality of the collected original audio and video data is evaluated through the deep learning model to obtain a quality evaluation result.

8. A monitoring system for audio and video transmission quality, characterized in that: include: A data acquisition module, used to acquire original audio and video data; A quality assessment module, used to perform quality assessment on the original audio and video data using a multimodal fusion deep learning algorithm to obtain a quality assessment result; An optimized transmission module, used for adjusting the transmission mode of the original audio and video data according to the quality evaluation result; A data analysis module, used to obtain user feedback data on the transmission quality of audio and video data transmitted by the adjusted transmission mode, analyze the feedback data using big data analysis technology to determine factors affecting the transmission quality, and adjust the adjusted transmission mode according to the factors; The execution module transmits the audio and video data in the second transmission mode.

9. An intelligent monitoring device for audio and video transmission quality, characterized in that: include: one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program codes, wherein the computer program codes include computer instructions, and the one or more processors call the computer instructions to enable the intelligent monitoring device for audio and video transmission quality to execute the method as described in any one of claims 1-7.

10. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on the intelligent monitoring device for audio and video transmission quality, the intelligent monitoring device for audio and video transmission quality executes the method as described in any one of claims 1 to 7.