An audio and video data acquisition method and system based on 5G terminal devices

The 5G-based audio-visual data collection method and system address mobility and efficiency issues by using API interfaces and quality scoring for optimized data transmission, improving collection and transmission efficiency.

CN118573959BActive Publication Date: 2025-06-17CHONGQING PINGKEJIE INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410670283.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-06-17
Estimated Expiration
2044-05-28

AI Technical Summary

Technical Problem

Traditional audio-visual data collection devices are limited by the need for fixed network connections, leading to reduced mobility and efficiency due to increased operational complexity and potential network bottlenecks, especially with outdated hardware and bandwidth constraints.

Method used

A method and system utilizing 5G terminal devices to control audio-visual data collection through API interfaces, incorporating quality scoring models for data optimization and encoding, enabling efficient data transmission over wireless connections.

Benefits of technology

Enhances audio-visual data collection efficiency by allowing flexible control, real-time processing, and optimized data transmission, overcoming mobility limitations and network constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118573959B_ABST
    Figure CN118573959B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio - video data acquisition method and system based on a 5G terminal device. Based on the 5G terminal device, the API interface of the audio - video acquisition device is called to send an audio - video acquisition instruction to the audio - video acquisition device, so that the audio - video acquisition device performs audio - video acquisition processing on the target object to obtain first audio - video data; the first audio - video data is input into a quality scoring model for quality scoring processing, and an audio - video quality detection score is output; when the audio - video quality detection score does not meet the preset audio - video quality detection score threshold, the first audio - video data is subjected to audio - video optimization processing to obtain first audio - video optimized data; the first audio - video optimized data is subjected to encoding processing to obtain first audio - video encoded data, and the first audio - video encoded data is transmitted to the target receiving device based on a preset network protocol; compared with the prior art, the technical solution of the present invention can improve the efficiency of audio - video data acquisition and transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio - video data processing, and particularly to an audio - video data acquisition method and system based on 5G terminal devices. Background Art

[0002] Traditional audio - video acquisition devices usually use Ethernet connections for networking, but this connection method has some limitations. First, Ethernet connections usually need to be connected to a fixed network access point and connected to the network in a wired manner, which limits the mobility of the devices. If audio - video acquisition is required at different locations, the connection needs to be frequently changed, increasing the operation complexity and time cost, resulting in low acquisition efficiency.

[0003] In addition, traditional audio - video acquisition devices usually adopt old network technologies or hardware facilities, which may be limited by network bandwidth and device performance. Insufficient network bandwidth or network congestion may lead to a decrease in data transmission rate, affecting the efficiency and quality of audio - video acquisition. At the same time, insufficient hardware device performance will also affect data processing and transmission speed, further reducing the acquisition efficiency. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an audio - video data acquisition method and system based on 5G terminal devices, which can improve the efficiency of audio - video data acquisition and transmission.

[0005] To solve the above - mentioned technical problem, the present invention provides an audio - video data acquisition method based on 5G terminal devices, including:

[0006] Invoking the API interface of the audio - video acquisition device based on the 5G terminal device, and sending an audio - video acquisition instruction to the audio - video acquisition device, so that after receiving the audio - video acquisition instruction, the audio - video acquisition device performs audio - video acquisition processing on the target object to obtain first audio - video data;

[0007] Inputting the first audio - video data into a pre - trained quality scoring model, so that the quality scoring model performs quality scoring processing on the first audio - video data and outputs an audio - video quality detection score;

[0008] When the audio - video quality detection score does not meet the preset audio - video quality detection score threshold, performing audio - video optimization processing on the first audio - video data to obtain first audio - video optimized data;

[0009] Performing encoding processing on the first audio - video optimized data to obtain first audio - video encoded data, and transmitting the first audio - video encoded data to the target receiving device based on a preset network protocol.

[0010] In a possible implementation, audio and video data of a target object is collected and processed to obtain first audio and video data, specifically including:

[0011] The audio and video collection device includes an audio collection device and a video collection device;

[0012] The audio and video collection instruction is parsed and processed to determine audio and video collection parameters, where the audio and video collection parameters include resolution, frame rate, and sampling frequency;

[0013] Based on the audio and video collection parameters, the device parameters of the audio collection device and the video collection device are set respectively to obtain audio collection device parameters and video collection device parameters;

[0014] Based on the audio collection device parameters, audio collection processing is performed on the target object to obtain first audio data, and based on the video collection device parameters, video collection processing is performed on the target object to obtain first video data;

[0015] The first audio data and the first video data are integrated to obtain first audio and video data.

[0016] In a possible implementation, the first audio and video data is input into a pre-trained quality scoring model so that the quality scoring model performs quality scoring processing on the first audio and video data and outputs an audio and video quality detection score, specifically including:

[0017] The first audio and video data is input into a pre-trained quality scoring model so that the quality scoring model performs adaptive filtering processing on the first audio and video data to obtain audio and video adaptive filtering data;

[0018] Signal energy calculation is performed on the first audio and video data to obtain a first signal energy, and signal energy calculation is performed on the audio and video adaptive filtering data to obtain a second signal energy. Based on the first signal energy and the second signal energy, a signal energy difference is determined;

[0019] Based on the signal energy difference, an audio and video quality detection score is determined.

[0020] In a possible implementation, encoding processing is performed on the first audio and video optimization data to obtain first audio and video encoded data, specifically including:

[0021] The first audio and video optimization data is subjected to data decomposition processing to obtain a plurality of audio and video decomposition data;

[0022] Feature extraction is performed on the plurality of audio and video decomposition data respectively to obtain a first feature corresponding to each audio and video decomposition data;

[0023] Based on the first feature, encode the multiple audio-visual decomposition data respectively to obtain multiple first audio-visual decomposition encoded data, and integrate the multiple first audio-visual decomposition encoded data to obtain first audio-visual encoded data.

[0024] In a possible implementation, perform data decomposition processing on the first audio-visual optimization data to obtain multiple audio-visual decomposition data; extract features from the multiple audio-visual decomposition data respectively to obtain the first feature corresponding to each audio-visual decomposition data, specifically including:

[0025] The first audio-visual optimization data includes first audio optimization data and first video optimization data;

[0026] Perform data decomposition processing on the first audio optimization data to obtain multiple audio decomposition data, and perform data decomposition processing on the first video optimization data to obtain multiple video decomposition data;

[0027] Based on the Mel filter bank, perform a convolution operation on each audio decomposition data to obtain the Mel filter bank output value corresponding to each audio decomposition data, and perform a logarithmic transformation on the Mel filter bank output value to obtain logarithmic Mel spectral coefficients, and obtain the first audio feature corresponding to each audio decomposition data based on the logarithmic Mel spectral coefficients;

[0028] Based on the optical flow algorithm, calculate the optical flow vector map corresponding to each video decomposition data, calculate the vector mean of the optical flow vector map to obtain the optical flow vector mean, and obtain the first video feature corresponding to each video decomposition data based on the optical flow vector mean.

[0029] In a possible implementation, based on the first feature, encode the multiple audio-visual decomposition data respectively to obtain multiple first audio-visual decomposition encoded data, specifically including:

[0030] Set audio bit position tags for each audio decomposition data respectively, and encode each audio decomposition data based on a preset audio encoding algorithm to obtain the first audio encoded data corresponding to each audio decomposition data;

[0031] Splice the audio bit position tag, the first audio feature, and the first audio encoded data corresponding to each audio decomposition data respectively to obtain the audio encoded data corresponding to each audio decomposition data;

[0032] Set video bit position tags for each video decomposition data respectively, and encode each video decomposition data based on a preset video encoding algorithm to obtain the first video encoded data corresponding to each video decomposition data;

[0033] For each video decomposition data, splice the corresponding video bit tags, the first video features, and the first video encoding data to obtain the video encoding data corresponding to each video decomposition data.

[0034] In a possible implementation, perform audio-video optimization processing on the first audio-video data to obtain first audio-video optimized data, specifically including:

[0035] Obtain the first signal energy corresponding to the first audio-video data, and based on the first signal energy, calculate the gain step of the filter;

[0036] And obtain the current gain of the filter, and adjust the current gain based on the gain step to obtain the adjusted gain of the filter;

[0037] Control the filter to perform filtering processing on the first audio-video data based on the adjusted gain to obtain first audio-video optimized data.

[0038] The present invention also provides an audio-video data acquisition system based on a 5G terminal device, including: an audio-video acquisition module, a quality scoring module, an audio-video optimization module, and an audio-video transmission module;

[0039] Among them, the audio-video acquisition module is used to call the api interface of the audio-video acquisition device based on the 5G terminal device, send an audio-video acquisition instruction to the audio-video acquisition device, so that after receiving the audio-video acquisition instruction, the audio-video acquisition device performs audio-video acquisition processing on the target object to obtain first audio-video data;

[0040] The quality scoring module is used to input the first audio-video data into a pre-trained quality scoring model, so that the quality scoring model performs quality scoring processing on the first audio-video data and outputs an audio-video quality detection score;

[0041] The audio-video optimization module is used to perform audio-video optimization processing on the first audio-video data to obtain first audio-video optimized data when the audio-video quality detection score does not meet a preset audio-video quality detection score threshold;

[0042] The audio-video transmission module is used to perform encoding processing on the first audio-video optimized data to obtain first audio-video encoding data, and transmit the first audio-video encoding data to a target receiving device based on a preset network protocol.

[0043] In a possible implementation, the audio-video acquisition module is used to perform audio-video acquisition processing on the target object to obtain first audio-video data, specifically including:

[0044] The audio - video acquisition device includes an audio acquisition device and a video acquisition device;

[0045] Parse and process the audio - video acquisition instruction to determine audio - video acquisition parameters, where the audio - video acquisition parameters include resolution, frame rate, and sampling frequency;

[0046] Based on the audio - video acquisition parameters, set the device parameters of the audio acquisition device and the video acquisition device respectively to obtain audio acquisition device parameters and video acquisition device parameters;

[0047] Based on the audio acquisition device parameters, perform audio acquisition processing on the target object to obtain first - audio data, and based on the video acquisition device parameters, perform video acquisition processing on the target object to obtain first - video data;

[0048] Integrate the first - audio data and the first - video data to obtain first - audio - video data.

[0049] In a possible implementation manner, the quality scoring module is used to input the first - audio - video data into a pre - trained quality scoring model, so that the quality scoring model performs quality scoring processing on the first - audio - video data and outputs an audio - video quality detection score, specifically including:

[0050] Input the first - audio - video data into a pre - trained quality scoring model, so that the quality scoring model performs adaptive filtering processing on the first - audio - video data to obtain audio - video adaptive filtering data;

[0051] Calculate the signal energy of the first - audio - video data to obtain a first signal energy, calculate the signal energy of the audio - video adaptive filtering data to obtain a second signal energy, and determine a signal energy difference based on the first signal energy and the second signal energy;

[0052] Determine the audio - video quality detection score based on the signal energy difference.

[0053] In a possible implementation manner, the audio - video transmission module is used to perform encoding processing on the first - audio - video optimized data to obtain first - audio - video encoded data, specifically including:

[0054] Perform data decomposition processing on the first - audio - video optimized data to obtain multiple audio - video decomposition data;

[0055] Extract features from each of the multiple audio - video decomposition data to obtain a first feature corresponding to each audio - video decomposition data;

[0056] Based on the first feature, encode the multiple audio-visual decomposition data respectively to obtain multiple first audio-visual decomposition encoded data, and integrate the multiple first audio-visual decomposition encoded data to obtain first audio-visual encoded data.

[0057] In a possible implementation, the audio-visual transmission module is configured to perform data decomposition processing on the first audio-visual optimization data to obtain multiple audio-visual decomposition data; perform feature extraction on the multiple audio-visual decomposition data respectively to obtain a first feature corresponding to each audio-visual decomposition data, specifically including:

[0058] The first audio-visual optimization data includes first audio optimization data and first video optimization data;

[0059] Perform data decomposition processing on the first audio optimization data to obtain multiple audio decomposition data, and perform data decomposition processing on the first video optimization data to obtain multiple video decomposition data;

[0060] Based on the Mel filter bank, perform a convolution operation on each audio decomposition data to obtain the Mel filter bank output value corresponding to each audio decomposition data, and perform a logarithmic transformation on the Mel filter bank output value to obtain logarithmic Mel spectrogram coefficients, and obtain a first audio feature corresponding to each audio decomposition data based on the logarithmic Mel spectrogram coefficients;

[0061] Based on the optical flow algorithm, calculate the optical flow vector map corresponding to each video decomposition data, calculate the vector mean of the optical flow vector map to obtain the optical flow vector mean, and obtain a first video feature corresponding to each video decomposition data based on the optical flow vector mean.

[0062] In a possible implementation, the audio-visual transmission module is configured to, based on the first feature, encode the multiple audio-visual decomposition data respectively to obtain multiple first audio-visual decomposition encoded data, specifically including:

[0063] Set an audio bit position label for each audio decomposition data respectively, and encode each audio decomposition data based on a preset audio encoding algorithm to obtain a first audio encoded data corresponding to each audio decomposition data;

[0064] Splice the audio bit position label, the first audio feature, and the first audio encoded data corresponding to each audio decomposition data respectively to obtain an audio encoded data corresponding to each audio decomposition data;

[0065] Set a video bit position label for each video decomposition data respectively, and encode each video decomposition data based on a preset video encoding algorithm to obtain a first video encoded data corresponding to each video decomposition data;

[0066] For each video decomposition data, splice the corresponding video bit tags, the first video feature, and the first video encoding data to obtain the video encoding data corresponding to each video decomposition data.

[0067] In a possible implementation manner, the audio-video optimization module is configured to perform audio-video optimization processing on the first audio-video data to obtain first audio-video optimized data, which specifically includes:

[0068] Obtain the first signal energy corresponding to the first audio-video data, and calculate the gain step of the filter based on the first signal energy;

[0069] And obtain the current gain of the filter, and adjust the current gain based on the gain step to obtain the adjusted gain of the filter;

[0070] Control the filter to perform filtering processing on the first audio-video data based on the adjusted gain to obtain first audio-video optimized data.

[0071] The present invention also provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the audio-video data acquisition method based on a 5G terminal device as described in any one of the above.

[0072] The present invention also provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the audio-video data acquisition method based on a 5G terminal device as described in any one of the above.

[0073] An audio-video data acquisition method and system based on a 5G terminal device according to an embodiment of the present invention have the following beneficial effects compared with the prior art:

[0074] Based on the API interface of the 5G terminal device to call the audio and video collection device, audio and video collection instructions can be remotely sent to the audio and video collection device, so that after receiving the audio and video collection instructions, the audio and video collection device performs audio and video collection processing on the target object to obtain the first audio and video data, realizing the intelligent control and scheduling of the audio and video collection device, improving the efficiency of audio and video data collection, and at the same time eliminating the limitation of relying on wired connection in the traditional method, enhancing the mobility of the device; by performing quality scoring processing on the first audio and video data through a pre-trained quality scoring model, quality problems in the audio and video data can be discovered and determined in a timely manner; when the quality score does not meet the preset threshold, the audio and video data is optimized to improve the quality of the audio and video transmitted to the target device; the first optimized audio and video data is encoded to obtain the first audio and video encoded data, which can further reduce the size of data transmission and improve the transmission efficiency; based on the preset network protocol for transmission, it can ensure the stable transmission of data and its efficient delivery to the target receiving device. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 is a schematic flowchart of an embodiment of a method for collecting audio and video data based on a 5G terminal device provided by the present invention;

[0076] Figure 2 is a schematic structural diagram of an embodiment of a system for collecting audio and video data based on a 5G terminal device provided by the present invention;

[0077] Figure 3 is a schematic structural diagram of a terminal device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] Embodiment 1, see Figure 1 , Figure 1 is a schematic flowchart of an embodiment of a method for collecting audio and video data based on a 5G terminal device provided by the present invention. As Figure 1 shown, the method includes steps 101-step 104, which are specifically as follows:

[0080] Step 101: Based on the 5G terminal device, call the API interface of the audio-video collection device, and send an audio-video collection instruction to the audio-video collection device, so that after receiving the audio-video collection instruction, the audio-video collection device performs audio-video collection processing on the target object to obtain the first audio-video data.

[0081] In one embodiment, obtain the interface document of the audio-video collection device, and based on the interface document, determine the API interface for controlling the audio-video collection device to collect audio-video data.

[0082] Preferably, through the API interface, the status control of the camera and the microphone can be realized, such as starting or stopping camera shooting or audio recording, etc.

[0083] In one embodiment, based on wireless technologies such as network connection or Bluetooth, establish a connection relationship between the 5G terminal device and the audio-video collection device. When it is detected that there is a connection relationship between the 5G terminal device and the audio-video collection device, call the API interface of the audio-video collection device based on the 5G terminal device, so that when the API interface is called, an audio-video collection instruction is sent to the audio-video collection device.

[0084] In one embodiment, the audio-video collection device includes an audio collection device and a video collection device.

[0085] In one embodiment, when performing audio-video collection processing on the target object to obtain the first audio-video data, by parsing and processing the audio-video collection instruction, determine the audio-video collection parameters, where the audio-video collection parameters include resolution, frame rate, and sampling frequency; based on the audio-video collection parameters, set the device parameters of the audio collection device and the video collection device respectively to obtain the audio collection device parameters and the video collection device parameters; based on the audio collection device parameters, perform audio collection processing on the target object to obtain the first audio data, and based on the video collection device parameters, perform video collection processing on the target object to obtain the first video data; integrate the first audio data and the first video data to obtain the first audio-video data.

[0086] In one embodiment, sending an audio-video collection instruction through the API interface can provide more flexible and precise control, and can adjust the collection parameters according to specific requirements, such as resolution, frame rate, sampling frequency, etc., to achieve better audio-video collection effects; at the same time, the high-speed transmission characteristics of 5G technology can ensure the real-time performance and transmission efficiency of the audio-video collection instruction, enabling the audio-video collection device to quickly respond and execute the instruction, realizing fast and real-time audio-video collection processing, and meeting the user's requirements for real-time performance.

[0087] Step 102: Input the first audio-video data into a pre-trained quality scoring model, so that the quality scoring model performs quality scoring processing on the first audio-video data and outputs an audio-video quality detection score.

[0088] In one embodiment, a historical audio-video data set is obtained, and adaptive filtering processing is performed on each historical audio-video data in the historical audio-video data set to obtain a historical audio-video adaptive filtering data set.

[0089] In one embodiment, the historical audio-video data set is used as a model training sample set, and a quality label is set for each model training sample in the model training sample set.

[0090] Specifically, the quality label is a quality detection score, where the quality detection score is obtained by performing signal energy on each historical audio-video data in the historical audio-video data set and each historical audio-video adaptive filtering data in the historical audio-video adaptive filtering data set.

[0091] In one embodiment, when training a quality scoring model constructed based on deep learning technology, the model training sample is used as the model input of the quality scoring model, and the quality label of the model training sample is used as the model output, and the quality scoring model is iteratively trained until the model fits or reaches a preset number of iterations to obtain a trained quality scoring model.

[0092] Specifically, in the iterative training process, the model training sample is input into the quality scoring model for forward propagation, the loss function is calculated, and the model parameters of the quality scoring model are updated by backpropagation based on the calculated loss function value until a quality scoring model with optimal model parameters is determined.

[0093] In one embodiment, the first audio-video data is input into a pre-trained quality scoring model, so that the quality scoring model performs adaptive filtering processing on the first audio-video data to obtain audio-video adaptive filtering data.

[0094] Specifically, an adaptive filter is set in the quality scoring model, and by obtaining the adaptive filter coefficients in the adaptive filter, the first audio-video data is subjected to adaptive filtering processing based on the adaptive filter coefficients.

[0095] Specifically, when performing adaptive filtering processing on the first audio-video data based on the adaptive filter coefficients, the adaptive filter coefficients and the first audio-video data are substituted into a preset adaptive filtering formula to obtain audio-video adaptive filtering data.

[0096] Specifically, the adaptive filtering formula is as follows:

[0097] ;

[0098] In the formula, is the audio - video adaptive filtering data, is the adaptive filter coefficient, is the first audio - video data.

[0099] In one embodiment, signal energy calculation is performed on the first audio - video data to obtain a first signal energy, and signal energy calculation is performed on the audio - video adaptive filtering data to obtain a second signal energy. Based on the first signal energy and the second signal energy, a signal energy difference is determined.

[0100] Specifically, the first audio data in the first audio - video data is converted into a frequency - domain audio signal, and sampling processing is performed on the frequency - domain audio signal to obtain a first discretized audio signal corresponding to each sampling point; the first audio signal amplitude corresponding to each sampling point of the first discretized audio signal is obtained, and the first audio signal amplitude is squared to obtain a first audio signal squared value corresponding to each first discretized audio signal. All the first audio signal squared values are integrated to obtain the first audio signal energy.

[0101] Specifically, the first video data in the first audio - video data is converted into a plurality of first image frames, the average value of the first pixel values corresponding to each first image frame is obtained, and the first pixel average value is squared to obtain a first pixel value squared value corresponding to each first image frame. All the first pixel value squared values are integrated to obtain the first video signal energy.

[0102] Specifically, the signal energy sum of the first audio signal energy and the first video signal energy is calculated to obtain the first signal energy of the first audio - video data.

[0103] Specifically, the first adaptive filtering audio data in the audio - video adaptive filtering data is converted into a second frequency - domain audio signal, and sampling processing is performed on the second frequency - domain audio signal to obtain a second discretized audio signal corresponding to each sampling point; the second audio signal amplitude corresponding to each sampling point of the second discretized audio signal is obtained, and the second audio signal amplitude is squared to obtain a second audio signal squared value corresponding to each second discretized audio signal. All the second audio signal squared values are integrated to obtain the first adaptive filtering audio signal energy.

[0104] Specifically, convert the first adaptive filtering video data in the audio-visual adaptive filtering data into multiple second image frames, obtain the average value of the second pixel values corresponding to each second image frame, and square the average value of the second pixels to obtain the squared value of the second pixel values corresponding to each second image frame. Integrate all the squared values of the second pixel values to obtain the energy of the first adaptive filtering video signal.

[0105] Specifically, calculate the sum of the signal energies of the first adaptive filtering audio signal energy and the first adaptive filtering video signal energy to obtain the second signal energy of the audio-visual adaptive filtering data.

[0106] Specifically, calculate the difference between the first signal energy and the second signal energy to determine the signal energy difference.

[0107] In one embodiment, when determining the audio-visual quality detection score based on the signal energy difference, directly use the signal energy difference as the audio-visual quality detection score corresponding to the first audio-visual data.

[0108] Step 103: When the audio-visual quality detection score does not meet the preset audio-visual quality detection score threshold, perform audio-visual optimization processing on the first audio-visual data to obtain the first audio-visual optimized data.

[0109] In one embodiment, by obtaining the first signal energy corresponding to the first audio-visual data, calculate the gain step of the filter based on the first signal energy.

[0110] Specifically, the calculation method of the first signal energy corresponding to the first audio-visual data is as shown in the above step 102, and will not be further elaborated here.

[0111] Specifically, input the first signal energy into a preset filter gain step calculation formula to obtain the gain step of the filter. The filter gain step calculation formula is as follows:

[0112] ;

[0113] In the formula, is the gain step, T and u are hyperparameters, where u is the learning rate, usually set to 0.1, and T is the hyperparameter threshold, usually set to 0.5. is the gain step, is the first signal energy.

[0114] In one embodiment, obtain the current gain of the filter, and adjust the current gain based on the gain step to obtain the adjusted gain of the filter.

[0115] Specifically, calculate the sum of the gain step and the current gain to obtain the adjusted gain of the filter, and use the adjusted gain as the current gain of the filter.

[0116] In one embodiment, control the filter to perform filtering processing on the first audio-visual data based on the adjusted gain to obtain first optimized audio-visual data.

[0117] Specifically, substitute the adjusted gain and the first audio-visual data into a preset filtering formula to obtain first optimized audio-visual data, where the preset filtering formula is as follows:

[0118] ;

[0119] In the formula, is the first optimized audio-visual data, is the adjusted gain, is the first audio-visual data.

[0120] Step 104: Perform encoding processing on the first optimized audio-visual data to obtain first encoded audio-visual data, and transmit the first encoded audio-visual data to a target receiving device based on a preset network protocol.

[0121] In one embodiment, perform data decomposition processing on the first optimized audio-visual data to obtain multiple decomposed audio-visual data.

[0122] Specifically, the first optimized audio-visual data includes first optimized audio data and first optimized video data.

[0123] Specifically, perform data decomposition processing on the first optimized audio data to obtain multiple decomposed audio data; where the data decomposition processing includes performing sliding processing on the first optimized audio data based on a preset time window to obtain multiple audio window data, and performing short-time Fourier transform processing on the multiple audio window data to obtain multiple decomposed audio data.

[0124] Specifically, perform data decomposition processing on the first optimized video data to obtain multiple decomposed video data; where the data decomposition processing includes decomposing the first optimized video data into multiple video image frames to obtain multiple decomposed video data.

[0125] In one embodiment, perform feature extraction on the multiple decomposed audio-visual data respectively to obtain first features corresponding to each decomposed audio-visual data.

[0126] In one embodiment, for the audio decomposition data, a convolution operation is respectively performed on each audio decomposition data based on the Mel filter bank to obtain the Mel filter bank output values corresponding to each audio decomposition data, and a logarithmic transformation process is performed on the Mel filter bank output values to obtain logarithmic Mel spectrogram coefficients, and a first audio feature corresponding to each audio decomposition data is obtained based on the logarithmic Mel spectrogram coefficients.

[0127] Specifically, when performing the logarithmic transformation process on the Mel filter bank output values, a logarithmic transformation is performed on the Mel filter bank output values based on the logarithmic function to obtain logarithmic values, and the logarithmic values are used as the logarithmic Mel spectrogram coefficients.

[0128] Specifically, when obtaining a first audio feature corresponding to each audio decomposition data based on the logarithmic Mel spectrogram coefficients, a discrete cosine transform is performed on the logarithmic Mel spectrogram coefficients to obtain Mel frequency cepstrum coefficients, and the first target number of coefficient values in the Mel frequency cepstrum coefficients are obtained to obtain the first audio feature.

[0129] In one embodiment, for the video decomposition data, based on the optical flow algorithm, an optical flow vector map corresponding to each video decomposition data is calculated, a vector mean value calculation is performed on the optical flow vector map to obtain an optical flow vector mean value, and a first video feature corresponding to each video decomposition data is obtained based on the optical flow vector mean value.

[0130] Specifically, based on the optical flow algorithm, an optical flow estimation is performed on each pixel point in each video decomposition data to obtain an optical flow vector corresponding to each pixel point, and the optical flow vector is converted into a visual optical flow vector map.

[0131] Specifically, when performing an optical flow estimation on each pixel point in each video decomposition data based on the optical flow algorithm, the first video decomposition data adjacent to each other in pairs among the multiple video decomposition data are associated to obtain multiple pairs of first video decomposition data, where each pair of first video decomposition data is divided into the previous frame of the first video decomposition data and the current frame of the first video decomposition data based on the acquisition time sequence, a pixel point matching process is performed on the current frame of the first video decomposition data and the previous frame of the first video decomposition data to obtain a first pixel point pair corresponding to each pixel point in the current frame of the first video decomposition data, based on the first pixel point pair, a position offset amount of each pixel point between the current frame of the first video decomposition data and the previous frame of the first video decomposition data is calculated, and the position offset amount is used as the optical flow vector corresponding to each pixel point.

[0132] Preferably, the optical flow algorithm includes but is not limited to the Lucas-Kanade optical flow algorithm, the Horn-Schunck optical flow algorithm, and the Farneback optical flow algorithm.

[0133] Specifically, when calculating the vector mean of the optical flow vector map to obtain the optical flow vector mean, integrate the optical flow vectors corresponding to each pixel point in the optical flow vector map to obtain the optical flow vector sum, and perform mean processing on the optical flow vector sum with the number of pixel points to obtain the optical flow vector mean.

[0134] In one embodiment, based on the first feature, perform encoding processing on the multiple audio-visual decomposition data respectively to obtain multiple first audio-visual decomposition encoded data, and integrate the multiple first audio-visual decomposition encoded data to obtain the first audio-visual encoded data.

[0135] Specifically, set an audio bit position tag for each audio decomposition data respectively, and perform encoding processing on each audio decomposition data based on a preset audio encoding algorithm to obtain the first audio encoding data corresponding to each audio decomposition data.

[0136] Preferably, set the audio bit position tag to 0.

[0137] Preferably, the preset audio encoding algorithms include but are not limited to pulse code modulation algorithm, adaptive differential pulse code modulation algorithm, MP3 algorithm, AAC algorithm, FLAC algorithm, and WAV algorithm.

[0138] Specifically, splice the audio bit position tag, the first audio feature, and the first audio encoding data corresponding to each audio decomposition data respectively to obtain the audio encoding data corresponding to each audio decomposition data.

[0139] Specifically, when splicing the audio bit position tag, the first audio feature, and the first audio encoding data corresponding to each audio decomposition data, use the audio bit position tag as the first sentence start encoding, the first audio encoding data as the first sentence middle encoding, and the first audio feature as the first sentence end encoding, and splice the audio bit position tag, the first audio encoding data, and the first audio feature in the order of sentence start encoding - sentence middle encoding - sentence end encoding to obtain the audio encoding data corresponding to each audio decomposition data.

[0140] Specifically, set a video bit position tag for each video decomposition data respectively, and perform encoding processing on each video decomposition data based on a preset video encoding algorithm to obtain the first video encoding data corresponding to each video decomposition data.

[0141] Preferably, set the video bit position tag to 1.

[0142] Preferably, the preset video encoding algorithms include but are not limited to H.264 / AVC algorithm, H.265 / HEV algorithm, VP9 algorithm, AV1 algorithm, and MPEG-2 algorithm.

[0143] Specifically, the video bit - tag, the first video feature, and the first video encoding data corresponding to each video decomposition data are respectively concatenated to obtain the video encoding data corresponding to each video decomposition data.

[0144] Specifically, when concatenating the video bit - tag, the first video feature, and the first video encoding data corresponding to each video decomposition data, the video bit - tag is used as the second sentence - start encoding, the first video encoding data is used as the second sentence - middle encoding, and the first video feature is used as the second sentence - end encoding. In the order of sentence - start encoding - sentence - middle encoding - sentence - end encoding, the video bit - tag, the first video encoding data, and the first video feature are concatenated to obtain the video encoding data corresponding to each video decomposition data.

[0145] In one embodiment, when transmitting the first audio - video encoding data to a target receiving device based on a preset network protocol, the first audio - video encoding data is encapsulated to obtain first audio - video encapsulated data, and in a 5G network environment, a connection relationship is established between the 5G terminal device and the target receiving device, and based on the preset network protocol, the first audio - video encapsulated data is transmitted to the target receiving device.

[0146] Specifically, the preset network protocol includes but is not limited to the HyperText Transfer Protocol, the Real - Time Streaming Protocol, the Web Real - Time Communication Protocol, the Secure Reliable Transport Protocol, and the QUIC (Quick UDP Internet Connections) protocol.

[0147] Embodiment 2, see Figure 2 , Figure 2 is a schematic structural diagram of an embodiment of an audio - video data acquisition system based on a 5G terminal device provided by the present invention. As Figure 2 shown, the system includes an audio - video acquisition module 201, a quality scoring module 202, an audio - video optimization module 203, and an audio - video transmission module 204, specifically as follows:

[0148] The audio - video acquisition module 201 is used to call the api interface of the audio - video acquisition device based on the 5G terminal device, send an audio - video acquisition instruction to the audio - video acquisition device, so that after receiving the audio - video acquisition instruction, the audio - video acquisition device performs audio - video acquisition processing on the target object to obtain first audio - video data.

[0149] The quality scoring module 202 is used to input the first audio - video data into a pre - trained quality scoring model, so that the quality scoring model performs quality scoring processing on the first audio - video data and outputs an audio - video quality detection score.

[0150] The audio - video optimization module 203 is configured to perform audio - video optimization processing on the first audio - video data to obtain first audio - video optimized data when the audio - video quality detection score does not meet the preset audio - video quality detection score threshold.

[0151] The audio - video transmission module 204 is configured to perform encoding processing on the first audio - video optimized data to obtain first audio - video encoded data, and transmit the first audio - video encoded data to a target receiving device based on a preset network protocol.

[0152] In one embodiment, the audio - video acquisition module 201 is configured to perform audio - video acquisition processing on a target object to obtain first audio - video data, specifically including: the audio - video acquisition device includes an audio acquisition device and a video acquisition device; parsing the audio - video acquisition instruction to determine audio - video acquisition parameters, where the audio - video acquisition parameters include resolution, frame rate, and sampling frequency; setting device parameters for the audio acquisition device and the video acquisition device respectively based on the audio - video acquisition parameters to obtain audio acquisition device parameters and video acquisition device parameters; performing audio acquisition processing on the target object based on the audio acquisition device parameters to obtain first audio data, and performing video acquisition processing on the target object based on the video acquisition device parameters to obtain first video data; integrating the first audio data and the first video data to obtain first audio - video data.

[0153] In one embodiment, the quality scoring module 202 is configured to input the first audio - video data into a pre - trained quality scoring model, so that the quality scoring model performs quality scoring processing on the first audio - video data and outputs an audio - video quality detection score, specifically including: inputting the first audio - video data into the pre - trained quality scoring model, so that the quality scoring model performs adaptive filtering processing on the first audio - video data to obtain audio - video adaptive filtering data; calculating the signal energy of the first audio - video data to obtain a first signal energy, and calculating the signal energy of the audio - video adaptive filtering data to obtain a second signal energy, and determining a signal energy difference based on the first signal energy and the second signal energy; determining the audio - video quality detection score based on the signal energy difference.

[0154] In one embodiment, the audio-video transmission module 204 is configured to perform encoding processing on the first audio-video optimized data to obtain first audio-video encoded data, specifically including: performing data decomposition processing on the first audio-video optimized data to obtain a plurality of audio-video decomposition data; respectively performing feature extraction on the plurality of audio-video decomposition data to obtain first features corresponding to each audio-video decomposition data; based on the first features, respectively performing encoding processing on the plurality of audio-video decomposition data to obtain a plurality of first audio-video decomposition encoded data, and integrating the plurality of first audio-video decomposition encoded data to obtain first audio-video encoded data.

[0155] In one embodiment, the audio-video transmission module 204 is configured to perform data decomposition processing on the first audio-video optimized data to obtain a plurality of audio-video decomposition data; respectively perform feature extraction on the plurality of audio-video decomposition data to obtain first features corresponding to each audio-video decomposition data, specifically including: the first audio-video optimized data includes first audio optimized data and first video optimized data; performing data decomposition processing on the first audio optimized data to obtain a plurality of audio decomposition data, and performing data decomposition processing on the first video optimized data to obtain a plurality of video decomposition data; respectively performing convolution operations on each audio decomposition data based on a Mel filter bank to obtain Mel filter bank output values corresponding to each audio decomposition data, performing logarithmic transformation processing on the Mel filter bank output values to obtain log Mel spectral coefficients, and obtaining first audio features corresponding to each audio decomposition data based on the log Mel spectral coefficients; based on an optical flow algorithm, calculating an optical flow vector map corresponding to each video decomposition data, performing vector mean calculation on the optical flow vector map to obtain an optical flow vector mean, and obtaining first video features corresponding to each video decomposition data based on the optical flow vector mean.

[0156] In one embodiment, the audio - video transmission module 204 is configured to perform encoding processing on the multiple audio - video decomposition data respectively based on the first feature, so as to obtain multiple first audio - video decomposition encoded data, specifically including: setting an audio bit - position label for each audio decomposition data respectively, performing encoding processing on each audio decomposition data respectively based on a preset audio encoding algorithm to obtain first audio encoded data corresponding to each audio decomposition data; splicing the audio bit - position label, the first audio feature, and the first audio encoded data corresponding to each audio decomposition data respectively to obtain audio encoded data corresponding to each audio decomposition data; setting a video bit - position label for each video decomposition data respectively, performing encoding processing on each video decomposition data respectively based on a preset video encoding algorithm to obtain first video encoded data corresponding to each video decomposition data; splicing the video bit - position label, the first video feature, and the first video encoded data corresponding to each video decomposition data respectively to obtain video encoded data corresponding to each video decomposition data.

[0157] In one embodiment, the audio - video optimization module 203 is configured to perform audio - video optimization processing on the first audio - video data to obtain first audio - video optimized data, specifically including: obtaining the first signal energy corresponding to the first audio - video data, calculating the gain step of the filter based on the first signal energy; and obtaining the current gain of the filter, adjusting the current gain based on the gain step to obtain the adjusted gain of the filter; controlling the filter to perform filtering processing on the first audio - video data based on the adjusted gain to obtain first audio - video optimized data.

[0158] The above - mentioned audio - video data acquisition device based on a 5G terminal device can implement the audio - video data acquisition method based on a 5G terminal device in the above - mentioned method embodiment. The optional items in the above - mentioned method embodiment are also applicable to this embodiment and will not be elaborated here.

[0159] Figure 3 A schematic structural diagram of a terminal device. As Figure 3 shown, the terminal device 3 in this embodiment includes: at least one processor 301 ( Figure 3 only one is shown in the figure), a memory 302, and a computer program 303 stored in the memory 302 and executable on at least one processor 301. When the processor 301 executes the computer program 303, the steps in any of the above - mentioned method embodiments are implemented.

[0160] The terminal device 3 can be a computing device such as a smart phone, a laptop computer, a tablet computer, and a desktop computer. The terminal device may include but is not limited to the processor 301 and the memory 302. Those skilled in the art can understand, Figure 3The terminal device 3 is only an example and does not limit the terminal device 3. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0161] The so-called processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0162] In some embodiments, the memory 302 may be an internal storage unit of the terminal device 3, such as the hard disk or memory of the terminal device 3. In other embodiments, the memory 302 may also be an external storage device of the terminal device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 3. Further, the memory 302 may also include both the internal storage unit and the external storage device of the terminal device 3. The memory 302 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory 302 may also be used to temporarily store data that has been output or is to be output.

[0163] In addition, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0164] In several embodiments provided in the present application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0165] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0166] In summary, a method and system for audio and video data collection based on a 5G terminal device provided by the present invention call the API interface of the audio and video collection device based on the 5G terminal device, send an audio and video collection instruction to the audio and video collection device, so that the audio and video collection device performs audio and video collection processing on the target object to obtain first audio and video data; input the first audio and video data into a quality scoring model for quality scoring processing, and output an audio and video quality detection score; when the audio and video quality detection score does not meet the preset audio and video quality detection score threshold, perform audio and video optimization processing on the first audio and video data to obtain first audio and video optimized data; perform encoding processing on the first audio and video optimized data to obtain first audio and video encoded data, and transmit the first audio and video encoded data to the target receiving device based on the preset network protocol; compared with the prior art, the technical solution of the present invention can improve the efficiency of audio and video data collection and transmission.

[0167] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present invention.

Claims

1. A method for collecting audio and video data based on 5G terminal equipment, characterized in that: include: Based on the 5G terminal device, the API interface of the audio and video acquisition device is called, and an audio and video acquisition instruction is sent to the audio and video acquisition device, so that the audio and video acquisition device performs audio and video acquisition processing on the target object after receiving the audio and video acquisition instruction to obtain the first audio and video data; Inputting the first audio and video data into a pre-trained quality scoring model so that the quality scoring model performs quality scoring processing on the first audio and video data and outputs an audio and video quality detection score; When the audio and video quality detection score does not meet the preset audio and video quality detection score threshold, performing audio and video optimization processing on the first audio and video data to obtain first audio and video optimization data; Encoding the first audio and video optimization data to obtain first audio and video encoded data, and transmitting the first audio and video encoded data to a target receiving device based on a preset network protocol; Inputting the first audio and video data into a pre-trained quality scoring model so that the quality scoring model performs quality scoring processing on the first audio and video data and outputs an audio and video quality detection score, specifically including: Inputting the first audio and video data into a pre-trained quality scoring model so that the quality scoring model performs adaptive filtering processing on the first audio and video data to obtain audio and video adaptive filtering data; Performing signal energy calculation on the first audio and video data to obtain a first signal energy, performing signal energy calculation on the audio and video adaptive filtering data to obtain a second signal energy, and determining a signal energy difference based on the first signal energy and the second signal energy; Determining an audio and video quality detection score based on the signal energy difference; Convert the first audio data in the first audio and video data into a frequency domain audio signal, and perform sampling processing on the frequency domain audio signal to obtain a first discretized audio signal corresponding to each sampling point; obtain a first audio signal amplitude of the first discretized audio signal corresponding to each sampling point, and perform square processing on the first audio signal amplitude to obtain a first audio signal square value corresponding to each first discretized audio signal, and integrate all first audio signal square values ​​to obtain first audio signal energy; Converting the first video data in the first audio and video data into a plurality of first image frames, obtaining an average first pixel value corresponding to each first image frame, squaring the first pixel value average to obtain a square first pixel value corresponding to each first image frame, and integrating all square first pixel values ​​to obtain first video signal energy; Calculating the signal energy sum of the first audio signal energy and the first video signal energy to obtain the first signal energy of the first audio and video data; Performing audio and video optimization processing on the first audio and video data to obtain first audio and video optimized data specifically includes: Acquire a first signal energy corresponding to the first audio and video data, and calculate a gain step size of a filter based on the first signal energy; and obtaining a current gain of the filter, and adjusting the current gain based on the gain step to obtain an adjusted gain of the filter; Controlling the filter to filter the first audio and video data based on the adjustment gain to obtain first audio and video optimized data; The first audio and video optimization data is encoded to obtain first audio and video encoded data, specifically including: Performing data decomposition processing on the first audio and video optimization data to obtain a plurality of audio and video decomposition data; Performing feature extraction on the plurality of audio and video decomposition data respectively to obtain a first feature corresponding to each audio and video decomposition data; Based on the first feature, respectively encode the plurality of audio and video decomposition data to obtain a plurality of first audio and video decomposition encoding data, and integrate the plurality of first audio and video decomposition encoding data to obtain first audio and video encoding data; Performing data decomposition processing on the first audio and video optimization data to obtain a plurality of audio and video decomposition data; performing feature extraction on the plurality of audio and video decomposition data to obtain a first feature corresponding to each audio and video decomposition data, specifically including: The first audio and video optimization data includes first audio optimization data and first video optimization data; Performing data decomposition processing on the first audio optimization data to obtain a plurality of audio decomposition data, and performing data decomposition processing on the first video optimization data to obtain a plurality of video decomposition data; Based on the Mel filter group, a convolution operation is performed on each audio decomposition data to obtain a Mel filter group output value corresponding to each audio decomposition data, and the Mel filter group output value is logarithmically transformed to obtain a logarithmic Mel spectrum coefficient, and a first audio feature corresponding to each audio decomposition data is obtained based on the logarithmic Mel spectrum coefficient; Based on the optical flow algorithm, an optical flow vector map corresponding to each video decomposition data is calculated, a vector mean calculation is performed on the optical flow vector map to obtain an optical flow vector mean, and a first video feature corresponding to each video decomposition data is obtained based on the optical flow vector mean; Based on the first feature, encoding the plurality of audio and video decomposition data is performed respectively to obtain a plurality of first audio and video decomposition encoding data, specifically including: respectively setting an audio bit label for each audio decomposition data, and respectively encoding each audio decomposition data based on a preset audio encoding algorithm to obtain first audio encoding data corresponding to each audio decomposition data; Respectively concatenate the audio bit label, the first audio feature, and the first audio encoding data corresponding to each audio decomposition data to obtain audio encoding data corresponding to each audio decomposition data; Setting a video bit label for each video decomposition data respectively, encoding each video decomposition data respectively based on a preset video encoding algorithm, and obtaining first video encoding data corresponding to each video decomposition data; The video bit label, the first video feature and the first video encoding data corresponding to each video decomposition data are respectively concatenated to obtain the video encoding data corresponding to each video decomposition data.

2. The method for collecting audio and video data based on a 5G terminal device according to claim 1, characterized in that: Performing audio and video acquisition processing on the target object to obtain first audio and video data specifically includes: The audio and video acquisition equipment includes an audio acquisition equipment and a video acquisition equipment; Parsing the audio and video acquisition instruction to determine audio and video acquisition parameters, wherein the audio and video acquisition parameters include resolution, frame rate, and sampling frequency; Based on the audio and video acquisition parameters, device parameters of the audio acquisition device and the video acquisition device are set respectively to obtain audio acquisition device parameters and video acquisition device parameters; Based on the audio acquisition device parameters, perform audio acquisition processing on the target object to obtain first audio data, and based on the video acquisition device parameters, perform video acquisition processing on the target object to obtain first video data; The first audio data and the first video data are integrated to obtain first audio and video data.

3. An audio and video data acquisition system based on a 5G terminal device, the audio and video data acquisition system is used to implement the audio and video data acquisition method according to claim 1, characterized in that: include: Audio and video acquisition module, quality scoring module, audio and video optimization module and audio and video transmission module; The audio and video acquisition module is used to call the API interface of the audio and video acquisition device based on the 5G terminal device, and send an audio and video acquisition instruction to the audio and video acquisition device, so that the audio and video acquisition device performs audio and video acquisition processing on the target object after receiving the audio and video acquisition instruction to obtain the first audio and video data; The quality scoring module is used to input the first audio and video data into a pre-trained quality scoring model so that the quality scoring model performs quality scoring processing on the first audio and video data and outputs an audio and video quality detection score; The audio and video optimization module is used to perform audio and video optimization processing on the first audio and video data to obtain first audio and video optimization data when the audio and video quality detection score does not meet the preset audio and video quality detection score threshold; The audio and video transmission module is used to encode the first audio and video optimization data to obtain first audio and video encoded data, and transmit the first audio and video encoded data to a target receiving device based on a preset network protocol; The quality scoring module is used to input the first audio and video data into a pre-trained quality scoring model so that the quality scoring model performs quality scoring processing on the first audio and video data and outputs an audio and video quality detection score, specifically including: Inputting the first audio and video data into a pre-trained quality scoring model so that the quality scoring model performs adaptive filtering processing on the first audio and video data to obtain audio and video adaptive filtering data; Performing signal energy calculation on the first audio and video data to obtain a first signal energy, performing signal energy calculation on the audio and video adaptive filtering data to obtain a second signal energy, and determining a signal energy difference based on the first signal energy and the second signal energy; Based on the signal energy difference, an audio and video quality detection score is determined.

4. The audio and video data acquisition system based on 5G terminal equipment as claimed in claim 3, characterized in that: The audio and video transmission module is used to encode the first audio and video optimization data to obtain first audio and video encoded data, specifically including: Performing data decomposition processing on the first audio and video optimization data to obtain a plurality of audio and video decomposition data; Performing feature extraction on the plurality of audio and video decomposition data respectively to obtain a first feature corresponding to each audio and video decomposition data; Based on the first feature, the plurality of audio and video decomposition data are respectively encoded to obtain a plurality of first audio and video decomposition encoding data, and the plurality of first audio and video decomposition encoding data are integrated to obtain first audio and video encoding data.

Citation Information

Patent Citations

  • Video processing method and device

    CN107645621A

  • Image processing method and device, electronic equipment and storage medium

    CN114387312A

  • Voice recognition method and system based on audio and video dual modes

    CN114974215A