Distributed audio and video control system and method

By adopting distributed audio and video control systems and methods in audio and video technology, audio and video data filtering and multi-objective optimization encoding and decoding are solved, the problem of low encoding and decoding efficiency in the existing technology is solved, and high-efficiency and low-energy audio and video stream processing is achieved, and high-resolution video is supported.

CN119946352AInactive Publication Date: 2025-05-06HUNAN CHUFENG TECH CO LTD

Patent Information

Application Number
CN202510115821.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing audio and video technology, the encoding and decoding process is inefficient, the storage space occupies a large amount, the playback quality is poor, the processor load and equipment energy consumption are increased, and the encoding and decoding parameters are cumbersome.

Method used

The distributed audio and video control system and method are adopted to collect audio and video data for filtering, and the encoding and decoding process is optimized in combination with a multi-objective optimization algorithm to output high-quality audio and video streams.

Benefits of technology

It improves encoding and decoding efficiency and quality, saves storage space, reduces system energy consumption, reduces processor load, extends device usage time, and adapts to different network conditions, supporting high-resolution video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946352A_ABST
    Figure CN119946352A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of audio and video, and discloses a distributed audio and video control system and method. Comprises: collecting audio and video data; the audio and video data comprises audio data and video data; filtering the audio data to obtain filtered audio data, and filtering the video data to obtain filtered video data; combining and aligning the filtered audio data and the filtered video data based on a time axis to obtain high-quality audio and video data; encoding the high-quality audio and video data, and optimizing the encoding process through a multi-target optimization algorithm to obtain an optimized data stream; decoding the optimized data stream by using a decoder, optimizing the decoding process through an algorithm, and outputting a perfect audio and video stream; the user plays the output perfect audio and video streams through the control interface; the storage space is saved, a large amount of redundant information is reduced, and the coding and decoding efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio and video technology, and more specifically, to a distributed audio and video control system and method. Background Art

[0002] The patent with application publication number CN110868552A discloses a distributed audio and video device and method, which relates to the field of audio and video technology. The device includes: an encoder, a decoder, a data network device, a control network device and a control management platform; the data end of the encoder and the decoder are connected to the data network device; the control end of the encoder, the decoder and the data network device are connected to the control network device; the control network device is connected to the control management platform. The method includes: the control management platform sends a control instruction to the control network device; the control network device modifies the parameters of the encoder, the decoder and the data network device according to the control instruction, so that the encoder and the decoder are connected accordingly. Separate the control network from the data network to avoid mutual interference between the control stream data and the data stream data, which is stable and reliable, and the control is precise and sensitive; the data network cannot access the control network to prevent intrusion into the control management platform and avoid causing the control management platform to crash.

[0003] With the popularization of the Internet and the enrichment of entertainment activities, more and more people nowadays have higher requirements for the quality of audio and video. In the existing field of audio and video technology, the encoding and decoding process always needs to deal with various restrictions, which can easily lead to low encoding and decoding efficiency, excessive storage space occupation, poor playback quality, increased processor load and excessive device energy consumption. In the encoding and decoding process, the parameters of the encoder and decoder need to be adjusted according to various restrictions, making the encoding and decoding process more cumbersome.

[0004] In view of this, the present invention proposes a distributed audio and video control system and method to solve the above problems. Summary of the invention

[0005] In order to overcome the above defects of the prior art and to achieve the above objectives, the present invention provides the following technical solutions: a distributed audio and video control method, comprising:

[0006] S1. Collect audio and video data; audio and video data includes audio data and video data;

[0007] S2. Filter the audio data to obtain filtered audio data, filter the video data to obtain filtered video data; align the filtered audio data and the filtered video data based on the time axis to obtain high-quality audio and video data;

[0008] S3. Encode high-quality audio and video data and optimize the encoding process through a multi-objective optimization algorithm to obtain an optimized data stream;

[0009] S4. A process of decoding the optimized data stream using a decoder and performing decoding through algorithm optimization to output a perfect audio and video stream; the user plays the output perfect audio and video stream through a control interface.

[0010] Furthermore, the method for obtaining the audio data and the video data includes:

[0011] The audio data V(t) is obtained by collecting and transmitting the audio to the encoder and recording the audio for t minutes. The video data R(t) is obtained by collecting and transmitting the video to the encoder and recording the video for t minutes. The start time of the recording is t0 and the end time is tx.

[0012] Furthermore, the method of filtering the audio data includes:

[0013] Initialize the filter parameters, including the cutoff frequency f c and filter coefficient a i ;

[0014] A transfer function H(z) of the filter is defined; based on the transfer function H(z), the audio data is preliminarily processed to obtain preliminary audio data;

[0015] Perform Laplace transform on the audio data V(t) to obtain a complex function V(z);

[0016] The formula for the Laplace transform is:

[0017] Among them, V(z) is a complex function, and the complex variable is z;

[0018] The formula for preliminary processing is:

[0019] Wherein, N represents the filter order, H(z) is the filter transfer function, V(z) is a complex function, and Y(z) represents preliminary audio data;

[0020] Filter order Where Ap is the maximum attenuation of the filter's passband, As is the minimum attenuation of the filter's stopband, fs is the filter's stopband cutoff frequency, fp is the filter's passband cutoff frequency, and BW is the device's network bandwidth;

[0021] The transfer function H(z) is calculated as:

[0022] Among them, f c represents the cutoff frequency, N represents the filter order, z is a complex variable, a i represents the filter coefficient;

[0023] Performing final processing on the preliminary audio data to obtain filtered audio data;

[0024] The final processing formula is: Among them, y(t) represents the audio data after filtering, Y(z) represents the output of the filter, represents the inverse Laplace operation.

[0025] Furthermore, the method of filtering the video data includes:

[0026] The pixel value of the filtered output video frame image is defined as I, and the calculation formula is:

[0027] Among them, (i, j) represents the position of any pixel block in the current video frame image, I(i, j) represents the grayscale value of the pixel block at position (i, j) in the current video frame image, (k, l) represents the position of the pixel block in the neighborhood, ω r represents the grayscale weight, ω s represents the spatial weight, and I represents the pixel value of the output video frame image after filtering;

[0028] Spatial weight ω s The calculation formula is:

[0029] Among them, ω s represents the spatial weight, (i, j) represents the position of the current pixel block, (k, l) represents the position of the pixel block in the neighborhood, σ s represents the standard deviation of the spatial Gaussian function, and exp represents the natural exponential function;

[0030] Gray weight ω r The calculation formula is:

[0031] Among them, ω r represents the grayscale weight, I(i,j) represents the grayscale value of the pixel block at position (i,j) in the current video frame image, I(k,l) represents the grayscale value of the pixel block in the neighborhood, σ r Represents the standard deviation of the grayscale Gaussian function;

[0032] The pixel value of the video frame image obtained by filtering is the pixel value of the optimized video frame image, and the set of optimized video frame images is the filtered video data r(t);

[0033] By adding time stamps to the filtered audio data and the filtered video data and combining and aligning the two types of data based on the time axis, high-quality audio and video data Q(t) with synchronized audio and video is obtained.

[0034] Furthermore, the optimized data stream is obtained by:

[0035] Use distributed encoding and decoding, and arrange x encoders and x decoders;

[0036] Construct a multi-objective optimization mathematical model and define four objective optimization functions, including encoding delay f1, bit rate f2, system power consumption f3, and audio and video quality f4;

[0037] Define ping0 as the encoding delay threshold, B0 as the preset target bit rate, and E0 as the preset target system power consumption;

[0038] Define the peak signal-to-noise ratio PSNR, and the peak signal-to-noise ratio threshold is P0;

[0039] Considering f1, f2 and f3 as constraints and f4 as the objective function, the mathematical model is constructed as follows: Among them, maxf4 means maximizing the audio and video quality, and fi means the constraint condition;

[0040] Wherein, f1 represents the encoding delay, ping0 represents the encoding delay threshold; f2 represents the bit rate, B0 represents the preset target bit rate; f3 represents the system power consumption, E0 represents the preset target system power consumption;

[0041] By dynamically adjusting the parameters of the constraint conditions, the constraint condition value is made less than or equal to the preset threshold, the objective function value is output, and the peak signal-to-noise ratio is calculated to determine whether the optimization is completed;

[0042] The formula for calculating the peak signal-to-noise ratio is:

[0043] Among them, PSNR represents peak signal-to-noise ratio, MAXI represents the maximum pixel value in the video frame image after encoding and compression, and MSE represents mean square error;

[0044] If the calculated peak signal-to-noise ratio is greater than or equal to the preset peak signal-to-noise ratio threshold P0, the optimization is completed and the encoder outputs the optimized data stream.

[0045] Furthermore, the coding delay is calculated by:

[0046] Encoding delay includes network delay and buffer delay;

[0047] The formula for calculating network delay is:

[0048] ping = ω1×p1+ω2×p2; where ping represents network delay, p1 represents data transmission delay, p2 represents data propagation delay, ω1 represents the weight of data transmission, and ω2 represents the weight of data propagation;

[0049] Data transmission delay Among them, BW refers to the network bandwidth of the device, and q(t) refers to the data volume of high-quality audio and video data Q(t);

[0050] Data propagation delay Among them, V0 refers to the speed of data signal transmission;

[0051] The formula for calculating buffer latency is:

[0052] Where buff represents buffer delay, Λ0 represents the inverse of the data signal transmission speed, maxi represents the maximum received frame size, avgi represents the average received frame size, noise represents the noise coefficient, and ω3 represents the noise coefficient weight;

[0053] Encoding delay f1 = ping + buff; where ping represents network delay and buff represents buffer delay;

[0054] If it is calculated that the coding delay f1 is greater than the coding delay threshold ping0, the data transmission weight ω1, the data propagation weight ω2 and the noise coefficient weight ω3 are adjusted until the coding delay f1 is less than or equal to ping0.

[0055] Furthermore, the bit rate is calculated as follows:

[0056] Bit rate Wherein, t represents time in minutes; q(t) represents the data volume of high-quality audio and video data Q(t); ω4 represents the bit rate weight; and μ represents the coding efficiency;

[0057] Coding efficiency Wherein, q(t) represents the data size of high-quality audio and video data Q(t), and q1 represents the size of the compressed data;

[0058] If the bit rate f2 is calculated to be greater than the preset target bit rate B0, the adjustment parameters are recalculated until the bit rate f2 is less than or equal to B0.

[0059] Furthermore, the system power consumption is calculated as follows:

[0060] Define the preliminary calculated system power consumption f0;

[0061] f0=P(t)=V×L×ω5; where V represents voltage, L represents current, ω5 represents system power consumption weight, and f0 represents the preliminarily calculated system power consumption;

[0062] Since the system power consumption varies with time, the system power consumption is expressed as the average system power consumption P avgTo express it, the formula is:

[0063] Among them, P avg represents the average system power consumption, T represents the time interval, and f3 represents the system power consumption;

[0064] If the calculated system power consumption f3 is greater than the preset target system power consumption E0, the adjustment parameters are recalculated until the system power consumption f3 is less than or equal to E0.

[0065] Furthermore, the decoder decodes the optimized data stream in the following manner:

[0066] Define coding units, including prediction blocks and transform blocks;

[0067] Use the prediction block to perform intra-frame prediction and output the predicted pixel value; use the transform block to perform inverse transform and output the residual signal;

[0068] Define the entropy decoding module, including the transform coefficient and reconstructed image modules;

[0069] The prediction block and the transformation block are entropy decoded by using an entropy decoding module to obtain inverse transformation input data and intra-frame prediction input data;

[0070] Receiving the optimized data stream output by the encoder, and parsing the data stream to obtain the encoding unit and related encoding information;

[0071] Loop through all prediction blocks and all transform blocks in the coding unit in turn; identify each transform block and determine whether it is an all-zero block. If it is an all-zero block, skip the subsequent steps and identify the next transform block. If it is a non-zero block, perform entropy decoding through the entropy decoding module; then perform inverse transform and intra-frame prediction;

[0072] The module separation algorithm is used to parallelize the inverse transformation and intra-frame prediction processes; the transformation coefficients in the entropy decoding module are derived and separated into the reconstructed image module; the inverse transformation input and intra-frame prediction input are obtained from the entropy decoding;

[0073] The entropy decoding output data is input into the image reconstruction module, the residual signal output by the inverse transform and the predicted pixel value output by the intra-frame prediction are combined to obtain the reconstructed image block, and the reconstructed image blocks are pieced together according to the index of the image block to form a complete frame image;

[0074] Combining each complete frame of image according to the time index and combining it with the audio will obtain a decoded complete audio and video stream; the user interacts with the system through the control interface to play the output complete audio and video stream.

[0075] A distributed audio and video control system, which is used to implement a distributed audio and video control method, comprising:

[0076] A data acquisition module is used to collect audio and video data, which includes audio data and video data;

[0077] A filtering processing module is used to filter the audio data to obtain filtered audio data, and filter the video data to obtain filtered video data; based on the time axis, the filtered audio data and the filtered video data are combined and aligned to obtain high-quality audio and video data;

[0078] The encoding optimization module is used to encode high-quality audio and video data and optimize the encoding process through a multi-objective optimization algorithm to obtain an optimized data stream;

[0079] The decoding optimization module is used to use the decoder to decode the optimized data stream and perform decoding through algorithm optimization to output a perfect audio and video stream; the user plays the output perfect audio and video stream through the control interface; and each module is connected by wired and / or wireless means.

[0080] The technical effects and advantages of a distributed audio and video control system and method of the present invention are as follows:

[0081] Distributed audio and video control is achieved by collecting audio and video data, optimizing the data, and improving the process by adding optimization algorithms during the audio and video stream encoding and decoding process. Compared with existing experience, the efficiency and quality of encoding and decoding are improved by pre-filtering the collected audio and video data and using multi-objective optimization algorithms and module separation algorithms to optimize encoding and decoding. The optimization of the encoding and decoding process saves storage space and reduces a large amount of redundant information. It reduces system energy consumption, reduces processor load, and extends the use time of the equipment. It enables the system to adapt to different network conditions and dynamically adjusts encoding parameters to adapt to different bandwidth and delay conditions to ensure the stability and continuity of audio and video streams. The efficient encoding and decoding technology is sufficient to support the encoding and decoding of high-resolution videos such as 4K or 8K, meeting the current user's pursuit of ultra-high-definition video content. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 A schematic diagram of a distributed audio and video control method of the present invention;

[0083] Figure 2 It is a schematic diagram of a distributed audio and video control system of the present invention. DETAILED DESCRIPTION

[0084] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0085] Example 1

[0086] See also Figure 1 As shown, a distributed audio and video control method described in this embodiment includes:

[0087] S1. Collect audio and video data; audio and video data includes audio data and video data;

[0088] S2. Filtering the audio data to obtain filtered audio data, filtering the video data to obtain filtered video data; combining and aligning the filtered audio data and the filtered video data based on the time axis to obtain high-quality audio and video data;

[0089] S3. Encode high-quality audio and video data and optimize the encoding process through a multi-objective optimization algorithm to obtain an optimized data stream;

[0090] S4. The decoder is used to decode the optimized data stream and the decoding process is performed through the algorithm optimization to output the perfect audio and video stream; the user plays the perfect audio and video stream output through the control interface;

[0091] The method for obtaining the audio data and the video data includes:

[0092] The audio data V(t) is obtained by collecting and transmitting the audio to the encoder and recording the audio data for t minutes. The video data R(t) is obtained by collecting and transmitting the video to the encoder and recording the video data for t minutes. The start time of the recording is t0 and the end time is tx.

[0093] The method of filtering the audio data includes:

[0094] Initialize the filter parameters, including the cutoff frequency f c and filter coefficient a i ;

[0095] A transfer function H(z) of the filter is defined; based on the transfer function H(z), the audio data is preliminarily processed to obtain preliminary audio data;

[0096] Perform Laplace transform on the audio data V(t) to obtain a complex function V(z);

[0097] The formula for the Laplace transform is:

[0098] Among them, V(z) is a complex function, and the complex variable is z;

[0099] The formula for preliminary processing is:

[0100] Wherein, N represents the filter order, H(z) is the filter transfer function, V(z) is a complex function, and Y(z) represents preliminary audio data;

[0101] Filter order Where Ap is the maximum attenuation of the filter's passband, As is the minimum attenuation of the filter's stopband, fs is the filter's stopband cutoff frequency, fp is the filter's passband cutoff frequency, and BW is the device's network bandwidth;

[0102] The transfer function H(z) is calculated as:

[0103] Among them, f c represents the cutoff frequency, N represents the filter order, z is a complex variable, a i represents the filter coefficient;

[0104] It should be noted that filtering the audio data will represent the audio signal as a set of its component frequencies and convert it into frequency domain data. Therefore, the final output needs to convert the frequency domain data into time domain data through inverse Laplace transform to obtain the filtered audio data;

[0105] Performing final processing on the preliminary audio data to obtain filtered audio data;

[0106] The final processing formula is: Among them, y(t) represents the audio data after filtering, Y(z) represents the output of the filter, represents the inverse Laplace operation;

[0107] The method of filtering the video data includes:

[0108] The pixel value of the filtered output video frame image is defined as I, and the calculation formula is:

[0109] Among them, (i, j) represents the position of any pixel block in the current video frame image, I(i, j) represents the grayscale value of the pixel block at position (i, j) in the current video frame image, (k, l) represents the position of the pixel block in the neighborhood, ω r represents the grayscale weight, ω s represents the spatial weight, and I represents the pixel value of the output video frame image after filtering;

[0110] Spatial weight ω s The calculation formula is:

[0111] Among them, ω s represents the spatial weight, (i, j) represents the position of the current pixel block, (k, l) represents the position of the pixel block in the neighborhood, σ s represents the standard deviation of the spatial Gaussian function, and exp represents the natural exponential function;

[0112] Gray weight ω r The calculation formula is:

[0113] Among them, ω r represents the grayscale weight, I(i,j) represents the grayscale value of the pixel block at position (i,j) in the current video frame image, I(k,l) represents the grayscale value of the pixel block in the neighborhood, σ r Represents the standard deviation of the grayscale Gaussian function;

[0114] The pixel value of the video frame image obtained by filtering is the pixel value of the optimized video frame image, and the set of optimized video frame images is the filtered video data r(t);

[0115] It should be noted that filtering the audio data and video data reduces the noise in the audio data and video data, filters the redundant information in the data, improves the overall quality of the data, and indirectly improves the efficiency of subsequent intelligent coding;

[0116] Since the recording start time of the filtered audio data and the filtered video data are both t0, the end time is tx, and the recording duration is t minutes; by adding timestamps to the filtered audio data and the filtered video data, and combining and aligning the two data based on the time axis, high-quality audio and video data Q(t) with synchronized audio and video is obtained;

[0117] The method for obtaining the optimized data stream includes:

[0118] Distributed encoding and decoding is adopted, and x encoders and x decoders are arranged so that the encoders and decoders can process and transmit data in parallel;

[0119] It should be noted that distributed coding and decoding technology refers to the technology of dividing audio and video signals into multiple data blocks, and then processing and transmitting them in parallel through multiple encoders and decoders. Compared with ordinary coding and decoding technology, distributed coding and decoding technology can distribute tasks to multiple nodes for processing, which improves the efficiency and quality of coding and decoding. Moreover, as the amount of processed data increases, the processing capacity of the system can be improved by adding nodes, which is convenient for optimization.

[0120] Construct a multi-objective optimization mathematical model and define four objective optimization functions, including encoding delay f1, bit rate f2, system power consumption f3, and audio and video quality f4;

[0121] Define ping0 as the encoding delay threshold, B0 as the preset target bit rate, and E0 as the preset target system power consumption

[0122] Define the peak signal-to-noise ratio PSNR, and the peak signal-to-noise ratio threshold is P0;

[0123] Considering f1, f2 and f3 as constraints and f4 as the objective function, the mathematical model is constructed as follows: Among them, maxf4 means maximizing the audio and video quality, and fi means the constraint condition;

[0124] Wherein, f1 represents the encoding delay, ping0 represents the encoding delay threshold; f2 represents the bit rate, b0 represents the preset target bit rate; f3 represents the system power consumption, E0 represents the preset target system power consumption;

[0125] By dynamically adjusting the parameters of the constraint conditions, the constraint condition value is made less than or equal to the preset threshold, the objective function value is output, and the peak signal-to-noise ratio is calculated to determine whether the optimization is completed;

[0126] The formula for calculating the peak signal-to-noise ratio is:

[0127] Among them, PSNR represents peak signal-to-noise ratio, MAXI represents the maximum pixel value in the video frame image after encoding and compression, and MSE represents mean square error;

[0128] If the calculated peak signal-to-noise ratio is greater than or equal to the preset peak signal-to-noise ratio threshold P0, the optimization is completed and the encoder outputs the optimized data stream;

[0129] The coding delay is calculated by:

[0130] Encoding delay includes network delay and buffer delay;

[0131] The formula for calculating network delay is:

[0132] ping = ω1×p1+ω2×p2; where ping represents network delay, p1 represents data transmission delay, p2 represents data propagation delay, ω1 represents the weight of data transmission, and ω2 represents the weight of data propagation;

[0133] Data transmission delay Among them, BW refers to the network bandwidth of the device, and q(t) refers to the data volume of high-quality audio and video data Q(t);

[0134] It should be noted that data transmission delay refers to the time it takes for data to propagate on the transmission medium, and data propagation delay refers to the time it takes for a data packet to propagate from the sender to the receiver;

[0135] Data propagation delay Among them, V0 refers to the speed of data signal transmission;

[0136] It should be noted that the data signal transmission speed is related to the type and material of the transmission medium. For example, the propagation speed of optical signals in optical fibers is much faster than that in copper wires. At the same time, the refractive index of certain materials in the medium will also affect the speed of optical signals.

[0137] The formula for calculating buffer latency is:

[0138] Where buff represents buffer delay, Λ0 represents the inverse of the data signal transmission speed, maxi represents the maximum received frame size, avgi represents the average received frame size, noise represents the noise coefficient, and ω3 represents the noise coefficient weight;

[0139] It should be noted that the average frame size refers to the average value of all frames received up to any moment;

[0140] Encoding delay f1 = ping + buff; where ping represents network delay and buff represents buffer delay;

[0141] If the calculated coding delay f1 is greater than the coding delay threshold ping0, the data transmission weight ω1, the data propagation weight ω2 and the noise coefficient weight ω3 are adjusted until the coding delay f1 is less than or equal to ping0;

[0142] The bit rate is calculated as:

[0143] Bit rate Wherein, t represents time in minutes; q(t) represents the data volume of high-quality audio and video data Q(t); ω4 represents the bit rate weight; and μ represents the coding efficiency;

[0144] Coding efficiency Wherein, q(t) represents the data size of high-quality audio and video data Q(t), and q1 represents the size of the compressed data;

[0145] If the bit rate f2 is calculated to be greater than the preset target bit rate B0, the adjustment parameters are recalculated until the bit rate f2 is less than or equal to B0.

[0146] It should be noted that compression is required when encoding the original data, thus generating compressed data; if the network bandwidth is insufficient, the bit rate can be reduced by adjusting the bit rate weight ω4 and the coding efficiency μ, so that more data can be transmitted with limited bandwidth in the same time;

[0147] The system power consumption is calculated as follows:

[0148] Define the preliminary calculated system power consumption f0;

[0149] f0=P(t)=V×L×ω5; where V represents voltage, L represents current, ω5 represents system power consumption weight, and f0 represents the preliminarily calculated system power consumption;

[0150] It should be noted that the system power consumption changes dynamically. The power consumption can be adjusted by adjusting the system power consumption weight ω5. When the amount of data processed increases, the power consumption is increased, and when the amount of data processed decreases, the power consumption is reduced. This dynamic control method can greatly reduce the power consumption of the overall encoding process and reduce the waste of resources.

[0151] Since the system power consumption varies with time, the system power consumption can be expressed as the average system power consumption P avg To express it, the formula is:

[0152] Among them, P avg represents the average system power consumption, T represents the time interval, and f3 represents the system power consumption;

[0153] If the calculated system power consumption f3 is greater than the preset target system power consumption E0, the adjustment parameters are recalculated until the system power consumption f3 is less than or equal to E0;

[0154] If the calculated peak signal-to-noise ratio is greater than or equal to the preset peak signal-to-noise ratio threshold value P0, the optimization is completed and the encoder outputs the optimized data stream; the encoding parameter combination at this time is the best parameter combination, the audio and video quality is the best, and the system power consumption is the lowest;

[0155] The decoder decodes the optimized data stream in the following manner:

[0156] Define coding units, including prediction blocks and transform blocks;

[0157] Use the prediction block to perform intra-frame prediction and output the predicted pixel value; use the transform block to perform inverse transform and output the residual signal;

[0158] Define the entropy decoding module, including the transform coefficient and reconstructed image modules;

[0159] The prediction block and the transformation block are entropy decoded by using an entropy decoding module to obtain inverse transformation input data and intra-frame prediction input data;

[0160] Receive the optimized data stream output by the encoder, and parse the data stream to obtain the encoding unit and related encoding information (such as resolution, frame rate, color coding, file size, etc.);

[0161] Loop through all prediction blocks and all transform blocks in the coding unit in turn; identify each transform block and determine whether it is an all-zero block. If it is an all-zero block, skip the subsequent steps and identify the next transform block. If it is a non-zero block, perform entropy decoding through the entropy decoding module; then perform inverse transform and intra-frame prediction;

[0162] The module separation algorithm is used to parallelize the inverse transformation and intra-frame prediction processes; the transformation coefficients in the entropy decoding module are derived and separated into the reconstructed image module; the inverse transformation input and intra-frame prediction input are obtained from the entropy decoding;

[0163] It should be noted that the module separation algorithm removes data dependency from the operation of separating the transform coefficients; in subsequent operations, since data dependency is removed, the input of the inverse transform only needs to obtain the coefficient information of the current transform block and the index value of the current block from entropy decoding, and the intra-frame prediction also only needs to obtain the mode information of the current prediction block and the prediction block position index from entropy decoding; the algorithm parallelizes each module, improves data throughput, reduces the dependency between each execution module, and improves decoding efficiency;

[0164] The output data is input to the image reconstruction module, the residual output of the inverse transform is combined with the predicted value output by the intra-frame prediction to obtain a reconstructed image block, and the reconstructed image blocks are stitched together according to the index of the image block to form a complete frame image;

[0165] Combining each complete frame of image according to the time index and combining it with the audio will obtain a decoded complete audio and video stream; the user interacts with the system through the control interface to play the output complete audio and video stream.

[0166] It should be noted that during the encoding and decoding process, users can pause or cancel and restore to the original data at any time through the control interface; during the audio and video playback process, the system collects feedback information from terminal users in real time (such as blurry image quality, high buffering times, frame drops, etc.), providing an information basis for further optimization of subsequent algorithms;

[0167] This embodiment realizes distributed audio and video control by collecting audio and video data, optimizing the data, and improving the process by adding optimization algorithms during the audio and video stream encoding and decoding process. Compared with existing experience, the efficiency and quality of encoding and decoding are improved by pre-filtering the collected audio and video data and using multi-objective optimization algorithms and module separation algorithms to optimize encoding and decoding. The optimization of the encoding and decoding process saves storage space and reduces a large amount of redundant information. It reduces system energy consumption, reduces processor load, and extends the use time of the device. It enables the system to adapt to different network conditions and dynamically adjusts encoding parameters to adapt to different bandwidth and delay conditions to ensure the stability and continuity of audio and video streams. The efficient encoding and decoding technology is sufficient to support the encoding and decoding of high-resolution videos such as 4K or 8K, meeting the current user's pursuit of ultra-high-definition video content.

[0168] Example 2

[0169] See also Figure 2 As shown, the part not described in detail in this embodiment is described in Example 1, which provides a distributed audio and video control system, including:

[0170] A data acquisition module is used to collect audio and video data, which includes audio data and video data;

[0171] A filtering processing module is used to filter the audio data to obtain filtered audio data, and filter the video data to obtain filtered video data; based on the time axis, the filtered audio data and the filtered video data are combined and aligned to obtain high-quality audio and video data;

[0172] The encoding optimization module is used to encode high-quality audio and video data and optimize the encoding process through a multi-objective optimization algorithm to obtain an optimized data stream;

[0173] The decoding optimization module is used to use the decoder to decode the optimized data stream and perform decoding through algorithm optimization to output a perfect audio and video stream; the user plays the output perfect audio and video stream through the control interface; and each module is connected by wired and / or wireless means.

[0174] Example 3

[0175] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the distributed audio and video control method provided above is implemented.

[0176] Since the electronic device introduced in this embodiment is an electronic device used to implement a distributed audio and video control method in the embodiment of the present application, based on the distributed audio and video control method introduced in the embodiment of the present application, a person skilled in the art can understand the specific implementation of the electronic device of the present embodiment and its various variations, so how the electronic device implements the method in the embodiment of the present application is not described in detail here. As long as a person skilled in the art implements the electronic device used in a distributed audio and video control method in the embodiment of the present application, it belongs to the scope of protection of this application.

[0177] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0178] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A distributed audio and video control method, characterized in that: include: S1. Collect audio and video data; audio and video data includes audio data and video data; S2. Filtering the audio data to obtain filtered audio data, and filtering the video data to obtain filtered video data; Combine and align the filtered audio data and the filtered video data based on the time axis to obtain high-quality audio and video data; S3. Encode high-quality audio and video data and optimize the encoding process through a multi-objective optimization algorithm to obtain an optimized data stream; S4. The process of decoding the optimized data stream by using a decoder and performing decoding through algorithm optimization is used to output a perfect audio and video stream; The user plays the output complete audio and video stream through the control interface.

2. A distributed audio and video control method according to claim 1, characterized in that: The method for obtaining the audio data and the video data includes: The audio data V(t) is obtained by collecting and transmitting the audio to the encoder and recording the audio for t minutes. The video data R(t) is obtained by collecting and transmitting the video to the encoder and recording the video for t minutes. The start time of the recording is t0 and the end time is tx.

3. A distributed audio and video control method according to claim 2, characterized in that: The method of filtering the audio data includes: Initialize the filter parameters, including the cutoff frequency f c and filter coefficient a i ; A transfer function H(z) of the filter is defined; based on the transfer function H(z), the audio data is preliminarily processed to obtain preliminary audio data; Perform Laplace transform on the audio data V(t) to obtain a complex function V(z); The formula for the Laplace transform is: Among them, V(z) is a complex function, and the complex variable is z; The formula for preliminary processing is: Wherein, N represents the filter order, H(z) is the filter transfer function, V(z) is a complex function, and Y(z) represents preliminary audio data; Filter order Where Ap is the maximum attenuation of the filter's passband, As is the minimum attenuation of the filter's stopband, fs is the filter's stopband cutoff frequency, fp is the filter's passband cutoff frequency, and BW is the device's network bandwidth; The transfer function H(z) is calculated as: Among them, f c represents the cutoff frequency, N represents the filter order, z is a complex variable, a i represents the filter coefficient; Performing final processing on the preliminary audio data to obtain filtered audio data; The final processing formula is: Among them, y(t) represents the audio data after filtering, Y(z) represents the output of the filter, represents the inverse Laplace operation.

4. A distributed audio and video control method according to claim 3, characterized in that: The method of filtering the video data includes: The pixel value of the filtered output video frame image is defined as I, and the calculation formula is: Among them, (i, j) represents the position of any pixel block in the current video frame image, I(i, j) represents the grayscale value of the pixel block at position (i, j) in the current video frame image, (k, l) represents the position of the pixel block in the neighborhood, ω r represents the grayscale weight, ω s represents the spatial weight, and I represents the pixel value of the output video frame image after filtering; Spatial weight ω s The calculation formula is: Among them, ω s represents the spatial weight, (i, j) represents the position of the current pixel block, (k, l) represents the position of the pixel block in the neighborhood, σ s represents the standard deviation of the spatial Gaussian function, and exp represents the natural exponential function; Gray weight ω r The calculation formula is: Among them, ω r represents the grayscale weight, I(i,j) represents the grayscale value of the pixel block at position (i,j) in the current video frame image, I(k,l) represents the grayscale value of the pixel block in the neighborhood, σ r Represents the standard deviation of the grayscale Gaussian function; The pixel value of the video frame image obtained by filtering is the pixel value of the optimized video frame image, and the set of optimized video frame images is the filtered video data r(t); By adding time stamps to the filtered audio data and the filtered video data and combining and aligning the two types of data based on the time axis, high-quality audio and video data Q(t) with synchronized audio and video is obtained.

5. A distributed audio and video control method according to claim 4, characterized in that: The method for obtaining the optimized data stream includes: Use distributed encoding and decoding, and arrange x encoders and x decoders; Construct a multi-objective optimization mathematical model and define four objective optimization functions, including encoding delay f1, bit rate f2, system power consumption f3, and audio and video quality f4; Define ping0 as the encoding delay threshold, B0 as the preset target bit rate, and E0 as the preset target system power consumption; Define the peak signal-to-noise ratio PSNR, and the peak signal-to-noise ratio threshold is P0; Considering f1, f2 and f3 as constraints and f4 as the objective function, the mathematical model is constructed as follows: Among them, maxf4 means maximizing the audio and video quality, and fi means the constraint condition; Wherein, f1 represents the encoding delay, ping0 represents the encoding delay threshold; f2 represents the bit rate, B0 represents the preset target bit rate; f3 represents the system power consumption, E0 represents the preset target system power consumption; By dynamically adjusting the parameters of the constraint conditions, the constraint condition value is made less than or equal to the preset threshold, the objective function value is output, and the peak signal-to-noise ratio is calculated to determine whether the optimization is completed; The formula for calculating the peak signal-to-noise ratio is: Among them, PSNR represents peak signal-to-noise ratio, MAXI represents the maximum pixel value in the video frame image after encoding and compression, and MSE represents mean square error; If the calculated peak signal-to-noise ratio is greater than or equal to the preset peak signal-to-noise ratio threshold P0, the optimization is completed and the encoder outputs the optimized data stream.

6. A distributed audio and video control method according to claim 5, characterized in that: The coding delay is calculated by: Encoding delay includes network delay and buffer delay; The formula for calculating network delay is: ping = ω1×p1+ω2×p2; where ping represents network delay, p1 represents data transmission delay, p2 represents data propagation delay, ω1 represents the weight of data transmission, and ω2 represents the weight of data propagation; Data transmission delay Among them, BW refers to the network bandwidth of the device, and q(t) refers to the data volume of high-quality audio and video data Q(t); Data propagation delay Among them, V0 refers to the speed of data signal transmission; The formula for calculating buffer latency is: Where buff represents buffer delay, Λ0 represents the inverse of the data signal transmission speed, maxi represents the maximum received frame size, avgi represents the average received frame size, noise represents the noise coefficient, and ω3 represents the noise coefficient weight; Encoding delay f1 = ping + buff; where ping represents network delay and buff represents buffer delay; If it is calculated that the coding delay f1 is greater than the coding delay threshold ping0, the data transmission weight ω1, the data propagation weight ω2 and the noise coefficient weight ω3 are adjusted until the coding delay f1 is less than or equal to ping0.

7. A distributed audio and video control method according to claim 6, characterized in that: The bit rate is calculated as: Bit rate Wherein, t represents time in minutes; q(t) represents the data volume of high-quality audio and video data Q(t); ω4 represents the bit rate weight; and μ represents the coding efficiency; Coding efficiency Wherein, q(t) represents the data size of high-quality audio and video data Q(t), and q1 represents the compressed data size; If the bit rate f2 is calculated to be greater than the preset target bit rate B0, the adjustment parameters are recalculated until the bit rate f2 is less than or equal to B0.

8. A distributed audio and video control method according to claim 7, characterized in that: The system power consumption is calculated as follows: Define the preliminary calculated system power consumption f0; f0=P(t)=V×L×ω5; Among them, V represents voltage, L represents current, ω5 represents system power consumption weight, and f0 represents the preliminarily calculated system power consumption; Since the system power consumption varies with time, the system power consumption is expressed as the average system power consumption P avg To express it, the formula is: Among them, P avg represents the average system power consumption, T represents the time interval, and f3 represents the system power consumption; If the calculated system power consumption f3 is greater than the preset target system power consumption E0, the adjustment parameters are recalculated until the system power consumption f3 is less than or equal to E0.

9. A distributed audio and video control method according to claim 8, characterized in that: The decoder decodes the optimized data stream in the following manner: Define coding units, including prediction blocks and transform blocks; Use the prediction block to perform intra-frame prediction and output the predicted pixel value; use the transform block to perform inverse transform and output the residual signal; Define the entropy decoding module, including the transform coefficient and reconstructed image modules; The prediction block and the transformation block are entropy decoded by using an entropy decoding module to obtain inverse transformation input data and intra-frame prediction input data; Receiving the optimized data stream output by the encoder, and parsing the data stream to obtain the encoding unit and related encoding information; Loop through all prediction blocks and all transform blocks in the coding unit in turn; identify each transform block and determine whether it is an all-zero block. If it is an all-zero block, skip the subsequent steps and identify the next transform block. If it is a non-zero block, perform entropy decoding through the entropy decoding module; then perform inverse transform and intra-frame prediction; The module separation algorithm is used to parallelize the inverse transformation and intra-frame prediction processes; the transformation coefficients in the entropy decoding module are derived and separated into the reconstructed image module; the inverse transformation input and intra-frame prediction input are obtained from the entropy decoding; The entropy decoding output data is input into the image reconstruction module, the residual signal output by the inverse transform and the predicted pixel value output by the intra-frame prediction are combined to obtain the reconstructed image block, and the reconstructed image blocks are pieced together according to the index of the image block to form a complete frame image; Combining each complete frame of image according to the time index and combining it with the audio will obtain a decoded complete audio and video stream; the user interacts with the system through the control interface to play the output complete audio and video stream.

10. A distributed audio and video control system, used to implement a distributed audio and video control method according to any one of claims 1 to 9, characterized in that: include: A data acquisition module is used to collect audio and video data, which includes audio data and video data; A filtering processing module is used to filter the audio data to obtain filtered audio data, and to filter the video data to obtain filtered video data; Combine and align the filtered audio data and the filtered video data based on the time axis to obtain high-quality audio and video data; The encoding optimization module is used to encode high-quality audio and video data and optimize the encoding process through a multi-objective optimization algorithm to obtain an optimized data stream; A decoding optimization module is used to decode the optimized data stream using a decoder and perform decoding through algorithm optimization to output a perfect audio and video stream; The user plays the output complete audio and video stream through the control interface; each module is connected by wire and / or wireless means.

Citation Information

Patent Citations

  • Distributed audio and video device and method

    CN110868552A

Cited By

  • Satellite channel high fault tolerance audio and video joint coding and decoding method

    CN120935358A