A Multimedia Data Encoding and Decoding Method for Low-Latency and High-Reliability Communication
The multimedia data encoding scheme addresses the challenge of low-latency, high-reliability communication by adapting to user perception and network variability, ensuring efficient and flexible multimedia data transmission through grain-size adaptive models and multi-path cooperative encoding.
Patent Information
- Application Number
- CN202211488181.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The existing multimedia data encoding methods are difficult to meet the low-latency and high-reliability requirements of emerging immersive multimedia communications, especially for video and tactile encoding. The existing methods cannot take into account high compression rate, personalized characteristics, scalability and flexibility, resulting in insufficient communication delay and reliability.
Design a granular adaptive ‘perceptual dead zone’ model, and flexible deconstruction and hierarchical encoding of multimedia signals are carried out based on Weber’s law. Through multi-path cooperative transmission and sliding window reception mechanism, combined with spatiotemporal correlation coding and entropy coding, the main description layer and enhanced description layer are generated to realize personalized encoding adaptation.
It realizes low latency and high reliability multimedia communication, adapts to the perception ability of different users, improves coding efficiency and network environment adaptability, reduces encoding and decoding delay and redundancy, and ensures the fluency and quality of multimedia services.
Smart Images

Figure CN115866267B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data encoding and decoding, and particularly relates to a multimedia data encoding and decoding method for low-latency and high-reliability communication. Background Art
[0002] Currently, multimedia services are developing towards ultra-high resolution and ultra-high fidelity. New emerging 360° video services, VR / AR services, and cross-modal communication services that include audio, video, and touch have pushed the immersion of the user experience of multimedia communication services to the extreme. However, the existing such immersive multimedia services are still limited to local areas, and it is still a long way to go to achieve remote immersive multimedia interaction. One important reason is that there is a contradiction between the ultra-high requirements of such services for communication latency and reliability and the interaction of large amounts of data, and it is difficult to guarantee ultra-high communication metrics.
[0003] It has been confirmed that significant improvement in communication reliability and reduction in transmission latency can be achieved through multipath concurrent transmission and flexible collaboration. However, existing multimedia data coding methods are difficult to provide effective support for low-latency and high-reliability communication. Taking video coding as an example, existing mainstream video coding schemes such as HEVC, SHVC, and MDC respectively pursue the compression coding efficiency of video, the scalability of video streams, and the flexibility of video streams. HEVC achieves high-efficiency video coding by exploiting the spatio-temporal correlation of video signals, but it has poor scalability and flexibility of video (see: Zhu Xiuchang, Li Xin, Chen Jie. The New Generation Video Coding Standard - HEVC [J]. Journal of Nanjing University of Posts and Telecommunications (Natural Science Edition), 2013, 33(3): 1-11.); Although SHVC can achieve the scalability of video coding while maintaining high coding efficiency, the dependence of its upper-layer data decoding on lower-layer data makes it difficult to provide support for flexible collaboration (see: Boyce J, Ye Y, Chen J, et al. Overview of SHVC: Scalable Extensions of the High Efficiency Video Coding Standard [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2016, 26(1): 20-34.); MDC is the most scalable and flexible by supporting independent encoding and decoding of each layer and combined encoding and decoding of multiple layers, but the additional coding redundancy introduced by it often leads to the problem of low coding efficiency (see: Kazemi M, Sadeghi K, Shirmohammadi S. A Mixed Layer Multiple Description Video Coding Scheme [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2012, 22(2): 202-215.). Therefore, these methods cannot adapt to emerging low-latency and high-reliability multimedia communication services. Especially for haptic coding, existing methods are still limited to one-dimensional haptic signal compression and reconstruction based on Weber's law (see: Antonakoglou K, Xu X, Steinbach E, et al. Toward Haptic Communications Over the 5G Tactile Internet [J]. IEEE Communications Surveys & Tutorials, 2018, 20(4): 3034–3059.), and cannot provide effective support for immersive cross-modal communication services.
[0004] The new generation of multimedia communication has put forward new requirements for data compression coding. Traditionally, data compression coding needs to have a high compression ratio to ensure transmission efficiency and save transmission bandwidth. In addition, it also needs to have 1) personalized characteristics to adapt to the differences in the perception capabilities of different users for different services and ensure service quality; 2) scalability and flexibility to adapt to the time-varying and fluctuating network environment, avoid a sudden drop in the quality of multimedia services, and support multi-path collaborative transmission to achieve low-latency and high-reliability communication; 3) simplicity to avoid introducing additional encoding and decoding computational complexity and latency. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to overcome the deficiencies of existing multimedia data encoding and decoding, and propose a compression encoding and decoding scheme for multimedia streams (such as video and tactile), which adapts to the new requirements of emerging multimedia services and provides support for low-latency and high-reliability multimedia communication.
[0006] Technical Solution: The present invention is a multimedia data encoding and decoding scheme for low-latency and high-reliability communication, including the following steps:
[0007] (1) Design a granularity-adaptive "perceptual dead zone" model based on Weber's Law to remove the imperceptible parts in the data.
[0008] (2) According to the spatio-temporal correlation of the multimedia signal, design a flexible deconstruction scheme for the data to deconstruct a single path of multimedia data into flexible and scalable multiple sub-path signals.
[0009] (3) Send the deconstructed multi-path signals into a spatio-temporal correlation encoder and an entropy encoder respectively to remove the spatio-temporal correlation redundancy and statistical redundancy in each path description, and generate a main description layer and multiple enhancement description layers.
[0010] (4) Assign different description layers to different paths for multi-path collaborative transmission.
[0011] (5) The receiving end receives the data of different layers in a sliding window manner.
[0012] (6) The receiving end performs combined decoding on the received multi-path signals.
[0013] (7) Reconstruct the multimedia information for the user, and obtain the user satisfaction information according to the user's feedback.
[0014] (8) According to the user satisfaction, adjust the encoder parameters to achieve personalized adaptation of the user perception ability and the encoding accuracy.
[0015] Furthermore, the design method of the granularity-adaptive "perceptual dead zone" model in step (1) is as follows:
[0016] Traditional data compression methods based on Weber's law have drawbacks because: ① Different types of multimedia services have different requirements for signal accuracy, and ② The perception abilities of different people are also different. Therefore, it is not possible to encode all types of data with a unified constant. In view of this, the present invention designs a "perceptual dead zone" model with adaptive granularity, which can be mathematically expressed as:
[0017]
[0018] where P is the three-dimensional vector (video or tactile) of the newly input multimedia signal, and P t-1 is the output vector of the encoder at the previous moment. Weber's law is used to determine whether the change amount of the signal is sufficient to be perceived. If so, the output of the encoder at the next moment is updated to P, otherwise it remains P t-1 . In other words, the present invention designs a "perceptual dead zone" model based on Weber's law to determine whether to update the input signal. Its key innovation is that considering the different requirements of different services and populations for signal accuracy, the present invention designs the diameter of the "perceptual dead zone" as a variable (1 + σ)·k that changes with σ, that is, the perceptual granularity is controlled by σ. At the start of encoding, according to the service type (fine-grained service, coarse-grained service, general granularity service), σ is initialized to σ f , σ c or σ m one of them. The update method of σ and the control method of granularity will be described in detail in the subsequent steps.
[0019] Furthermore, the specific steps of the flexible deconstruction scheme of the data in step (2) are as follows:
[0020] (201) Periodically downsample the input multimedia signal (video or tactile) with spatio-temporal correlation properties from the time domain and the space domain.
[0021] Record the original signal as P o , after periodic downsampling, the original signal generates N sub-signals, denoted as P n , 1 ≤ n ≤ N, where is the "pixel" value of the m-th frame with coordinates (x, y) in the n-th signal. And record the total number of frames of each sub-signal as M, and the resolution of each frame as (X, Y), so 1 ≤ m ≤ M, 1 ≤ x ≤ X, 1 ≤ y ≤ Y.
[0022] (202) Extract the relevant components in each sub-signal to avoid redundant repeated encoding and improve the encoding efficiency.
[0023] In order to balance the encoding efficiency and flexibility, the present invention proposes a common information extraction method to extract the relevant components in each sub-signal to generate a main "description", denoted as L MainOne implementation method is to average the corresponding sampled values of each sub-path signal:
[0024]
[0025] Then, this correlation component is removed from the sub-path signal P n to generate N independent sub-descriptions, denoted as L n :
[0026]
[0027] Furthermore, in step (3), the generated one main description and N sub-descriptions (a total of N + 1 descriptions) are respectively sent to the spatio-temporal correlation encoder and the entropy encoder (compatible with HEVC and AVC encoders) to remove the spatio-temporal correlation redundancy and statistical redundancy in each description, generating 1 main description layer and N enhancement description layers.
[0028] Furthermore, in step (4), different description layers are assigned to different paths for transmission for multi-path cooperative transmission. The specific assignment method is: the main description layer has the highest importance, so it is assigned to one or more channels with the best quality for transmission. Other enhancement description layers can be decoded by any combination, so these layers are evenly assigned to other available channels for transmission.
[0029] Furthermore, in step (5), the data of different layers is received in a sliding window manner. Specifically, it includes the following steps:
[0030] (501) The multimedia data is first divided into several groups of slices (SG) with a certain duration in time order.
[0031] (502) When the user requests the multimedia service through interaction with the intelligent device, the user's device sends a request for the corresponding SG number.
[0032] (503) The requester pre-caches as many data slices as possible within the waiting time Δτ, or until all the data slices of the corresponding SG are received.
[0033] (504) The playback of the multimedia data starts, and at the same time, the transmission of the next SG is carried out. The transmission deadline t 0 is set at the starting position of the currently transmitted SG. The playback time t p is tracked in real time by the user client. If t p has not reached t 0 , the transmission is carried out in the order of the main description layer first and then the enhancement description layers.
[0034] (505) To ensure the smoothness of playback, if t p reaches t 0, the SG currently being transmitted needs to be decoded. At this time, two situations will occur: 1) If the current data slice being transmitted belongs to the enhancement description layer, the transmission of the current SG will immediately terminate. Decode all the received data slices, and at the same time start transmitting the next SG to avoid interruption of video playback for the current and the next SG. 2) If the current data slice being transmitted is the main description layer, the playback time has arrived, but the received data is still insufficient for decoding, and interruption of video playback is inevitable.
[0035] (506) When a playback interruption occurs, the most critical issue is to decide whether to continue waiting for the current transmission to complete or skip the current transmission and execute the transmission of the next SG. At this time, the transmission rate of the link and the size of the remaining data can be used to predict the respective waiting times of the two, and the SG with the shorter waiting time is selected for transmission to shorten the duration of the playback interruption.
[0036] Further, in step (6), the receiving end decodes all the data of the current SG that has been received. Specifically, it includes the following steps:
[0037] (601) Decode the main description layer (compatible with HEVC and AVC decoders) and store it in the decoding buffer.
[0038] (602) Refer to the data of the main description layer in the buffer, decode all the received enhancement description layer data, which is the inverse process of encoding, and store it in their respective decoding buffers.
[0039] (603) Fuse all the decoded main description layer and enhancement description layer to improve the resolution of the multimedia data.
[0040] Further, the method for obtaining the user's satisfaction information in step (7) includes the following steps:
[0041] (701) Train a multi-modal satisfaction recognition metric model.
[0042] (70101) Input database where and represent the satisfaction feature vectors composed of touch, video, and audio respectively, and e i represents the corresponding satisfaction label, and e i ∈{1,2,3,4,5}, 1≤i≤l.
[0043] (70102) For all i and j, solve Ω through the following expression h , Ω v , Ω a , α, β, γ, δ.
[0044]
[0045] wherein is the square of the projection distance of the i-th and j-th eigenvectors in the subspace Λ, and Λ = Ω T Ω. α, β, and γ are the optimization parameters of different modal eigenvectors, α, β, γ ≥ 0 and α + β + γ = 1. is for the imposed boundary greater than 1, that is, if the eigenvector Φ i , Φ j do not belong to the same class, then Ξ is a logical operator, if the condition holds, its value is 0, otherwise it is 1.
[0046] (702) Reconstruct and play back the decoded multimedia data to the user, and obtain from the user the eigenvector, and substitute it into the above model to identify the user satisfaction Q. Where u i represents the user receiving the multimedia service.
[0047] Furthermore, step (8) adjusts the encoder parameter σ according to the user satisfaction index to achieve the personalized adaptation of user satisfaction and coding accuracy. Specifically, it includes the following steps:
[0048] (801) Set an expected value Q E of user satisfaction and the update step size Δσ of the encoder parameter σ.
[0049] (802) If Q < Q E and σ > -1, then σ = σ - Δσ.
[0050] (803) Repeat steps 1 to step (802) until Q ≥ Q E .
[0051] Advantageous effects: Compared with the prior art, the significant advantages of the present invention are:
[0052] 1. The present invention adaptively adjusts the encoder parameter through the feedback of the user satisfaction index, and matches the corresponding user perception ability with different coding granularities, thereby realizing personalized coding. In this way, it can not only ensure the user perception level but also avoid the waste of coding bits.
[0053] 2. The present invention realizes the scalability and flexibility of the multimedia data stream by developing the spatio-temporal correlation in the multimedia data and through flexible data deconstruction and reconstruction. The hierarchical encoding and decoding can, on the one hand, flexibly adapt to the changing network environment; on the other hand, it can support multi-path collaborative transmission, thus facilitating the realization of low-latency and high-reliability communication. In particular, each enhancement description layer is not mutually dependent, and decoding does not need to wait for each other, shortening the decoding delay.
[0054] 3. The codec of the present invention is executed in the spatio-temporal dimension without complex mathematical calculations. On the one hand, it ensures simplicity and avoids additional encoding and decoding delays. On the other hand, spatio-temporal related encoding can greatly remove the redundancy of multimedia signals and improve the compression efficiency of multimedia data. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the encoder block diagram of the present invention;
[0056] Figure 2 is the schematic diagram of multi-path collaborative transmission and data reception of the present invention;
[0057] Figure 3 is the decoder block diagram of the present invention (taking three-channel signals as an example);
[0058] Figure 4 is the performance of the codec of the present invention;
[0059] Figure 5 is the comparison of user satisfaction of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In this embodiment, we established a cross-modal communication demonstration platform that includes multi-modal (audio, video, tactile) data acquisition, encoding, transmission, and decoding. The acquisition module of tactile data is realized by using a robotic arm, a robotic dexterous hand, and a pressure sensor. And by combining this module with a high-speed camera and a microphone, we collected 15,000 pairs of multi-modal data sequences. These sequences contain synchronous data of tactile, video, and audio generated by the interaction between the robotic hand and several different materials (including spandex, wooden board, stone, etc.). Then, these data are encoded by the method proposed in the present invention and transmitted to the user through a wireless network. Finally, these data are reconstructed at the user end, and a 3D virtual environment is established for the user to achieve cross-modal remote interaction. The list of devices required for the construction of the demonstration system related to the present invention is shown in Table 1.
[0061] Table 1 List of devices for the demonstration system
[0062]
[0063] As Figure 1 shown in the flowchart, taking the spatio-domain deconstruction and reconstruction of tactile signals as an example, the time-domain deconstruction and reconstruction of tactile signals and the spatio-temporal domain deconstruction and reconstruction of video signals have similar properties and will not be elaborated. A multimedia data encoding and decoding technology for low-latency and high-reliability communication described in this embodiment includes the following steps:
[0064] (1) Collect multimedia data and perform compression encoding on the collected data. In this embodiment, the tactile encoding and decoding in a typical cross-modal communication system including audio, video, and touch are taken as an example.
[0065] Let P be the three-dimensional vector of the newly input tactile signal, and P t-1 be the output vector of the encoder at the previous moment (the initial value is 0). Use Weber's law to determine whether the change amount of the signal is sufficient to be perceived. If so, update the output of the encoder at the next moment to P, otherwise keep it as P t-1 . That is:
[0066]
[0067] The key innovation of the present invention is that considering the different requirements of different services and populations for signal accuracy, the present invention designs the diameter of the "perceptual dead zone" as a variable (1 + σ)·k that changes with σ, that is, the perceptual granularity is controlled by σ. At the start of encoding, according to the service type (fine-grained service, coarse-grained service, general-grained service), σ is initialized to σ f , σ c or σ m one of them. Here, according to the nature of the tactile signal, we set σ f , σ c , σ m to 0, 1, 2 respectively, and k = 0.05.
[0068] (2) According to the spatio-temporal correlation of the tactile signal, flexibly decompose the tactile signal. Endow the tactile data stream with flexible and scalable properties, and at the same time, take into account the compression efficiency of the data.
[0069] (201) Periodically downsample the input tactile signal with spatio-temporal correlation properties from the time domain and the space domain.
[0070] Denote the original tactile signal as P o , after periodic downsampling, the original signal generates 4 sub-signals, denoted as P n , 1 ≤ n ≤ 4, where is the "pixel" value of the m-th frame at coordinates (x, y) in the n-th sub-signal. And denote the total number of frames M of each sub-signal as 600, and the sampling frequency as 2000 Hz.
[0071] (202) Extract the relevant components in each sub-signal, avoid redundant repeated encoding, and improve the encoding efficiency.
[0072] To balance coding efficiency and flexibility, the present invention proposes a common information extraction method, which extracts relevant components from each sub-path signal to generate a main "description". The implementation method adopted in this embodiment is to average the corresponding sampled values of each sub-path signal:
[0073]
[0074] (3) Feed the generated one main description and four sub-descriptions (a total of five descriptions) into the spatio-temporal correlation encoder and entropy encoder (compatible with HEVC and AVC encoders) respectively to remove the spatio-temporal correlation redundancy and statistical redundancy in each description. Generate one main description layer and four tactile enhanced description layers.
[0075] (4) Allocate different description layers to different paths for transmission to perform multi-path collaborative transmission, as Figure 2 shown. The specific allocation method is as follows: The main description layer has the highest importance, so it is allocated to the best-quality one or more channels for transmission. In this embodiment, it is the cellular channel. Other enhanced description layers can be decoded in any combination, so these layers are evenly allocated to other available channels for transmission. In this embodiment, they are D2D channels.
[0076] (5) The receiving end receives data of different layers in a sliding window manner. As Figure 2 shown, it specifically includes the following steps:
[0077] (501) Divide the tactile data into several groups of slices (SGs) with a certain duration in time order. Each SG lasts for 10 ms.
[0078] (502) When the user requests the multimedia service through interaction with the smart device, the user's device sends a request for the corresponding SG number.
[0079] (503) The requester pre-caches as many data slices as possible within the waiting time Δτ = 10 ms, or until all data slices of the corresponding SG are received.
[0080] (504) The playback of the tactile data starts, and at the same time, the transmission of the next SG is carried out. The transmission deadline t 0 is set at the starting position of the currently transmitted SG. The playback time t p is tracked in real time by the user client. If t p has not reached t 0 , the transmission is carried out in the order of the main description layer first and then the enhanced description layers.
[0081] (505) To ensure the smoothness of playback, if t p reaches t 0, the SG currently being transmitted is decoded. At this time, two situations may occur: ① If the current data slice being transmitted belongs to the enhanced description layer, the transmission of the current SG will be immediately terminated. All received data slices will be decoded, and the transmission of the next SG will be started simultaneously to avoid interruption of video playback for the current and the next SG. ② If the current data slice being transmitted is the main description layer, the playback time has arrived, but the received data is still insufficient for decoding, and interruption of video playback is inevitable.
[0082] (506) When a playback interruption occurs, the transmission rate of the link and the size of the remaining data are used to predict their respective waiting times, and the SG with the shorter waiting time is selected for transmission to shorten the duration of the playback interruption.
[0083] (6) The receiving end decodes all the data of the currently received SG. As Figure 3 shown, it specifically includes the following steps:
[0084] (601) Decode the main description layer (compatible with HEVC and AVC decoders) and store it in the decoding buffer.
[0085] (602) Referring to the data of the main description layer in the buffer, decode all received enhanced description layer data. Decoding is the inverse process of encoding and is stored in their respective decoding buffers. If the data of a certain / some description layers is not completely received, the data of that layer will be abandoned for decoding.
[0086] (603) Fuse all decoded main description layers and enhanced description layers to improve the haptic reconstruction resolution.
[0087] (7) Reconstruct the haptics for the user and obtain the user's satisfaction information. It specifically includes the following steps:
[0088] (701) Pre-train the multi-modal satisfaction recognition metric model.
[0089] (70101) Input the database where and respectively represent the satisfaction feature vectors composed of haptics, video, and audio, e i represents the corresponding satisfaction label, e i ∈{1,2,3,4,5}, 1≤i≤l. This database is composed of the user's facial expression information, voice information, and haptic information collected by a video camera, a microphone, and a haptic device. The satisfaction label is given by the user's score. The pre-training process of the model is executed before cross-modal communication.
[0090] (70102) For all i and j, solve Ω through the following formula h , Ω v , Ω a, α, β, γ, δ.
[0091]
[0092] Wherein, is the square of the projection distance of the i-th and j-th eigenvectors in the subspace Λ, and Λ = Ω T Ω. α, β, γ are the optimization parameters of different modal eigenvectors, α, β, γ ≥ 0 and α + β + γ = 1. is a boundary greater than 1 imposed on , that is, if Φ i , Φ j do not belong to the same class, then Ξ is a logical operator, if the condition holds, its value is 0, otherwise it is 1.
[0093] Since Ω h , Ω v , Ω a , the values of α, β, γ, δ will vary significantly according to the application scenario and the signal type, and these parameters are intermediate parameters for model training - application and have no reference value for a specific application. Therefore, the program does not output the actual values.
[0094] (702) Reconstruct the touch for the user through the tactile device and obtain the eigenvector from the user and bring it into the above model to identify the user satisfaction Q. It is also obtained by collecting the user's information in an imperceptible manner by video, audio, and tactile devices.
[0095] (8) According to the user's satisfaction index, adjust the encoder parameter σ to achieve personalized adaptation of user satisfaction and coding accuracy. Specifically, it includes the following steps:
[0096] (801) Set the expected value Q of user satisfaction E = 4.5, and the update step size Δσ of the encoder parameter σ = 0.5.
[0097] (802) If Q < Q E and σ > -1, then σ = σ - Δσ.
[0098] (803) Repeat steps 1 to step (8 - 2) until Q ≥ Q E .
[0099] The experimental method of this embodiment will be further described below.
[0100] In this embodiment, the performance metrics for evaluating the multimedia data encoding and decoding scheme proposed by the present invention are divided into two categories: accuracy metrics and non-accuracy metrics. The accuracy metrics include the mean squared error (MSE) and the bit rate, and the non-accuracy metric is the average satisfaction level.
[0101] Mean Squared Error (MSE): It refers to the average of the sum of the squares of the differences between the original signal and the reconstructed signal after encoding and decoding, which reflects the distortion degree of the codec. The larger the MSE value, the greater the encoding and decoding distortion, and vice versa.
[0102] Bit Rate: It refers to the amount of bandwidth occupied by the data. Furthermore, the ratio of the bit rate of the encoded data to the original data is the compression ratio. The lower the compression ratio, the stronger the compression ability of the encoder, and vice versa.
[0103] Satisfaction Level: It refers to the evaluation of the service quality by users when participating in multimedia services. It is divided into five levels: very satisfied (5 points), relatively satisfied (4 points), generally satisfied (3 points), dissatisfied (2 points), and extremely dissatisfied (1 point).
[0104] The experimental results are analyzed below in conjunction with the drawings and tables.
[0105] In this embodiment, we recorded the haptic signal encoding results with σ = -1 to 2 and Δσ = 0.5. Figure 4 Some representative experimental results are shown. Among them, when σ = -1, (1 + σ)·k = 0, that is, no haptic sampling values are discarded. Therefore, Figure 4 (a) shows the original signal. At this time, the encoding and decoding distortion MSE = 0, and the encoding bit rate is equal to the sampling frequency. From Figure 4 (a) to Figure 4 (e), it can be seen that as σ increases, more and more haptic details are discarded. Therefore, the encoding and decoding distortion MSE increases accordingly, and the encoding bit rate decreases, while the compression efficiency increases. Figure 4 (f) shows the negative correlation between MSE and the encoding bit rate. It should be noted that as σ increases, the high-value details of the haptic signal are first discarded, which is consistent with the theoretical analysis of Weber's law.
[0106] Figure 5Shows the user satisfaction comparison experiment of this embodiment. The experiment compares the solution in the present invention with the traditional HEVC encoding and SHVC encoding schemes under multi-path cooperation and non-multi-path cooperation methods respectively. As shown in the figure, when multi-path cooperative transmission is not used, the user satisfaction of the solution proposed in the present invention is basically the same as that of the SHVC scheme. The reason is that under poor network conditions, the decoding lower limits of this solution and the SHVC scheme are the same, that is, the base description layer is received. The reason for the lower user satisfaction of HEVC is that it does not introduce a scalable solution, has poor adaptability to the network, and is prone to playback interruption. When multi-path cooperative transmission is used, this solution can achieve the best user satisfaction fastest because it is the most flexible and adaptable while taking into account the encoding efficiency.
Claims
1. A multimedia data encoding and decoding method for low-latency and high-reliability communication, characterized in that, It includes the following steps: (1) Remove the imperceptible part of the data; (2) Decompose the multimedia data of one path into several flexible and scalable sub-path signals; (3) Send the decomposed multi-path signals into the spatio-temporal correlation encoder and the entropy encoder respectively, remove the spatio-temporal correlation redundancy and statistical redundancy in each path description, and generate a main description layer and several enhancement description layers; (4) According to the different importance of the main description layer and the enhancement description layers, allocate different description layers to different transmission paths for multi-path cooperative transmission; (5) The receiving end receives the data of different layers in the form of a sliding window; (6) The receiving end performs combined decoding on the received multi-path signals; (7) Reconstruct the multimedia information for the user, and obtain the user satisfaction information according to the user's feedback; (8) Adjust the encoder parameters according to the user's satisfaction to achieve personalized adaptation of the user perception ability and the coding accuracy; In step (1), the "perceptual dead zone" model with adaptive granularity is used to remove the imperceptible part of the data, which is expressed mathematically as: Among them, P is the three-dimensional vector of the newly input multimedia signal, P t-1 is the output vector of the encoder at the previous moment; Weber's law is used to determine whether the change amount of the signal is sufficient to be perceived. If so, the output of the encoder at the next moment is updated to P, otherwise it remains P t-1 . The diameter of the "perceptual dead zone" is designed as a variable (1 + σ)·k that changes with σ, where k is the traditional Weber fraction; the above formula sets the perceptual granularity as a variable controlled by σ. At the start of encoding, according to the service type, fine-grained service, coarse-grained service, general granularity service, σ is initialized to σ f , σ c or σ m one of them; σ adaptively updates and controls the encoding granularity according to the user's experience and service requirements; Step (2) adopts a flexible decomposition method of the data, which specifically includes: (201) Periodically downsample the input multimedia signal with spatio-temporal correlation properties, video or tactile, from the time domain and the space domain; Denote the original signal as P o , after periodic downsampling, the original signal generates N sub-signals, denoted as P n , 1 ≤ n ≤ N, where is the "pixel" value at the coordinate (x, y) of the m-th frame in the n-th signal; and denote the total number of frames of each sub-signal as M, and the resolution of each frame as (X, Y), so 1 ≤ m ≤ M, 1 ≤ x ≤ X, 1 ≤ y ≤ Y; (202)Extract the relevant components from each sub-path signal, and generate a main "description" by extracting the relevant components from each sub-path signal, denoted as L Main ; One implementation method is to average the corresponding sampled values of each sub-path signal: Then, relevant components are removed from the Zilu signal P n to generate N mutually independent sub-descriptions, denoted as L n :
2. The multimedia data encoding and decoding method for low-latency and high-reliability communication according to claim 1, characterized in that The specific steps of step (5) are: (501) The multimedia data is first divided into several slice groups SG of a certain duration in time order; (502) When the user requests the multimedia service through interaction with the intelligent device, the user's device sends a request for the corresponding SG number; (503) The requester pre-caches as many data slices as possible within the waiting time Δτ, or until all data slices of the corresponding SG are received; (504) Multimedia data playback starts, and at the same time, the transmission of the next SG is carried out; the transmission deadline is t 0 Set at the starting position of the currently transmitted SG; playback time is t p Tracked in real time by the user client. If t p has not reached t 0 , then the transmission is carried out in the order of the main description layer first and then the enhancement description layer; (505) If t p reaches t 0 , then decode the currently transmitted SG; At this time: If the currently transmitted data slice belongs to the enhancement description layer, the transmission of the current SG is immediately terminated; and all received data slices are decoded, and at the same time, the transmission of the next SG is started; If the currently transmitted data slice is the main description layer, but the playback time has arrived and the received data is still insufficient for decoding, the video playback interruption is inevitable; (506) When a playback interruption occurs, use the transmission rate of the link and the size of the remaining data to predict their respective waiting times, and select the SG with the shorter waiting time for transmission to shorten the playback interruption duration.
3. The multimedia data encoding and decoding method for low-latency and high-reliability communication according to claim 1, wherein, The specific steps of step (6) are: (601) Decode the main description layer to be compatible with HEVC and AVC decoders, and store them in the decoding buffer; (602) Refer to the data of the main description layer in the reference buffer, decode all received enhancement description layer data, which is the inverse process of encoding, and store them in their respective decoding buffers; (603) Integrate all decoded main description layers and enhancement description layers to improve the resolution of the multimedia data.
4. The multimedia data encoding and decoding method for low-latency and high-reliability communication according to claim 1, characterized in that, The specific steps of step (7) are: (701) Train a multi-modal satisfaction recognition metric model; (70101) Input database where and represent the satisfaction feature vectors composed of touch, video, and audio respectively, and e i represents the corresponding satisfaction label, and e i ∈{1,2,3,4,5}, 1≤i≤l; (70102) For all i and j, solve for Ω through the following expressions h , Ω v , Ω a , α, β, γ, δ; Among them, is the square of the projection distance of the i-th and j-th feature vectors in the subspace Λ, where Λ = Ω T Ω; α, β, γ are optimization parameters of different modal feature vectors, α, β, γ ≥ 0 and α + β + γ = 1; is for the imposed boundary greater than 1, that is, if the feature vector Φ i , Φ j do not belong to the same class, then Ξ is a logical operator, if the condition holds, its value is 0, otherwise 1; (702) Reconstruct and play back the decoded multimedia data to the user, and obtain from the user feature vectors, and input them into the above model to identify the user satisfaction Q; where u i represents the user who receives the multimedia service.
5. The multimedia data encoding and decoding method for low-latency and high-reliability communication according to claim 1, characterized in that, Step (8) specifically includes the following steps: (801) Set an expected value Q of user satisfaction E , and the update step size Δσ of the encoder parameter σ; (802) If Q < Q E and σ > -1, then σ = σ - Δσ; (803) Repeat steps 1 to (802) until Q ≥ Q E .