A method and system for rate-delay scalable coding of vibrotactile signals
By using a non-stationary autoencoder model and a non-uniform quantization method, the contradiction between perception quality and real-time performance in existing vibration tactile signal encoding methods is resolved. A coding scheme with scalable bit rate and latency is provided, enabling efficient, real-time, and high-fidelity tactile feedback in multimedia services.
Patent Information
- Application Number
- CN202511233847.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-01
AI Technical Summary
Existing vibration tactile signal encoding methods struggle to achieve efficient compression while maintaining perception quality and meeting real-time requirements, and they do not offer options for scalable latency and scalable bit rate.
A non-stationary autoencoder model based on finite scalar quantization and attention is used to preprocess the triaxial vibration tactile signal, convert it into a one-dimensional frame using the DFT321 algorithm, compress the residual using non-uniform quantization, and generate the output bitstream by combining entropy coding, providing scalability in bit rate and latency.
It enables real-time, high-fidelity haptic feedback that adapts to different latency switching requirements and time-varying network bandwidth in multimedia services, enhancing the user's immersive experience.
Smart Images

Figure CN120729333B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vibration haptic signal coding, in particular to a code rate delay scalable coding method and system for vibration haptic signals. BACKGROUND
[0002] With the development of tactile internet (TI), multimedia communication and sensor hardware, the trend of integrating human senses into traditional audiovisual services is increasingly evident. As an important mode of human senses, haptics can obtain useful information through physical interaction, including attributes such as hardness, texture, temperature and pressure; haptic feedback helps to quickly distinguish visually similar but textured objects (such as stone or plaster), thereby enhancing the user's sense of reality and providing an immersive experience. Therefore, haptic feedback technology has great potential in various services such as augmented reality (AR), virtual reality (VR), remote medical care, multimedia education and robot remote operation. Currently, haptic feedback has attracted widespread attention, and related research covers various fields from signal acquisition to processing and compression. In order to create an immersive user experience, haptic coding must emphasize two key factors: high-fidelity feedback and strict delay requirements.
[0003] Specifically, on the one hand, the sampling frequency of haptic signals is usually 2.8 kHz, and the intensity is represented by 16-bit data, so the uncompressed haptic stream of a single interaction point requires a bit rate of 44.8 kbps; in the case of supporting 5G technology, the bandwidth occupied by a single interaction point can be negligible compared to video streams. However, achieving high-fidelity haptic feedback also requires multiple interaction points to be set on the surface, and previous research on haptic gloves has deployed more than 46 interaction points, resulting in a total bit rate of about 2.0 Mbps. Due to the bandwidth competition between video streams and haptic streams, directly transmitting uncompressed haptic streams can damage video streams and reduce user perception quality, so this problem cannot be ignored. Therefore, developing a haptic coding method that does not affect human perception is crucial for providing high-fidelity feedback.
[0004] On the other hand, existing multimedia services can be divided into three types according to delay sensitivity: delay-intensive, delay-sensitive and delay-critical. Delay-intensive services are mainly video-based, with haptic feedback playing a supporting role in this case. Users passively perceive haptic feedback and tolerate a delay of about 45 milliseconds. Delay-sensitive services attach equal importance to video and haptic feedback, especially in XR / AR services. These services require more stringent delay requirements than passive interaction, usually between 10 and 30 milliseconds. Finally, delay-critical services are mainly haptic feedback, such as remote surgery, which involves human life and property. The delay must be strictly controlled below 10 milliseconds, and this delay requirement is crucial to haptic coding, as it directly affects user experience and the practicality of services.
[0005] The prior research develops various haptic coding methods to reduce the bit rate of haptic stream, and the prior art has VPC-DS method, also has PVC-SLP, VC-PWQ and RNVC method, VPC-DS adopts SPIHT lossless compression algorithm to realize a high compression ratio of vibration haptic signal coding method, to compress the real-time haptic stream to low bit rate (about 3.5 kbps), but it will cause unacceptable perceptual quality distortion; PVC-SLP adopts sparse linear prediction for signal synthesis, and cooperates with the application of DCT transform to residual signal, and further compresses the DCT coefficient by using acceleration sensitivity function, but it needs to buffer 200 samples (about 71.4 ms) to obtain the assumed short-term stationarity, so as to realize efficient and perceptual lossless compression, which conflicts with the strict low delay requirement; VC-PWQ dynamically extracts key signal features through perceptual absolute threshold, and utilizes DWT to efficiently decouple input correlation, and under the premise of maintaining subjective perceptual quality, the haptic bit stream is compressed, but its calculation complexity is high, and the real-time performance is limited; further, RNVC adopts gated recurrent unit (GRU) to predict subsequent haptic signals, which meets the real-time requirement and exhibits excellent compression performance, but the non-stationarity of vibration haptic signals is not fully considered.
[0006] The above-mentioned existing haptic coding method also has the following defects: due to the non-stationarity of the haptic signal, it is difficult for these methods to realize efficient compression while maintaining perceptual quality and meeting real-time requirements; none of the haptic coding methods provides options for delay scalability and code rate scalability. SUMMARY
[0007] To solve the problems mentioned in the background art, the purpose of the present application is to provide a code rate delay scalable coding method and system for vibration haptic signals.
[0008] In a first aspect, the purpose of the present application can be realized by the following technical scheme: a code rate delay scalable coding method for vibration haptic signals, the method comprising the following steps:
[0009] Receiving a three-axis vibration haptic signal, pre-processing the three-axis vibration haptic signal to obtain a pre-processed three-axis vibration haptic signal, inputting the pre-processed three-axis vibration haptic signal into a pre-established non-stationary autoencoder model based on finite scalar quantization and attention, and outputting the normalized reconstruction frame and the difference between the normalized frame and the FSQ code word;
[0010] Based on the difference between the normalized reconstruction frame and the normalized frame, the residual is calculated, the non-uniform quantization method is used, the residual is quantized according to the quantization bit width, the residual quantization index is generated, the FSQ code word and the residual quantization index are losslessly compressed by using entropy coding, and the output bit stream is generated.
[0011] With reference to the first aspect, in some implementations of the first aspect, the method further includes that the process of pre-processing the tri-axial vibration haptic signal includes:
[0012] The tri-axial vibration haptic signal is buffered in a buffer, and then segmented by a sliding window. According to the control of the overlap rate, the signal is divided into overlapping or non-overlapping frames. The DFT321 algorithm is applied to convert the haptic frame in the three-dimensional vibration haptic signal into a one-dimensional frame, to obtain the pre-processed tri-axial vibration haptic signal;
[0013] According to the DFT321 algorithm, the frequency spectrum and the phase θ are expressed as follows:
[0014]
[0015] where f represents the frequency, A1 represents the discrete Fourier transform of the X-axis of the original haptic frame, A2 represents the discrete Fourier transform of the Y-axis of the original haptic frame, and A3 represents the discrete Fourier transform of the Z-axis of the original haptic frame. The one-dimensional frame x can be derived by the preset frequency spectrum and phase, and the expression is as follows:
[0016]
[0017] where iDFT(·) represents the inverse discrete Fourier transform, and real(·) is used to extract the real part from the complex value, represents the frequency spectrum, θ represents the phase; after the DFT321 algorithm, the three-dimensional vibration haptic frame is converted into a one-dimensional frame x.
[0018] With reference to the first aspect, in some implementations of the first aspect, the method further includes that the process of inputting the pre-processed tri-axial vibration haptic signal into the pre-established non-stationary autoencoder model based on finite scalar quantization and attention includes:
[0019] The pre-processed one-dimensional frame x is initially normalized to generate a normalized frame x', and then the attention-based non-stationary autoencoder model NAE compresses x' into a latent compact haptic representation z The FSQ layer further quantizes z into a codebook C as an FSQ code word. The core parameters including the mean µ x , the standard deviation σ x and the code word C are transmitted to the receiving end through a wireless communication network.
[0020] The receiving end first restores the code word C into the compressed representation Normalized reconstructed frames are obtained through the NAE decoder. Then, the reconstructed frames are output through inverse normalization processing. In the end-to-end training process, an additional multi-scale discriminator loss function was constructed.
[0021] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the process of initially normalizing the preprocessed one-dimensional frame x:
[0022]
[0023] In the formula, x = [x1, x2, ..., x... L ] T Represents a sample of a one-dimensional frame. µ x This represents the mean. σ x Let x' represent the standard deviation, L represent the frame length, and x' represent the normalized frame, where x i Let i represent the sample at position i, for i = 1, 2, ..., L;
[0024] The potential compact tactile representation FSQ will z Based on quantization level l = [ l 1, l 2, ..., l d Quantized into codewords C = [C1, C2, ..., C Lc [This converts codewords into a compressed representation.] L c Indicates codeword length, d represents the length of the quantization level, and is a compressed scalar. Obtain it in the following ways:
[0025]
[0026] in, This indicates a round-down operation. Indicates the first j One quantification level, A compact tactile representation of the position (i, j).
[0027] A pass-through estimator is used to pass gradients through rounding operations:
[0028]
[0029] Where sg(·) represents forcing the gradient to zero, and round(·) represents rounding. Indicates the first jquantization level, representing (a) i , j ) compact haptic representation, compression scalar scaled to the range (-1, 1), i.e. ; the haptic representation is then quantized to a positive integer codeword in C :
[0030]
[0031] where l 0 = 1, l t denotes the t quantization level, denotes the j quantization level, is the compression scalar, the compressed haptic representation can be obtained by the codeword C in the following way:
[0032]
[0033] where denotes a cyclic modulo operation with modulus , denotes a floor operation, C i denotes the i-th positive integer codeword in C, l t denotes the t quantization level, denotes the j quantization level.
[0034] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the pre-established training of the non-stationary autoencoder model based on limited scalar quantization and attention is as follows:
[0035] A multi-scale discriminator is adopted to improve the quality of the reconstructed haptic signal and the generalization performance of the non-stationary autoencoder model FSQ-NAE based on limited scalar quantization and attention, a loss function is designed to optimize the training, and the multi-scale discriminator and the loss function of the FSQ-NAE model are defined as follows:
[0036] Multi-scale discriminator loss: the training target of the multi-scale discriminator is to minimize the following hinge adversarial loss function:
[0037]
[0038] where K denotes the number of discriminator models, x' denotes the normalized frame, denotes the normalized reconstructed frame, D k (·) denotes the output of the k-th discriminator at the last layer;
[0039] FSQ-NAE model loss: The training objective of the FSQ-NAE model includes mean squared error loss, structural similarity loss, adversarial loss, and feature loss:
[0040] Mean squared error loss Ensuring the reconstructed haptic frame has excellent fidelity at the sample level, i.e.,
[0041]
[0042] where x denotes the input one-dimensional frame, denotes the output reconstructed frame, and L denotes the frame length.
[0043] Structural similarity loss Ensuring the reconstructed haptic frame maintains structural similarity, i.e.,
[0044]
[0045] where ssim(·) denotes one-dimensional SSIM calculation between two haptic frames, x denotes the input one-dimensional frame, denotes the output reconstructed frame; Adversarial loss Aims to overcome the discriminative effect of the discriminator, i.e.,
[0046]
[0047] where denotes the normalized reconstructed frame, K denotes the number of discriminator models, D k (k) denotes the output of the kth discriminator at the last layer; Feature loss is calculated by averaging the L1 distance of different inputs at the internal layer output of the discriminator, i.e.,
[0048]
[0049] where K denotes the number of discriminator models, M denotes the number of layers of each discriminator model, x' denotes the normalized frame, denotes the normalized reconstructed frame, D m k (k) denotes the output of the kth discriminator at the mth layer; Overall FSQ-NAE model loss L G is the weighted sum of the four different components:
[0050]
[0051] where denotes the mean squared error loss, denotes a structural similarity loss, denotes an adversarial loss, denotes a feature loss, λ mse denotes a weight of the mean square error loss in the model loss, λ ssim denotes a weight of the structural similarity loss in the model loss, λ adv denotes a weight of the adversarial loss in the model loss, λ feat denotes a weight of the feature loss in the model loss;
[0052] With reference to the first aspect, in some implementations of the first aspect, the method further comprises that a process of a non-stationary attention mechanism of the FSQ-NAE model is as follows:
[0053] Starting from a self-attention mechanism, an original non-stationary haptic frame x is received as an input:
[0054]
[0055] wherein Q, K, V represent Query, Key, Value respectively, wherein the lengths of Q, K, V are L, and the dimensions of Q, K, V are d k , Attn(·) represents an attention mechanism, Softmax(·) represents a normalized exponential function, an embedding layer is represented as F(·), it is assumed that F has a linear property and is applied to each input sample, therefore, Q = [q1, q2,…q L ] T ] i ] i ]
[0056] Using the generated normalized haptic frame x' as an input, Q' = [F (x1'), F (x2'),…, F (x L ') T Based on the linear property, it is derived that wherein is an average value of Q in a time dimension, σ x denotes a standard deviation;
[0057] Based on a connection between the attention weight of the original haptic frame and the normalized frame:
[0058]
[0059] wherein Q and K represent Query and Key with lengths of L and dimensions of d k , and V represents Value with a length of L and a dimension of d µQ denotes the average of Q over the time dimension, µ K denotes the average of K over the time dimension, σ x denotes the standard deviation, , , , , respectively, summing over each column and each element in the matrix , and performing a Softmax(·) operation on both sides of the equation, where Softmax(·) denotes a normalized exponential function, to obtain the direct connection between the attention weights:
[0060]
[0061] where d k denotes the dimensionality, and the approximation is defined as a positive scale of the non-stationary factor τ = σ x 2 and the shift vector , the non-stationary factor τ and are learned directly from the original haptic frame x using a multi-layer perceptron (MLP) layer µ x and σ x ; the non-stationary attention is calculated as follows:
[0062]
[0063] where , µ V denotes the average of V over the time dimension; the output of the non-stationary attention block is mapped to a latent haptic representation of size by setting the convolution kernel and stride of a one-dimensional convolution block z .
[0064] In combination with the first aspect, in some implementations of the first aspect, the method further includes: in the process of employing the non-uniform quantization method, quantizing the residual according to the quantization bit width to generate a residual quantization index, including:
[0065] clipping all samples in the residual to make the residual fall within the interval [-1, 1], and assigning a 1-bit sign code P c to represent the sign of e:
[0066]
[0067] After taking the absolute value, all samples are constrained in the range of [0, 1];
[0068] The allocated quantization code P v to represent the non-uniform quantization result with a preset bit width:
[0069] The sign code P c and the quantization code P v form a residual quantization index I.
[0070] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the allocated quantization code P v to represent the non-uniform quantization result with a preset bit width, according to the parity of the quantization bit width, there are two different cases of the quantization code, including:
[0071] When P v represented by 2n bits, the quantized residual R all values are as follows:
[0072]
[0073] wherein and are scaling factors to ensure that the maximum value in R is 1;
[0074] When P v represented by 2n+1 bits, all values of the quantized residual R can be obtained in the following way:
[0075]
[0076] wherein , .
[0077] The second aspect, in order to achieve the above object, the present application discloses a kind of code rate delay scalable coding system for vibration haptic signal, comprising:
[0078] Signal processing module, for receiving three-axis vibration haptic signal, pre-processes three-axis vibration haptic signal, obtains pre-processed three-axis vibration haptic signal, inputs pre-processed three-axis vibration haptic signal into pre-established non-stationary self-encoder model based on finite scalar quantization and attention, output obtains the difference between normalized reconstruction frame and normalized frame and FSQ code word;
[0079] The compression encoding module is configured to calculate a residual based on a difference between the normalized reconstructed frame and the normalized frame, quantize the residual by using a non-uniform quantization method according to a quantization bit width, generate a residual quantization index, and perform lossless compression on the FSQ code word and the residual quantization index by using entropy coding to generate an output bit stream.
[0080] In another aspect of the present application, in order to achieve the above-mentioned object, a terminal device is disclosed, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, a code rate and delay scalable coding method for vibration haptic signals is used.
[0081] The present application has the following beneficial effects:
[0082] The present application designs a non-stationary autoencoder framework to achieve efficient compression of haptic signals while ensuring perceptual quality; by using the sliding window method, the overlap coding is used to reduce the buffer delay, and by adjusting the overlap ratio, various delay options are provided to adapt to different requirements; by using a non-uniform quantization method to compress the residual and dynamically enhance the fidelity, the method provides multiple bit rate levels by adjusting the quantization bit width to adapt to different network bandwidths. The present application well realizes a vibration haptic signal coding method which provides code rate and delay scalable options, solves the challenge that the existing vibration haptic signal coding method cannot simultaneously adapt to various delay switching requirements in multimedia services and time-varying network bandwidth, provides real-time and high-fidelity haptic feedback, and enhances the immersion of users in multimedia services. BRIEF DESCRIPTION OF DRAWINGS
[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description, and obviously, other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings;
[0084] Figure 1 is a method flowchart of the present application;
[0085] Figure 2 is a complete network structure diagram of the present application;
[0086] Figure 3 is a model architecture diagram of the present application;
[0087] Figure 4 is an SNR comparison result diagram of the present application and other methods;
[0088] Figure 5This is a comparison chart of the PSNR results of this invention with other methods;
[0089] Figure 6 This is a comparison diagram of the ST-SIM results between the present invention and other methods;
[0090] Figure 7 This is a flowchart of the workflow of the present invention;
[0091] Figure 8 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0092] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0093] Example 1:
[0094] like Figure 1 As shown, a rate-delay scalable coding method for vibration tactile signals includes the following steps:
[0095] S101: Receives triaxial vibration tactile signals, preprocesses the triaxial vibration tactile signals to obtain preprocessed triaxial vibration tactile signals, inputs the preprocessed triaxial vibration tactile signals into a pre-established non-stationary autoencoder model based on finite scalar quantization and attention, and outputs the difference between the normalized reconstructed frame and the normalized frame, as well as the FSQ codeword.
[0096] The preprocessing of triaxial vibration tactile signals includes:
[0097] Step 1-1: First, the triaxial vibration tactile signal is buffered in a buffer, and then segmented using a sliding window. Figure 2 In Phase 1 of the process, the signal is divided into overlapping or non-overlapping frames based on the overlap rate to select the desired delay level. Then, the DFT321 algorithm is applied to convert the tactile frames within the three-dimensional vibration tactile signal into one-dimensional frames without causing a significant degradation in perceived quality.
[0098] First, the triaxial vibration tactile signals are buffered. Then, a sliding window method is used for frame segmentation to achieve scalable latency. The window length is fixed at 64 samples (approximately 22.9 ms). The overlap ratio 0 controls the degree of overlap between adjacent sliding windows to manage buffer latency in a scalable manner.
[0099] In the present application, the selection of the overlap ratio O must take into account the coding efficiency and the delay requirement. O can be set to 0, 1 / 4, 1 / 2 or 7 / 8. When O is set to 0, the encoding is carried out in a non-overlapping manner to meet the basic delay requirement (such as delay-intensive services) and provide efficient encoding. Increasing the overlap ratio O (such as 1 / 4 and 1 / 2) can reduce the buffer delay, meet the delay requirement of delay-sensitive services, and enhance user immersion, especially setting the overlap ratio O to 7 / 8, the buffer delay is about 2.9 ms, which can meet the delay requirement of delay-critical services, and the frame after sliding window segmentation will be used in the next DFT321 algorithm stage;
[0100] Step 1-2, secondly, due to the insensitivity of human hands to the directionality of tactile sensation, the original three-dimensional signal is converted into a one-dimensional signal without causing significant perceptual degradation. Therefore, the DFT321 algorithm is performed on the frame after sliding window segmentation, as follows:
[0101] To ensure perceptual quality, both spectral matching and time-domain matching must be met. Spectral matching ensures that the energy of the one-dimensional frame in each frequency band matches the combined energy of the original frame; time-domain matching aligns the peak value and mutation characteristics of the one-dimensional frame with the corresponding characteristics in the original frame. According to the DFT321 algorithm, to meet the above matching requirements, the frequency spectrum and phase θ expressions are as follows:
[0102]
[0103] where f represents the frequency, A1 represents the discrete Fourier transform of the X-axis of the original tactile frame, A2 represents the discrete Fourier transform of the Y-axis of the original tactile frame, and A3 represents the discrete Fourier transform of the Z-axis of the original tactile frame. The one-dimensional frame x can be derived by the pre-set frequency spectrum and phase, as follows:
[0104]
[0105] where iDFT(·) represents the inverse discrete Fourier transform, real(·) is used to extract the real part from the complex value, represents the frequency spectrum, θ represents the phase; after the DFT321 algorithm, the three-dimensional vibration tactile frame is converted into a one-dimensional frame x.
[0106] The pre-processed vibration tactile signal is used as the input of the FSQ-NAE model in the next stage.
[0107] A FSQ-NAE model is designed, which is the stage 2 part of Figure 2 The detailed structure diagram of the model is shown in Figure 3
[0108] The pre-processed one-dimensional frame x is initially normalized to generate a normalized frame x'. The NAE encoder then compresses x' into a latent compact haptic representation z , the FSQ layer further quantizes z the latent representation into a finite codebook C, the FSQ codebook C will be used in the bitstream generation of step (4). The core parameters (mean µ x , standard deviation σ x and codebook C) are transmitted to the receiving end through a wireless communication network;
[0109] The receiving end first restores the codebook C to the compressed representation , obtains the normalized reconstructed frame through the NAE decoder (mirror-symmetry with the encoder), and outputs the reconstructed frame through the inverse normalization process. In the end-to-end training process, a multi-scale discriminator is used to improve the quality of the reconstructed haptic signal and the generalization performance of the non-stationary autoencoder model FSQ-NAE based on limited scalar quantization and attention, that is, a multi-scale discriminator loss function is additionally constructed in the training process.
[0110] The specific implementation steps are as follows:
[0111] Step 2-1, after preprocessing, the one-dimensional frame x is modeled. Due to the inherent non-stationarity of the haptic signal, a simple and effective normalization method is used to eliminate the non-stationarity within each input haptic frame; the mean µ x shift and the standard deviation σ x scaling operation are used to transform it. The normalized frame
[0112] x' can be obtained by the following expression:
[0113]
[0114] where x = [x1, x2,…, x L ] T represents the samples of the one-dimensional frame, µ x represents the mean, σ x represents the standard deviation, L represents the frame length, and x' represents the normalized frame, where x i represents the sample at position i, for i = 1,2,…,L. Note that through the normalization module, the autoencoder model receives a more stable input distribution, making it suitable for subsequent signal modeling;
[0115] Step 2-2, In the non-stationary attention module, the present application aims to reintegrate the non-stationary characteristics of the original haptic frame into the attention mechanism to avoid the problem of over-smoothing. First, the relationship between the attention weights of the original frame and the normalized frame is explored, and then the reintegration of non-stationary characteristics is achieved through compensation weights, as follows:
[0116] Step 221, First, starting from the self-attention mechanism, which receives the original non-stationary haptic frame x as input:
[0117]
[0118] where Q, K, V represent Query, Key, Value with length L and dimension d k . Attn(·) represents the attention mechanism, Softmax(·) represents the normalized exponential function, and the embedding layer is represented as F(·), which is responsible for mapping the input frame to the feature space. To simplify the analysis, it is assumed that F has linear properties and is applied to each input sample. Therefore, each query in Q = [q1, q2, … q L ] T can be calculated as q i = F(x i );
[0119] Step 222, second, using the generated normalized haptic frame x' as input,
[0120] Q' = [F(x1'), F(x2'), …, F(x L ')] T Based on the assumption of linear properties, it can be deduced that where is the average value of Q in the time dimension, σ x and
[0121] Step 223, then, the present application proposes a relationship between the attention weights of the original haptic frame and the normalized frame:
[0122]
[0123] where Q and K represent Query and Key with length L and dimension d k , µ Q is the average value of Q in the time dimension, µ K is the average value of K in the time dimension, σ x is the standard deviation, , , , , respectively Summing is performed on each column and each element within the formula. A Softmax(·) operation is then applied to both sides of the formula. Softmax(·) represents the normalized exponential function, and due to its translation invariance, a direct relationship between the attention weights can be obtained:
[0124]
[0125] Where d k Indicates dimension;
[0126] Step 224: Finally, the key lies in approximating the positive scaling scalar, which is defined as a non-stationary factor. τ = σ x 2 and shift vector Due to the challenge of the function F strictly preserving the linear property, scalars τ sum vector It is difficult to use directly to compensate for non-stationary features, and obtaining K and µ Q This requires significant cost. To address this issue, a simple yet effective multilayer perceptron (MLP) layer is used to directly learn the non-stationary factors from the original tactile frame x. τ and and its statistics µ x and σ x The calculation of non-stationary attention is as follows:
[0127]
[0128] in , µ V Let V represent the average value of V over the time dimension; four one-dimensional convolutional blocks are used after the non-stationary attention block. By setting appropriate convolutional kernels and strides, these blocks map the output of the non-stationary attention block to the value of V. Potential tactile representation z Potential tactile representation z This will be used in the next FSQ quantization phase;
[0129] Steps 2-3: For potential tactile representations FSQ bases it on quantization level l = [ l 1, l 2, ..., l d Quantized into codewords C = [C1, C2, ..., C Lc]. Subsequently, these codewords can be converted into compressed representations where L c denotes the codeword length, d denotes the length of quantization levels. In particular, the compressed scalar is obtained in the following way:
[0130]
[0131] where denotes the floor operation, denotes the j quantization level, denotes the compact haptic representation at position (i,j);
[0132] The gradients are passed through the straight-through estimator by a rounding operation:
[0133]
[0134] where sg(·) denotes the forced gradient to zero, round(·) denotes the rounding, denotes the j quantization level, denotes the compact haptic representation at position i , j After this operation, the compressed scalar is scaled to the range (-1,1), i.e. . Subsequently, the haptic representation is quantized to the positive integer codewords in C:
[0135]
[0136] where l 0=1, l t denotes the t quantization level, denotes the j quantization level, is the compressed scalar, the compressed haptic representation can be obtained from the codewords C in the following way:
[0137]
[0138] where denotes the cyclic modulo operation with modulus , denotes the floor operation, C i denotes the i-th positive integer codeword in C, l t denotes the t quantization level, denotes the j quantization level;
[0139] In the present application, the FSQ code word length L c is fixed as 6, as a hyper-parameter to balance the base bit rate and the reconstructed signal quality. The quantization levels l are set as [8, 5, 5, 5], which means that about bits are needed to transmit a single code word. In addition, 24 bits are allocated per frame for transmitting the mean µ x and standard deviation σ x . Finally, these bits constitute the base bit rate, and the theoretical code rate can reach 3.85 kbps;
[0140] After FSQ layer quantization, the generated FSQ code word C and other core parameters (mean x , standard deviation σ x ) will be transmitted to the receiving end through the wireless communication network;
[0141] Step 2-4, the receiving end restores the code word C to the compressed representation , obtains the normalized reconstructed frame through the NAE decoder, and finally performs inverse normalization processing on the mean µ x , standard deviation σ x to output the reconstructed frame ;
[0142] Step 2-5, FSQ-NAE model training: in the end-to-end training process of the present application, a multi-scale discriminator is used to improve the quality of the reconstructed haptic signal and the generalization performance of the FSQ-NAE model, and an effective loss function is designed to optimize the training. The detailed loss function definitions of the multi-scale discriminator and the FSQ-NAE model are as follows:
[0143] Multi-scale discriminator loss: the training target of the multi-scale discriminator is to minimize the following hinge adversarial loss function:
[0144]
[0145] wherein K denotes the number of discriminator models (in the present application K = 3), x' denotes the normalized frame, denotes the normalized reconstructed frame, D k (k) denotes the output of the kth discriminator at the last layer; this encourages the FSQ-NAE model to reconstruct more realistic haptic frames, with the goal of deceiving the discriminator;
[0146] FSQ-NAE model loss: the training target of this model contains four independent loss components:
[0147] Mean square error loss To ensure that the reconstructed haptic frames have excellent fidelity at the sample level, namely:
[0148]
[0149] Where x represents the input one-dimensional frame. Indicates the output reconstructed frame;
[0150] Structural similarity loss Ensure that the reconstructed haptic frames maintain structural similarity, that is:
[0151]
[0152] Where ssim(·) represents the one-dimensional SSIM calculation between two haptic frames, and x represents the input one-dimensional frame. Indicates the output reconstructed frame; adversarial loss. The aim is to overcome the discriminant effect, namely:
[0153]
[0154] in Indicates a normalized reconstructed frame. K D represents the number of discriminator models. k (·) represents the output of the k-th discriminator in the last layer; feature loss The average L1 distance at the discriminator's inner layer output for different inputs is used. To calculate, that is:
[0155]
[0156] in K The number of discriminator models is represented by M, where M represents the number of layers in each discriminator model (M = 4 in this invention), and x' represents the normalized frame. D represents the normalized reconstructed frame. m k (·) represents the output of the k-th discriminator at layer m; the overall loss of the FSQ-NAE model is a weighted sum of four different components:
[0157]
[0158] This invention employs a grid search to ensure that the contributions of the four components are of the same order of magnitude. This represents the mean squared error loss. Represents structural similarity loss. Indicating resistance to loss, Represents feature loss, λmse denotes the weight of the mean squared error loss in the model loss, λ ssim denotes the weight of the structural similarity loss in the model loss, λ adv denotes the weight of the adversarial loss in the model loss, λ feat denotes the weight of the feature loss in the model loss, which is set to λ mse = 1, λ ssim = 0.1, λ adv = 1 and λ feat = 10.
[0159] S102: Calculate the residual based on the difference between the normalized reconstructed frame and the normalized frame, use a non-uniform quantization method to quantize the residual according to the quantization bit width, generate a residual quantization index, use entropy coding to losslessly compress the FSQ codeword and the residual quantization index, and generate an output bit stream.
[0160] The process of using a non-uniform quantization method to quantize the residual according to the quantization bit width to generate a residual quantization index includes:
[0161] A non-uniform quantization method is used to quantize the residual e according to the quantization bit width B, which can be set to 0, 3, 4, 5, or 6. When B is 0, it means that no residual quantization is performed, and only the basic code rate level is used. Increasing the value of the quantization bit width B will increase the code rate level, thereby improving the fidelity of the haptic signal and generating a quantization index I of a specified bit width, i.e. Figure 2 Stage 3 part.
[0162] The specific steps are as follows:
[0163] Step 3-1, first, clip all samples in the residual to fall within the interval [-1, 1]. Next, assign a 1-bit sign code P c to represent the sign of e:
[0164]
[0165] After taking the absolute value, all samples are constrained within the range [0, 1];
[0166] Step 3-2, second, assign a quantization code P v to represent the non-uniform quantization result with a preset bit width. According to the parity of the quantization bit width, the quantization code will be different in two cases:
[0167] Step 321, when P v quantized residual error R all possible values are as follows:
[0168]
[0169] where and are scaling factors to ensure the maximum value in R is 1. Taking 4-bit non-uniform quantization as an example, the present application obtains , , and for p 0 and p 1 all (2 4 = 16) combinations have . Next, the goal is to find the quantized value closest to the current input value, i.e. At this time, 2n bits are assigned to j and k, respectively, to specify the specific values of p 0 and p 1 For e = 0.38, the present application has p 0,1 = 2 -4 and p 1,3 = 2 -1 . Thus R e = 0.375 and P v represent binary bits 0111;
[0170] Step 322, when P v quantized residual error R all possible values can be obtained in the following way:
[0171]
[0172] where , . Taking 3-bit non-uniform quantization as an example, each quantized residual error is a scaled sum of , and , where , and = 8 / 5. For e = 0.38, there are p 0,2 = 2-2 and =0, then R e =0.4 and P v denotes the binary bit 100;
[0173] Step 3-3, finally, the symbol code P c and the quantization code P v form a quantization index I, which will be used in the bitstream generation phase.
[0174] The process of lossless compression of FSQ codewords and residual quantization indices using entropy coding to generate an output bitstream is as follows:
[0175] Step 4-1, lossless compression of FSQ codewords C and residual quantization indices I using Huffman coding and generating an output bitstream, Huffman coding relies on known a priori probability models, and Huffman codebooks need to be pre-transmitted, which are obtained by statistical analysis of the occurrence of codewords in the training data set. The output bitstream includes two parts: frame header and payload;
[0176] Step 4-2, the frame header includes overlap ratio O, quantization bit width B, average value µ x and standard deviation σ x These parameters are crucial for decoding;
[0177] Step 4-3, the payload contains 6 FSQ codewords C and multiple quantization indices I, all of which are individually Huffman coded based on different codebooks.
[0178] Specifically, the following embodiments further illustrate the present application:
[0179] The following experimental results show that, compared with existing methods, the present application realizes the rate-delay scalable coding of vibrotactile signals using non-stationary attention mechanism and sliding window method, and achieves better generation effect.
[0180] The present application uses TUM standard vibrotactile signal data set to conduct extensive experiments to evaluate the performance of the present application compared with the state-of-the-art method, and further verify the scalability of the rate and delay of the present application.
[0181] In this embodiment, the performance indicators for evaluating the vibrotactile signal coding scheme proposed by the present application are divided into three categories: signal-to-noise ratio, peak signal-to-noise ratio, and ST-SIM.
[0182] Signal-to-noise ratio: Signal-to-noise ratio (SNR) is an evaluation index of signal quality. It measures the purity of the signal or the degree of noise interference by calculating the ratio of signal power to noise power (usually expressed in decibels). The higher the SNR value, the better the signal quality and the smaller the noise interference.
[0183] Peak signal-to-noise ratio: Peak signal-to-noise ratio (PSNR) is an evaluation index of signal quality, which calculates the ratio of the energy of the peak signal to the average energy of the noise. It is the most common and widely used objective evaluation index of signal. The higher the PSNR value, the smaller the distortion.
[0184] ST-SIM: ST-SIM is used to evaluate the perceptual quality, and ST-SIM has a high correlation with the average subjective score, which can be used as a feasible alternative to subjective testing. ST-SIM measures the perceptual fidelity between the original x and the reconstructed The formula is as follows:
[0185]
[0186] Where S-SIMj and T-SIMj represent the spectral similarity and temporal similarity of the jth block (block length 512) respectively. N b represents the total number of blocks contained in the signal.
[0187] Table 1 Performance comparison results of haptic encoding schemes at a basic bit rate of 3.46kbps
[0188]
[0189] Table 2 Delay comparison results of haptic encoding schemes
[0190]
[0191] From Tables 1, 2 and Figure 4 It can be seen that the method proposed in the present application has obvious advantages compared with the above-mentioned competitive methods. Compared with the four existing methods, the method of the present application shows better performance under the same bit rate and the same stream buffer length. The results show that the non-stationary autoencoder architecture proposed in the present application realizes efficient compression of haptic signals while ensuring perceptual quality; the sliding window method with overlap coding reduces the buffer delay; the non-uniform quantization method dynamically enhances the fidelity of haptic signals.
[0192] As Figure 2 shown, Figure 2For the complete network structure diagram of the application, the left side is stage 1: preprocessing, tactile signal input, frame division through a sliding window, the left side is a buffer, a variety of delay options are provided by controlling the ratio of the overlap length to the frame length, that is, the overlap rate, the DFT321 algorithm is applied to the divided frame to obtain the preprocessed frame; the upper middle part is stage 2: FSQ-NAE model, the preprocessed frame x is normalized to obtain the mean, variance and x', which is compressed into a latent compact tactile representation through an encoder z , and is further processed by limited scalar quantization to , and , the difference between x' and non-overlapping residual e is taken as the output of the next stage; the lower middle part is stage 3: residual quantization, the non-overlapping residual e is non-uniformly quantized according to the quantization bit width, and the quantized residual is output; the right part is stage 4 bit stream generation: the FSQ code word and the quantization index are respectively entropy encoded and compressed as codebook 1 and codebook 2, the encoded FSQ code word and the encoded quantization index are taken as the payload of the data packet, the frame header of the data packet includes the overlap rate, the quantization bit width and the mean variance, the frame header and the payload constitute the data packet, and the output bit stream is generated.
[0193] As shown in Figure 3 , the model architecture diagram of the application is shown in Figure 3 , the left side is the training process, the preprocessed three-axis vibration tactile signal x is taken as the input frame, and x' obtained after normalization is input into the encoder; in the encoder, first pass through the embedding layer, then enter the non-stationary attention module, extract Q', K' and V', Q', K' and V' represent the average values of Query, Key and Value in the time dimension, and the multi-layer perception MLP directly obtains the non-stationary factor τ and , and the mean µ x and the standard deviation σ x , Q' and K' which perform matrix multiplication together perform Rescale function operation, then pass through the Softmax function, and output matrix multiplication with V', which passes through residual connection & layer normalization, convolution block, residual connection & layer normalization; after the non-stationary attention module, four one-dimensional convolutions are used to obtain the latent compact tactile representation z ; the latent compact tactile representation z is quantized into , , which enters the decoder, passes through four transposed one-dimensional convolutions, and then passes through the non-stationary attention module, in which Rescale function and Softmax function operations are first performed, then residual connection & layer normalization, convolution block, residual connection & layer normalization are performed to generate the normalized reconstructed frame The reconstructed frame is obtained after denormalization. During end-to-end training, a multi-scale discriminator is used to improve the quality of the reconstructed tactile signal and the generalization performance of the FSQ-NAE model; the input frame x and the reconstructed frame x are calculated. Mean squared error and structural similarity, mean squared error loss Structural similarity loss Combating losses and feature loss The four independent components trained for this model The image shows the discriminator loss. The right side of the image illustrates the deployment process: at the transmitting end, the input frame is first normalized, then passes through the encoder, and the output undergoes finite scalar quantization. In the network, the FSQ codeword and mean-variance are input to the receiving end, and after passing through the decoder, are denormalized to obtain the reconstructed frame. .
[0194] like Figure 4 As shown, Figure 4 This is a comparison chart of the signal-to-noise ratio (SNR) of this invention with five other comparison methods. The horizontal axis of the chart represents the bit rate in kbps, and the vertical axis represents the signal-to-noise ratio (SNR) in dB. The legend is in the upper left corner. From top to bottom, the methods are VPC-DS (buffer = 512), PVC-SLP (buffer = 200), VC-PWQ (buffer = 512), RNVC (buffer = 512), RNVC (buffer = 64), and this invention (buffer = 64). Each point on the curve corresponding to this invention in the chart corresponds to the bit rate and SNR of this invention when the quantization bit width B = 3, 4, 5, and 6, respectively.
[0195] like Figure 5 As shown, Figure 5 This is a comparison chart of the peak signal-to-noise ratio (PSNR) of this invention with five other comparison methods. The horizontal axis of the chart represents the bit rate in kbps, and the vertical axis represents the peak signal-to-noise ratio (PSNR) in dB. The legend is in the upper left corner, and from top to bottom, the methods are VPC-DS (buffer = 512), PVC-SLP (buffer = 200), VC-PWQ (buffer = 512), RNVC (buffer = 512), RNVC (buffer = 64), and this invention (buffer = 64).
[0196] like Figure 6 As shown, Figure 6This is a comparison chart of the present invention with five other ST-SIM methods. The horizontal axis represents the bit rate in kbps, the vertical axis represents ST-SIM, and the lower right corner is the legend. From top to bottom, the methods are VPC-DS (buffer = 512), PVC-SLP (buffer = 200), VC-PWQ (buffer = 512), RNVC (buffer = 512), RNVC (buffer = 64), and the present invention (buffer = 64).
[0197] like Figure 7 As shown, Figure 7 This is a flowchart of the workflow of this invention. The TUM standard vibration and tactile signal dataset is divided into two parts: a training set and a test set. The training set is used to train the FSQ-NAE model. The vibration and tactile signals in the training set are first subjected to sliding window framing and DFT321 algorithm. In the FSQ-NAE model, the signals are first normalized, then encoded using an attention-based non-stationary autoencoder, then finite scalar quantization, then decoded by an attention-based non-stationary autodecoder, and finally denormalized. During the training process, a multi-scale discriminator is used to improve the quality of the reconstructed tactile signals and the generalization performance of the FSQ-NAE model. The trained FSQ-NAE model will be used in the testing phase. The test set includes 280 vibration and tactile signals. These 280 vibration and tactile signals are input into the trained FSQ-NAE model, the output is subjected to residual quantization, and finally bitstream generation is achieved.
[0198] Example 2: To achieve the above objective, such as Figure 8 As shown, based on Embodiment 1, this invention discloses a rate-delay scalable coding system for vibration tactile signals, comprising:
[0199] Signal processing module 11 is used to receive triaxial vibration tactile signals, preprocess the triaxial vibration tactile signals to obtain preprocessed triaxial vibration tactile signals, input the preprocessed triaxial vibration tactile signals into a pre-established non-stationary autoencoder model based on finite scalar quantization and attention, and output the difference between the normalized reconstructed frame and the normalized frame as well as the FSQ codeword.
[0200] The compression coding module 12 is used to calculate the residual based on the difference between the normalized reconstructed frame and the normalized frame, and to use a non-uniform quantization method to quantize the residual according to the quantization bit width to generate a residual quantization index. The FSQ codeword and the residual quantization index are then losslessly compressed using entropy coding to generate an output bit stream.
[0201] Based on the same inventive concept, the present application further provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the program comprises program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are configured to implement one or more instructions, and are specifically configured to load and execute one or more instructions in the computer storage medium to implement the above method.
[0202] It needs to be further explained that, based on the same inventive concept, the present application further provides a computer storage medium, which stores a computer program, and the computer program is executed by the processor to perform the above method. The storage medium can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0203] In the description of the present application, the description of the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0204] The foregoing presents and describes the basic principles, main features and advantages of the present disclosure. It should be understood by those skilled in the art that the present disclosure is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, various changes and improvements can be made to the present disclosure, and all these changes and improvements fall within the scope of the present disclosure.
Claims
1. A rate-delay scalable coding method for vibration tactile signals, characterized in that, The method includes the following steps: The system receives triaxial vibration tactile signals, preprocesses the triaxial vibration tactile signals to obtain preprocessed triaxial vibration tactile signals, inputs the preprocessed triaxial vibration tactile signals into a pre-established non-stationary autoencoder model based on finite scalar quantization and attention, and outputs the difference between the normalized reconstructed frame and the normalized frame, as well as the FSQ codeword. The process of inputting the preprocessed triaxial vibration tactile signal into a pre-established non-stationary autoencoder model based on finite scalar quantization and attention, and outputting the difference between the normalized reconstructed frame and the normalized frame, as well as the FSQ codeword, includes: The preprocessed one-dimensional frame x is initially normalized to generate a normalized frame x'. Then, the attention-based non-stationary autoencoder model NAE compresses x' into a latent compact tactile representation. z The FSQ layer will further z The code is quantized into a codebook C, which serves as the FSQ codeword. Core parameters are transmitted to the receiving end via a wireless communication network. These core parameters include the mean value. µ x Standard deviation σ x and codeword C; The receiving end first restores the codeword C to its compressed representation. Normalized reconstructed frames are obtained through the NAE decoder. Then, the reconstructed frames are output through inverse normalization processing. In this process, a multi-scale discriminator loss function is additionally constructed during end-to-end training. The training of the pre-established non-stationary autoencoder model based on finite scalar quantization and attention is as follows: The loss function is designed to optimize training. The loss function for the multi-scale discriminator and the FSQ-NAE model is defined as follows: Multi-scale discriminator loss: The discriminator training objective is to minimize the following hinge adversarial loss function. : in K This indicates the number of discriminator models, mean (·) indicates the calculation of the average value, and x' represents the normalized frame. D represents the normalized reconstructed frame. k (·) represents the output of the k-th discriminator in the last layer; FSQ-NAE Model Loss: The training objectives of the FSQ-NAE model include mean squared error loss, structural similarity loss, adversarial loss, and feature loss. Mean square error loss To ensure that the reconstructed haptic frames have excellent fidelity at the sample level, namely: Where x represents the input one-dimensional frame. This indicates the output reconstructed frame, where L represents the frame length. Structural similarity loss Ensure that the reconstructed haptic frames maintain structural similarity, that is: Where ssim(·) represents the one-dimensional SSIM calculation between two haptic frames, and x represents the input one-dimensional frame. Indicates the output reconstructed frame; Combating losses The aim is to overcome the discriminant effect, namely: in Indicates a normalized reconstructed frame. K D represents the number of discriminator models. k (·) represents the output of the k-th discriminator in the last layer; Feature loss The average L1 distance at the discriminator's inner layer output for different inputs is used. To calculate, that is: Where M represents the number of layers in each discriminator model, K This represents the number of discriminator models, and x' represents the normalized frame. D represents the normalized reconstructed frame. m k (·) represents the output of the k-th discriminator in the m-th layer; the loss of the FSQ-NAE model. L G It is a weighted sum of four different components: in Indicates the mean square error loss. Represents structural similarity loss. Indicating resistance to loss, Represents feature loss, λ mse This indicates the weight of the mean squared error loss in the model loss. λ ssim This represents the weight of structural similarity loss in the model loss. λ adv This indicates the weight of the adversarial loss in the model loss. λ feat This represents the weight of the feature loss in the model loss; The non-stationary attention mechanism of the FSQ-NAE model is as follows: Starting from the self-attention mechanism, the original non-stationary tactile frame x is received as input: Where Q, K, and V represent Query, Key, and Value, respectively, and the length of Q, K, and V is L, and the dimension is d. k Attn(·) represents the attention mechanism, and Softmax(·) represents the normalized exponential function; the embedding layer is represented by F(·), which is assumed to have linear properties and is applied to each input sample. Therefore, Q = [q1, q2, ... q L ] T Each query in the table can be computed as q. i = F(x) i ); Using the generated normalized haptic frame x' as input, Q' = [F(x1'), F(x2'), ..., F(x... L ')] T Derived based on linear properties ,in It is the average value of Q over the time dimension. σ x Indicates standard deviation; The relationship between attention weights based on the original tactile framework and the normalized framework: Where Q and K represent length L and dimension d, respectively. k Query and Key µ Q This represents the average value of Q over the time dimension. µ K This represents the average value of K over the time dimension. σ x Indicates standard deviation. , , , , respectively Summing is performed on each column and each element within the formula. Performing a Softmax(·) operation on both sides, where Softmax(·) represents the normalized exponential function, reveals the direct relationship between the attention weights: Where d k The dimension is represented by a positive scaling scalar defined as a non-stationary factor. τ = σ x 2 and shift vector The non-stationary factor is learned directly from the original tactile frame x using a multilayer perceptron (MLP) layer. τ and and statistics µ x and σ x Non-stationary attention The calculation is as follows: in , µ V This represents the average value of V over the time dimension; by setting the convolution kernel and stride through a one-dimensional convolution block, the output of the non-stationary attention block is mapped to the size. Potential tactile representation z ; The residual is calculated based on the difference between the normalized reconstructed frame and the normalized frame. A non-uniform quantization method is used to quantize the residual according to the quantization bit width to generate a residual quantization index. The FSQ codeword and the residual quantization index are then losslessly compressed using entropy coding to generate the output bit stream.
2. The rate-delay scalable coding method for vibration tactile signals according to claim 1, characterized in that, The preprocessing of the triaxial vibration tactile signal includes: The triaxial vibration tactile signal is buffered in a buffer and then segmented through a sliding window. Based on the overlap rate, the signal is divided into overlapping or non-overlapping frames. The DFT321 algorithm is applied to convert the tactile frames in the three-dimensional vibration tactile signal into one-dimensional frames, thus obtaining the preprocessed triaxial vibration tactile signal. According to the DFT321 algorithm, the spectrum and phase θ The expression is as follows: Where f represents frequency, A1 represents the discrete Fourier transform of the original tactile frame along the X-axis, A2 represents the discrete Fourier transform of the original tactile frame along the Y-axis, and A3 represents the discrete Fourier transform of the original tactile frame along the Z-axis. The one-dimensional frame x can be derived from the preset spectrum and phase, as shown in the following expression: Where iDFT(·) denotes the inverse discrete Fourier transform, and real(·) is used to extract the real part from the complex value. Represents the spectrum. θ The phase is represented; after the DFT321 algorithm, the three-dimensional vibration tactile frame is converted into a one-dimensional frame x.
3. The rate-delay scalable coding method for vibration tactile signals according to claim 2, characterized in that, The process of initial normalizing the preprocessed one-dimensional frame x: Where x = [x1, x2, ..., x] L ] T Represents a sample of a one-dimensional frame. µ x This represents the mean. σ x Let x' represent the standard deviation, L represent the frame length, and x' represent the normalized frame, where x i Let i represent the sample at position i, where i = 1, 2, ..., L; The potential compact tactile representation FSQ will z Based on quantization level l = [ l 1, l 2, ..., l d Quantized into codewords C = [C1, C2, ..., C Lc [This converts codewords into a compressed representation.] L c Indicates codeword length, d represents the length of the quantization level, and is a compressed scalar. Obtain it in the following ways: in, This indicates a round-down operation. Indicates the first j One quantification level, express( i, j Compact tactile representation of location; A pass-through estimator is used to pass gradients through rounding operations: Where sg(·) represents forcing the gradient to zero, and round(·) represents rounding. Indicates the first j One quantification level, express( i , j Compact tactile representation of location, compressed scalar It is scaled to the range (-1, 1), i.e. Subsequently, the tactile representation was quantized into positive integer codewords in C. : in l 0=1, l t Indicates the first t One quantification level, Indicates the first j One quantification level, To compress scalars, compress tactile representations It can be obtained from codeword C in the following way: in Indicates For the modulus, perform a cyclic modulo operation. Indicates the floor operation, C i Let l represent the i-th positive integer codeword in C. t Indicates the first t One quantification level, Indicates the first j Each quantification level.
4. The rate-delay scalable coding method for vibration tactile signals according to claim 1, characterized in that, The process of using a non-uniform quantization method to quantize the residual based on the quantization bit width and generate a residual quantization index includes: Clamp all samples within the residuals so that the residuals fall within the interval [−1, 1], and assign a 1-bit sign code. P c The symbol to represent 'e': After taking the absolute value, all samples are constrained to the range [0,1]. Assign quantization code P v To represent non-uniform quantization results with a preset bit width; Symbol code P c With allocation quantization code P v The combination of these elements forms the residual quantization index I.
5. A rate-delay scalable coding method for vibration tactile signals according to claim 4, characterized in that, The allocation quantization code P v The process of representing a non-uniform quantization result with a preset bit width is described. Depending on the parity of the quantization bit width, there are two different cases for the quantization code: when P v When represented using 2n bits, the quantization residual R All values are as follows: in and It is to ensure R The maximum value of the proportionality coefficient is 1; when P v When using 2n+1 bits for representation, the quantization residual R All possible values can be obtained in the following ways: in , .
6. A rate-delay scalable coding system for vibration-tactile signals, employing the rate-delay scalable coding method for vibration-tactile signals as described in any one of claims 1 to 5, characterized in that, include: The signal processing module is used to receive triaxial vibration tactile signals, preprocess the triaxial vibration tactile signals to obtain preprocessed triaxial vibration tactile signals, input the preprocessed triaxial vibration tactile signals into a pre-established non-stationary autoencoder model based on finite scalar quantization and attention, and output the difference between the normalized reconstructed frame and the normalized frame as well as the FSQ codeword. The compression coding module is used to calculate the residual based on the difference between the normalized reconstructed frame and the normalized frame. It uses a non-uniform quantization method to quantize the residual according to the quantization bit width, generates a residual quantization index, and performs lossless compression of the FSQ codeword and the residual quantization index using entropy coding to generate the output bit stream.
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it employs a rate-delay scalable coding method for vibration tactile signals as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method and a system used for touch data encoding and stream transfer
CN104184721A
Method, system and equipment for visually impaired person to recognize virtual texture
CN117763430A