Cross-modal information source coding and decoding method based on semantic association

By constructing semantic prior information between the primary and secondary modalities, the secondary modal signal is divided into semantically relevant and irrelevant parts. Cross-modal cooperative coding and standard single-modal coding are used to solve the problem of low efficiency in multimodal coding and achieve efficient multimodal data communication.

CN121125014APending Publication Date: 2025-12-12NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511211429.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing multimodal coding methods fail to effectively utilize semantic redundancy between modalities, resulting in low coding efficiency and insufficient utilization of cross-modal semantic redundancy, making it difficult to meet the high-efficiency compression requirements of multimodal data communication.

Method used

By constructing semantic prior information between the primary and secondary modes, the secondary mode signal is divided into semantically relevant and irrelevant parts. Cross-modal cooperative coding and standard single-modal coding are used to process them separately. Combined with Slepian-Wolf coding and existing coding standards, a heterogeneous bitstream is generated for transmission, and the signal is reconstructed at the decoding end.

Benefits of technology

It significantly reduces redundant data, improves overall compression efficiency, and ensures reconstruction quality even under unstable channel conditions, achieving synergistic optimization of compression efficiency and reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125014A_ABST
    Figure CN121125014A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information processing, and discloses a cross-modal information source encoding and decoding method based on semantic association, and the method specifically comprises the following steps: obtaining information source signals of a plurality of modals, inputting the information source signals into an encoding end, determining a main modal according to the service quality index parameters of each modal, and enabling the other modals to be secondary modals; according to the type of the main modal signal, adopting a corresponding preset coding strategy to perform feature retention type coding processing on the main modal signal source signal so as to retain semantic features necessary for cross-modal decoding; semantic prior information between the primary mode and each secondary mode is constructed and used for describing a semantic label of a primary mode information source signal and a semantic consistency relation between the primary mode and the secondary modes; dividing the sub-modal signal source signal into a semantic correlation signal part and a semantic irrelevant signal part based on semantic prior information; the problems of low multi-modal coding efficiency and insufficient cross-modal semantic redundancy utilization in the prior art are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information processing technology, specifically relating to a cross-modal source encoding and decoding method based on semantic association. Background Technology

[0002] Multimodal services that integrate auditory, visual, and tactile experiences can significantly enhance user immersion and have shown broad application prospects in various fields such as healthcare, virtual reality, education, and remote training. However, multimodal data communication requires simultaneous low latency, high reliability, and large capacity, which places a significant burden on existing wireless bearer networks. Therefore, higher technical requirements are placed on the efficient compression of multimodal data. Existing video, audio, and tactile signal source encoding and decoding compression methods mainly employ their own independent dedicated coding frameworks: video typically uses techniques such as spatiotemporal prediction, transform coding, and entropy coding (e.g., HEVC, VVC) to remove redundant inter- and intra-frame information; audio compression relies on frequency domain analysis, perceptual models, and entropy coding (e.g., AAC, Opus) to achieve efficient encoding; tactile signals often use compression methods based on perceptual dead zones, reducing the bit rate by discarding weak force feedback that is imperceptible to humans. These methods can effectively compress data redundancy within their respective modalities, but they generally ignore the high-level semantic redundancy that may exist between modalities, thus making it difficult to achieve optimal overall compression performance in multimodal scenarios. In typical multimodal applications, different modal sources usually represent the same object from different perceptual dimensions, thus exhibiting natural semantic correlations between modalities. With the empowerment of artificial intelligence, joint learning of multimodal data samples can uncover deep semantic relationships between different modalities and construct a unified shared semantic space. Based on the mapping relationships between different modal data in this space, it is expected to achieve intermodal information collaboration and semantic-level fusion, providing a new path for efficient multimodal data communication. Although existing cross-modal semantic analysis techniques have been widely applied in knowledge retrieval, content generation, and signal reconstruction, there is currently a lack of universal and effective solutions for introducing them into communication systems to fully utilize semantic relationships to compress redundancy between multimodal sources and reduce overall transmission overhead. Therefore, existing technologies suffer from low multimodal coding efficiency and insufficient utilization of cross-modal semantic redundancy. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to provide a cross-modal source coding and decoding method based on semantic association, which solves the problems of low multimodal coding efficiency and insufficient utilization of cross-modal semantic redundancy in existing technologies.

[0004] The objective of this invention can be achieved through the following technical solutions:

[0005] A cross-modal source encoding and decoding method based on semantic association specifically includes the following steps:

[0006] Multiple modal source signals are acquired and input into the encoding terminal. A primary modality is determined based on the quality of service index parameters of each modality, and the remaining modalities are secondary modalities.

[0007] Based on the type of the main modal signal, a corresponding preset coding strategy is adopted to perform feature-preserving coding processing on the main modal source signal in order to preserve the semantic features necessary for cross-modal decoding;

[0008] Construct semantic prior information between the main modality and each submodality to describe the semantic labels of the main modality source signal and the semantic consistency relationship between the main modality and the submodality;

[0009] Based on semantic prior information, the submodal source signal is divided into a semantically relevant signal part and a semantically irrelevant signal part;

[0010] Perform cross-modal cooperative coding on the semantically related signal portion;

[0011] Perform standard single-mode coding on semantically irrelevant signal components;

[0012] A multiplexing mechanism is adopted to merge the encoding results of the main modal source signal, the encoding results of the semantically related signal part, the encoding results of the semantically unrelated signal part, and the semantic prior information to form a heterogeneous bitstream. The heterogeneous bitstream is the final output of the encoding end and is transmitted to the decoding end.

[0013] At the decoding end, the primary mode source signal is reconstructed and semantic association information is extracted to assist in generating a reference signal for the secondary mode source signal;

[0014] Based on the reconstruction results of the primary modal source signal and semantic prior information, cross-modal collaborative decoding is performed on the semantically relevant signal part, and standard single-modal decoding is used for the semantically irrelevant signal part to recover the secondary modal source signal.

[0015] The acquired source signals for each modality include at least two of the following: video signals, audio signals, tactile signals, and text signals;

[0016] Based on the service quality index parameters of each modality, a primary modality is determined, and the remaining modalities are secondary modalities. The specific steps include:

[0017] Determine the service quality indicator parameters, including latency sensitivity and reliability sensitivity;

[0018] Delay sensitivity is determined based on the maximum end-to-end transmission delay allowed for the source signal of that mode;

[0019] The reliability sensitivity is determined based on the maximum allowable packet loss rate of the source signal in that mode;

[0020] Based on preset weighted scoring rules, a comprehensive score is given for the latency sensitivity and reliability sensitivity of each modality type;

[0021] Sort all modal types from highest to lowest based on their overall score;

[0022] The modality with the highest score is identified as the primary modality, and the remaining modalities are designated as secondary modalities. The semantic prior information between the primary modality and each secondary modality is constructed, specifically including the following steps:

[0023] Semantic analysis is performed on the main modal source signal to extract its semantic tags. The semantic tags include at least one of object category, material attribute, action type and event state.

[0024] For each submodal source signal, a semantic consistency relationship is established between it and the main modal source signal. The semantic consistency relationship includes local segment alignment relationship, semantic attribute similarity score and cross-modal semantic mapping function.

[0025] Local segment alignment relationships are used to identify the correspondence between submodal signal segments and primary modal semantic tags;

[0026] Semantic attribute similarity scoring is used to quantify the degree of matching between the submodality and the main modality in terms of semantic features;

[0027] Cross-modal semantic mapping functions are used to describe the transformation relationship of semantic features between different modalities;

[0028] Semantic tags and semantic consistency relationships are encoded into structured metadata to form semantic prior information.

[0029] Preset encoding strategies include:

[0030] When the main modal source signal is a video signal, a compression algorithm based on inter-frame prediction is used in combination with perceptual quantization and entropy coding techniques.

[0031] When the main modal source signal is a tactile signal, non-uniform quantization and predictive coding based on Weber's law are used.

[0032] When the main modal source signal is an audio signal, a modified discrete cosine transform or parameterized coding is used;

[0033] When the primary modal source signal is a text signal, context-independent compression or semantic keyword extraction is used.

[0034] Performing cross-modal cooperative coding on the semantically related signal portion includes the following steps:

[0035] The semantically relevant signal components are quantized to preserve signal features and control distortion.

[0036] Convert the quantization result into multiple bit planes;

[0037] Based on the coding results of the main mode source signal and semantic prior information, Slepian-Wolf coding is performed on each bit plane to generate a parser as the compression result;

[0038] Slepian-Wolf coding uses LDPC codes, Polar codes, or Turbo codes as channel codes.

[0039] Performing standard single-mode coding on semantically irrelevant signal components includes the following steps:

[0040] When the submodal source signal is a video signal, intra-frame or inter-frame coding is performed using HEVC or H.264 standards.

[0041] When the submodal source signal is an audio signal, MP3 or AAC standards are used for compression encoding.

[0042] When the submodal source signal is a tactile signal, an entropy coding method based on the sensory dead zone is adopted.

[0043] The primary mode source signal is reconstructed and semantic association information is extracted at the decoding end to assist in generating a reference signal for the secondary mode source signal. This process includes the following steps:

[0044] Separate the encoding result of the main mode source signal from the received heterogeneous bitstream;

[0045] The encoded results of the separated main mode source signals are decoded and reconstructed.

[0046] Semantic prior information is extracted from heterogeneous bitstreams to obtain the main modality semantic labels and the semantic consistency relationship between the main modality and the submodality;

[0047] The reconstructed primary modal source signal and the extracted semantic prior information are input into the cross-modal generator to generate the reference signal for the secondary modal source signal.

[0048] Cross-modal generators include any of the following: conditional variational autoencoders, diffusion models, or cross-modal neural networks.

[0049] The recovery of the submodal source signal includes the following steps:

[0050] Perform cross-modal cooperative decoding on the semantically related signal portion, perform Slepian-Wolf decoding on the received checksum, and recover the semantically related signal content by combining the reference signal;

[0051] Perform standard single-mode decoding on semantically irrelevant signal components:

[0052] When the submode is a video signal, it is decoded using a HEVC or H.264 standard decoder.

[0053] When the submode is an audio signal, it is decoded using an MP3 or AAC standard decoder.

[0054] When the submodal is a tactile signal, interpolation reconstruction based on the sensory dead zone is used.

[0055] The beneficial effects of this invention are:

[0056] This invention constructs semantic prior information between the primary modality and the secondary modality, divides the secondary modality signal into semantically relevant and irrelevant parts, and processes them using cross-modal cooperative coding and standard single-modal coding respectively, effectively solving the problems of low efficiency of multimodal coding and insufficient utilization of cross-modal semantic redundancy in the prior art;

[0057] Since the semantically relevant part uses the main modality signal as a reference for Slepian-Wolf encoding, the amount of redundant data is significantly reduced. At the same time, the semantically irrelevant part uses the traditional encoding standard to ensure compatibility, thereby improving the overall compression efficiency.

[0058] At the decoding end, a secondary mode reference signal is generated based on the reconstructed primary mode signal and semantic prior information. Combined with a cross-modal collaborative decoding mechanism, the signal can be recovered with high quality even when the channel conditions are unstable, thus solving the problem of unstable reconstruction quality in real-time communication scenarios. This method not only fully explores the semantic correlation between modes, but also achieves synergistic optimization of compression efficiency and reconstruction quality through modular design to be compatible with existing coding standards. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a schematic diagram of the encoding and decoding process of the present invention;

[0061] Figure 2 This is a schematic diagram of the core components of the cross-modal source encoder of the present invention;

[0062] Figure 3 This is a schematic diagram of the core components of the cross-modal source decoder of the present invention;

[0063] Figure 4This is a schematic diagram of the remote operation platform in an embodiment of the present invention;

[0064] Figure 5 This is a schematic diagram of the video signal in an embodiment of the present invention;

[0065] Figure 6 This is a schematic diagram of the audio signal in an embodiment of the present invention;

[0066] Figure 7 This is a schematic diagram of tactile signals in an embodiment of the present invention;

[0067] Figure 8 This is a schematic diagram illustrating the video coding efficiency analysis results in an embodiment of the present invention;

[0068] Figure 9 This is a schematic diagram illustrating the audio coding efficiency analysis results in an embodiment of the present invention. Detailed Implementation

[0069] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0070] like Figure 1 As shown, a cross-modal source encoding and decoding method based on semantic association specifically includes the following steps:

[0071] Multiple modal source signals are acquired and input into the encoding terminal. A primary modality is determined based on the Quality of Service (QoS) index parameters of each modality, and the remaining modalities are secondary modalities.

[0072] Based on the type of the main modal signal, a corresponding preset coding strategy is adopted to perform feature-preserving coding processing on the main modal source signal in order to preserve the semantic features necessary for cross-modal decoding;

[0073] Construct semantic prior information between the main modality and each submodality to describe the semantic labels of the main modality source signal and the semantic consistency relationship between the main modality and the submodality;

[0074] Based on semantic prior information, the submodal source signal is divided into a semantically relevant signal component and a semantically irrelevant signal component, as follows:

[0075] When the semantic consistency label of a local segment meets the preset relevance threshold, the segment is divided into a semantically relevant signal part;

[0076] When the semantic consistency label of a local segment does not reach the preset relevance threshold, the segment is classified as a semantically irrelevant signal part;

[0077] Perform cross-modal cooperative coding on the semantically related signal portion;

[0078] Perform standard single-mode coding on semantically irrelevant signal components;

[0079] A multiplexing mechanism is adopted to merge the encoding results of the main modal source signal of the multimodal multiplexer, the encoding results of the semantically related signal part, the encoding results of the semantically unrelated signal part, and the semantic prior information to form a heterogeneous bitstream. The heterogeneous bitstream is the final output of the encoding end and is transmitted to the decoding end.

[0080] At the decoding end, the primary mode source signal is reconstructed and semantic association information is extracted to assist in generating a reference signal for the secondary mode source signal;

[0081] Based on the reconstruction results of the primary modal source signal and semantic prior information, cross-modal collaborative decoding is performed on the semantically relevant signal part, and standard single-modal decoding is used for the semantically irrelevant signal part to recover the secondary modal source signal.

[0082] The acquired source signals for each modality include at least two of the following: video signals, audio signals, tactile signals, and text signals;

[0083] Based on the service quality index parameters of each modality, a primary modality is determined, and the remaining modalities are secondary modalities. The specific steps include:

[0084] Determine the service quality indicator parameters, including latency sensitivity and reliability sensitivity;

[0085] Delay sensitivity is determined based on the maximum end-to-end transmission delay allowed for the source signal of that mode;

[0086] The reliability sensitivity is determined based on the maximum allowable packet loss rate of the source signal in that mode;

[0087] Based on preset weighted scoring rules, a comprehensive score is given for the latency sensitivity and reliability sensitivity of each modality type;

[0088] Sort all modal types from highest to lowest based on their overall score;

[0089] The modality with the highest score is identified as the primary modality, and the remaining modalities are designated as secondary modalities. The semantic prior information between the primary modality and each secondary modality is constructed, specifically including the following steps:

[0090] Semantic analysis is performed on the main modal source signal to extract its semantic tags. The semantic tags include at least one of object category, material attribute, action type and event state.

[0091] Preferably, semantic labels can be automatically obtained using modules such as classifiers, semantic segmenters, and Transformers;

[0092] For each submodal source signal, a semantic consistency relationship is established between it and the main modal source signal. The semantic consistency relationship includes local segment alignment relationship, semantic attribute similarity score and cross-modal semantic mapping function.

[0093] Local segment alignment relationships are used to identify the correspondence between submodal signal segments and primary modal semantic tags;

[0094] Semantic attribute similarity scoring is used to quantify the degree of matching between the submodality and the main modality in terms of semantic features;

[0095] Cross-modal semantic mapping functions are used to describe the transformation relationship of semantic features between different modalities, and are used to guide subsequent cross-modal encoding and decoding;

[0096] Semantic tags and semantic consistency relationships are encoded into structured metadata to form semantic prior information.

[0097] Preset encoding strategies include:

[0098] When the main modal source signal is a video signal, a compression algorithm based on inter-frame prediction (such as HEVC) combined with perceptual quantization and entropy coding techniques is used to achieve a high compression ratio while maintaining structural and texture fidelity.

[0099] When the main modal source signal is a tactile signal, non-uniform quantization and predictive coding based on Weber's law are used to adapt to the low-latency transmission requirements.

[0100] When the main modal source signal is an audio signal, modified discrete cosine transform (such as MDCT) or parametric coding (such as CELP) is used to compress redundancy while maintaining intelligibility and naturalness.

[0101] When the main modal source signal is a text signal, context-independent compression or semantic keyword extraction is used to reduce redundant information.

[0102] Performing cross-modal cooperative coding on the semantically related signal portion includes the following steps:

[0103] The semantically relevant signal components are quantized to preserve signal features and control distortion.

[0104] Convert the quantization result into multiple bit planes;

[0105] Based on the coding results of the main mode source signal and semantic prior information, Slepian-Wolf coding is performed on each bit plane to generate a parser as the compression result;

[0106] Slepian-Wolf coding uses LDPC codes, Polar codes, or Turbo codes as channel codes.

[0107] Performing standard single-mode coding on semantically irrelevant signal components includes the following steps:

[0108] When the submodal source signal is a video signal, intra-frame or inter-frame coding is performed using HEVC or H.264 standards.

[0109] When the submodal source signal is an audio signal, MP3 or AAC standards are used for compression encoding.

[0110] When the submodal source signal is a tactile signal, an entropy coding method based on the sensory dead zone is adopted.

[0111] The primary mode source signal is reconstructed and semantic association information is extracted at the decoding end to assist in generating a reference signal for the secondary mode source signal. This process includes the following steps:

[0112] Separate the encoding result of the main mode source signal from the received heterogeneous bitstream;

[0113] The encoded results of the separated main mode source signals are decoded and reconstructed.

[0114] Semantic prior information is extracted from heterogeneous bitstreams to obtain the main modality semantic labels and the semantic consistency relationship between the main modality and the submodality;

[0115] The reconstructed primary modal source signal and the extracted semantic prior information are input into the cross-modal generator to generate the reference signal for the secondary modal source signal.

[0116] Cross-modal generators include any of the following: conditional variational autoencoders, diffusion models, or cross-modal neural networks.

[0117] The recovery of the submodal source signal includes the following steps:

[0118] Perform cross-modal cooperative decoding on the semantically related signal portion, perform Slepian-Wolf decoding on the received checksum, and recover the semantically related signal content by combining the reference signal;

[0119] Perform standard single-mode decoding on semantically irrelevant signal components:

[0120] When the submode is a video signal, it is decoded using a HEVC or H.264 standard decoder.

[0121] When the submode is an audio signal, it is decoded using an MP3 or AAC standard decoder.

[0122] When the submodal is a tactile signal, interpolation reconstruction based on the sensory dead zone is used.

[0123] This embodiment is deployed on Figure 4 The remote operation platform shown has a two-way communication link between the master and slave terminals:

[0124] Main control domain: includes computer and force feedback haptic device; receives audio-visual-tactile feedback signals from subordinate domain via computer, processes them through force feedback haptic device and outputs haptic control signals to subordinate domain;

[0125] Preferably, the Geomagic Touch force feedback device (sampling rate 1kHz) can be used to achieve tactile interaction;

[0126] Subordinate domain: Equipped with a UR3 robotic arm featuring a pressure sensor (tactile), camera (video), and microphone (audio);

[0127] Communication link: Transmission is achieved through a wireless network, with the uplink (slave end to master end) enabling simultaneous transmission of audio, video, and touch data.

[0128] Input signals: Multimodal data streams generated by the interaction of the robotic arm with different materials (brass, silk, spandex, rags, wood), including video signals (such as...). Figure 5 As shown), audio signals (such as...) Figure 6 (as shown) and tactile signals (such as) Figure 7 (as shown);

[0129] Tactile signal: Time-series force feedback data collected by a pressure sensor;

[0130] Video signal: RGB frame sequence (resolution 1920×1080);

[0131] Audio signal: 16-bit / 48kHz sampling rate sound wave data;

[0132] Modal priority determination: Based on QoS sensitivity analysis (latency requirements: tactile < audio < video; reliability requirements: tactile > video > audio), tactile is determined to be the primary modality, and video / audio is the secondary modality;

[0133] Tactile signal compression coding: The tactile signal undergoes "quantization perception & dead-zone coding" processing. Specifically, a quantization and perception dead-zone coding scheme based on Weber's law is adopted. After converting the original signal into discrete integer values ​​through linear quantization, only signal changes exceeding the human-perceptible difference threshold need to be encoded; other changes can be ignored, thus saving bit rate. First, the tactile signal undergoes high-precision quantization processing, and then based on the current tactile pressure value F... t Set the minimum perceived change threshold Δ t =k·|F tIn this example, k is 0.1; only when the difference between the current value and the previous encoded value is greater than the threshold Δ. t Encoding is triggered only when the current sample is selected; otherwise, the current sample is skipped and no encoding is required. For tactile data that needs to be encoded, Huffman coding is used to entropy encode it.

[0134] The semantic dependencies of video, audio, and tactile signals are described in text form, specifically including: the semantic consistency between each macroblock in the video frame and the tactile material: determining whether the video content of each macroblock is consistent with the current tactile material, with consistency marked as 1 and otherwise as 0; the consistency between the audio frame and the tactile material: determining whether the current audio segment expresses a physical event or sound source material that corresponds to tactile sensation (such as rubbing a wooden board), with consistency represented by 0 / 1; the compression of semantic prior information uses Huffman coding to output a semantic text bitstream;

[0135] Audio signal compression encoding: such as Figure 2 As shown, the audio signal undergoes "MDCT transform & quantization & bit plane extraction" processing. The specific steps include: performing an MDCT (Modified Discrete Cosine Transform) on the audio signal to obtain a frequency domain coefficient vector, and then using 4-bit scalar uniform quantization; after quantization, each value is extracted into a bit plane to obtain four binary planes. Where 'l' represents the bit level, arranged from high to low as bit plane 1 to bit plane 4; each bit plane Each bit plane is a sequence of binary random variables, which is used as the input source data for the Slepian-Wolf encoder of LDPC; an LDPC code is assigned to each bit plane, and its parity-check matrix is... Through Calculate the checksum Finally, only the checksum As the compressed output of this bit plane, the original bit data is sent. Not sent directly, but by Composed of audio bitstream.

[0136] Video signal compression coding: Based on the semantic consistency labels in the semantic prior information, video frames are classified into macroblocks to locate macroblocks in video I-frames that have the same semantic type as the tactile signal; for video macroblocks related to tactile semantics, "DCT transform & quantization & bit plane extraction" processing is performed, specifically including the following steps: first, the DCT coefficients of the image values ​​are obtained through DCT transform, and then 8-bit scalar uniform quantization is selected; after quantization, each value is extracted into a bit plane to obtain 8 binary planes. Where 'l' represents the bit level, arranged from high to low as bit plane 1 to bit plane 8; each bit plane Each bit plane is a sequence of binary random variables, which is used as the input source data for an LDPC-based Slepian-Wolf encoder; an LDPC code is assigned to each bit plane, and its parity-check matrix is... Through Calculate the checksum Finally, only the checksum As the compressed output of this bit plane, the original bit data b is sent. l Not sent directly, but by The semantically related video stream 2 is composed of video I-frame macroblocks that are not related to tactile semantics. HEVC intra-frame coding is used for video B-frames and P-frames. Their encoded output is video stream 1.

[0137] After receiving the heterogeneous bitstream, the decoding end first performs demultiplexing processing to separate the heterogeneous bitstream into video bitstream 1, video bitstream 2, audio bitstream, etc.

[0138] Decoding of tactile signals and semantic prior information: Both decoding processes are the inverse of encoding. Specifically, for the received tactile bitstream, Huffman decoding is first used to recover the encoded tactile sample values ​​and their corresponding timestamps. For missing moments that are not encoded (i.e., within the perceptual dead zone), linear interpolation reconstruction or higher-order interpolation methods are used to reconstruct the missing samples based on the time position and value of adjacent received tactile samples, thereby restoring the complete tactile temporal signal. The tactile reconstruction signal generated in this process needs to undergo perceptual quality assessment to ensure that the reconstructed force feedback curve maintains physical accuracy within the human tactile perception threshold (Weber coefficient k = 0.1). Semantic prior information is directly recovered using Huffman decoding.

[0139] Haptic-assisted reference information generation: By identifying 500 tactile sampling data points that correspond temporally to video and audio frames, the corresponding material categories are extracted. This information is then combined with received semantic prior information to obtain semantically relevant macroblocks in the video frames and material text information corresponding to the audio frames. This text result is input into a cross-modal generator of the corresponding modality, which includes a text-to-video generator and a text-to-audio generator. The text-to-video generator converts the material category text description into a semantically consistent keyframe sequence. The text-to-audio generator converts the material action description into the corresponding acoustic feature waveform. The two cross-modal generators are used to generate video reference signals and audio reference signals, respectively, as auxiliary information in the submodal signal encoding stage. It should be noted that this patent does not involve the design and training of tactile material recognition algorithms or cross-modal generator models, but rather integrates them as system modules based on existing publicly available technologies.

[0140] Before generating the reference signal, the system first performs tactile recognition processing on the reconstructed tactile signal. By analyzing the time-frequency characteristics and transient response patterns of the pressure curve, it accurately identifies the type of interactive material (such as metal, wood, fabric, etc.). This recognition process is implemented using MIT GelSight's Touch-based Material Classification model, which is built on a deep convolutional network and can extract discriminative material features from the tactile signal. The semantic labels generated by the recognition, together with the reconstructed signal, constitute the input of the cross-modal generator, ensuring that the generated reference signal maintains a high degree of consistency with the original tactile experience in terms of physical properties.

[0141] In this example, the following mainstream open-source models are adopted: the tactile recognition classifier adopts the Touch-based Material Classification model from MIT GelSight, which can automatically classify material types using sensor data; the text-to-video generator adopts CogVideo proposed by Tsinghua University, which supports the generation of semantically consistent video reference segments using natural language descriptions; the text-to-audio generator adopts AudioLDM proposed by Oxford, which can convert material class descriptions (such as "rubbing on a rough wooden board") into semantically corresponding audio segments; the generated cross-modal reference signals have structural features that are highly consistent with tactile semantics, and can assist submodals (such as video and audio) in collaborative compression and accuracy recovery after the main modality signal (i.e., tactile) is reconstructed.

[0142] Audio decoding: A Slepian-Wolf decoder based on LDPC is used for decoding. Specifically, based on the current reference audio sample and the contextual temporal structure, the distribution model of the error term between the reference signal and the real audio signal is estimated online using a Laplace distribution, with parameters obtained adaptively from maximum likelihood estimation or historical data. A layer-by-layer decoding strategy is employed to recover the 8-bit plane b of the quantized video. l Specifically, starting from the most significant bit plane, proceed sequentially towards the least significant bit; for each bit plane b l The decoder calculates the log-likelihood ratio of the bit plane by combining the current reference signal value, the decoded higher-level bit planes, and the current error term distribution; then, the decoder uses the LDPC parity-check matrix H... l Seeking satisfaction: binary sequence As b lThe reconstruction process is as follows: If the CRC check of the current layer fails, an additional checksum is requested from the encoder via the feedback channel to achieve progressive decoding. If multiple consecutive decoding attempts fail (e.g., reaching the maximum number of iterations or consecutive errors), an early stop mechanism is initiated to abandon the layer and save computational resources; or the last piece of redundant information is requested, which the decoder utilizes. The inverse matrix is ​​used to recover the bit plane; the audio reconstruction process begins after all bit planes have been decoded, the frequency band coefficients are combined and recovered, and inverse quantization and inverse MDCT transformation are performed to obtain the reconstructed audio data; the reconstructed segment is spliced ​​with the segment that did not use the cross-modal reference signal to restore the complete audio.

[0143] Video Decoding: The bitstream of semantically relevant macroblocks in the I-frame is decoded using an LDPC-based Slepian-Wolf decoder, while the bitstreams of semantically unrelated parts of the I-frame, B-frames, and P-frames are reconstructed using HEVC intra-frame decoding and inter-frame decoding; for example Figure 3 As shown, for the semantically related part of the I-frame, the error term distribution parameters of the reference signal are first estimated online based on the video reference signal and the surrounding spatiotemporal context for subsequent soft decoding; next, a layer-by-layer decoding strategy is used to recover the 8-bit plane b of the quantized video. l Specifically, starting from the most significant bit plane, proceed sequentially towards the least significant bit; for each bit plane b l The decoder calculates the log-likelihood ratio of the bit plane by combining the current reference signal value, the decoded higher-level bit planes, and the current error term distribution; then, the decoder uses the LDPC parity-check matrix H... l Seeking satisfaction: binary sequence As b l The reconstruction is performed; if the CRC check of the current layer fails, an additional checksum is requested from the encoder via the feedback channel to achieve progressive decoding; if multiple consecutive decoding attempts fail (e.g., reaching the maximum number of iterations or consecutive errors), then: an early stop mechanism is initiated, abandoning the layer to save computational resources; or the last piece of redundant information is requested, which the decoder utilizes. The inverse matrix is ​​used to recover the bit plane; after all bit planes are decoded, video reconstruction is performed. First, the frequency band coefficients are combined and recovered, and then inverse quantization and inverse transform processing are performed (using the fast IDCT algorithm to implement the inverse DCT transform) to obtain the reconstructed video pixel blocks; after the reconstructed region is stitched together with the region that did not use the cross-modal reference signal, the complete video frame is restored, and the I-frame reconstruction process is completed.

[0144] The rate-distortion performance of the cross-modal audio and video encoding / decoding methods described in this invention will be compared with traditional audio and video coding schemes to evaluate and measure the effectiveness and accuracy of the methods described in this invention.

[0145] Audio rate-distortion performance for example Figure 8 As shown, the horizontal axis reflects the bit rate required for the audio signal, and the vertical axis is the average noise-mask ratio (ANMR) corresponding to different bit rates. ANMR is a commonly used metric for evaluating audio distortion, and a smaller value is better.

[0146] from Figure 8 It can be seen that the cross-modal audio coding scheme involved in this invention has a smaller average noise-masking ratio and better rate-distortion performance at the same bit rate.

[0147] The rate-distortion performance of video is, for example Figure 9 As shown, the horizontal axis reflects the bit rate required for the video signal; the vertical axis is the peak signal-to-noise ratio (PSNR) corresponding to different bit rates, which is a commonly used metric for video distortion assessment, and the higher the value, the better.

[0148] from Figure 9 It can be seen that the cross-modal video coding scheme involved in this invention has a larger peak signal-to-noise ratio and better rate-distortion performance at the same bit rate.

[0149] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0150] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the present invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A cross-modal source encoding and decoding method based on semantic association, characterized in that, Specifically, the following steps are included: Multiple modal source signals are acquired and input into the encoding terminal. A primary modality is determined based on the quality of service index parameters of each modality, and the remaining modalities are secondary modalities. Based on the type of the main modal signal, a corresponding preset coding strategy is adopted to perform feature-preserving coding processing on the main modal source signal in order to preserve the semantic features necessary for cross-modal decoding; Construct semantic prior information between the main modality and each submodality to describe the semantic labels of the main modality source signal and the semantic consistency relationship between the main modality and the submodality; Based on semantic prior information, the submodal source signal is divided into a semantically relevant signal part and a semantically irrelevant signal part; Perform cross-modal cooperative coding on the semantically related signal portion; Perform standard single-mode coding on semantically irrelevant signal components; A multiplexing mechanism is adopted to merge the encoding results of the main modal source signal, the encoding results of the semantically related signal part, the encoding results of the semantically unrelated signal part, and the semantic prior information to form a heterogeneous bitstream. The heterogeneous bitstream is the final output of the encoding end and is transmitted to the decoding end. At the decoding end, the primary mode source signal is reconstructed and semantic association information is extracted to assist in generating a reference signal for the secondary mode source signal; Based on the reconstruction results of the primary modal source signal and semantic prior information, cross-modal collaborative decoding is performed on the semantically relevant signal part, and standard single-modal decoding is used for the semantically irrelevant signal part to recover the secondary modal source signal.

2. The cross-modal source encoding and decoding method based on semantic association according to claim 1, characterized in that, The acquired source signals for each modality include at least two of the following: video signals, audio signals, tactile signals, and text signals.

3. The cross-modal source encoding and decoding method based on semantic association according to claim 2, characterized in that, Based on the service quality index parameters of each modality, a primary modality is determined, and the remaining modalities are secondary modalities. The specific steps include: Determine the service quality indicator parameters, including latency sensitivity and reliability sensitivity; Delay sensitivity is determined based on the maximum end-to-end transmission delay allowed for the source signal of that mode; The reliability sensitivity is determined based on the maximum allowable packet loss rate of the source signal in that mode; Based on preset weighted scoring rules, a comprehensive score is given for the latency sensitivity and reliability sensitivity of each modality type; Sort all modal types from highest to lowest based on their overall score; The modality with the highest score is designated as the primary modality, and the remaining modalities are designated as secondary modalities.

4. The cross-modal source encoding and decoding method based on semantic association according to claim 3, characterized in that, Constructing semantic prior information between the primary modality and each submodality includes the following steps: Semantic analysis is performed on the main modal source signal to extract its semantic tags. The semantic tags include at least one of object category, material attribute, action type and event state. For each submodal source signal, a semantic consistency relationship is established between it and the main modal source signal. The semantic consistency relationship includes local segment alignment relationship, semantic attribute similarity score and cross-modal semantic mapping function. Local segment alignment relationships are used to identify the correspondence between submodal signal segments and primary modal semantic tags; Semantic attribute similarity scoring is used to quantify the degree of matching between the submodality and the main modality in terms of semantic features; Cross-modal semantic mapping functions are used to describe the transformation relationship of semantic features between different modalities; Semantic tags and semantic consistency relationships are encoded into structured metadata to form semantic prior information.

5. The cross-modal source encoding and decoding method based on semantic association according to claim 4, characterized in that, Preset encoding strategies include: When the main modal source signal is a video signal, a compression algorithm based on inter-frame prediction is used in combination with perceptual quantization and entropy coding techniques. When the main modal source signal is a tactile signal, non-uniform quantization and predictive coding based on Weber's law are used. When the main modal source signal is an audio signal, a modified discrete cosine transform or parameterized coding is used; When the primary modal source signal is a text signal, context-independent compression or semantic keyword extraction is used.

6. The cross-modal source encoding and decoding method based on semantic association according to claim 5, characterized in that, Performing cross-modal cooperative coding on the semantically related signal portion includes the following steps: The semantically relevant signal components are quantized to preserve signal features and control distortion. Convert the quantization result into multiple bit planes; Based on the coding results of the main mode source signal and semantic prior information, Slepian-Wolf coding is performed on each bit plane to generate a parser as the compression result; Slepian-Wolf coding uses LDPC codes, Polar codes, or Turbo codes as channel codes.

7. The cross-modal source encoding and decoding method based on semantic association according to claim 6, characterized in that, Performing standard single-mode coding on semantically irrelevant signal components includes the following steps: When the submodal source signal is a video signal, intra-frame or inter-frame coding is performed using HEVC or H.264 standards. When the submodal source signal is an audio signal, MP3 or AAC standards are used for compression encoding. When the submodal source signal is a tactile signal, an entropy coding method based on the sensory dead zone is adopted.

8. The cross-modal source encoding and decoding method based on semantic association according to claim 7, characterized in that, The primary mode source signal is reconstructed and semantic association information is extracted at the decoding end to assist in generating a reference signal for the secondary mode source signal. This process includes the following steps: Separate the encoding result of the main mode source signal from the received heterogeneous bitstream; The encoded results of the separated main mode source signals are decoded and reconstructed. Semantic prior information is extracted from heterogeneous bitstreams to obtain the main modality semantic labels and the semantic consistency relationship between the main modality and the submodality; The reconstructed primary modal source signal and the extracted semantic prior information are input into the cross-modal generator to generate the reference signal for the secondary modal source signal.

9. The cross-modal source encoding and decoding method based on semantic association according to claim 8, characterized in that, Cross-modal generators include any of the following: conditional variational autoencoders, diffusion models, or cross-modal neural networks.

10. The cross-modal source encoding and decoding method based on semantic association according to claim 9, characterized in that, The recovery of the submodal source signal includes the following steps: Perform cross-modal cooperative decoding on the semantically related signal portion, perform Slepian-Wolf decoding on the received checksum, and recover the semantically related signal content by combining the reference signal; Perform standard single-mode decoding on semantically irrelevant signal components: When the submode is a video signal, it is decoded using a HEVC or H.264 standard decoder. When the submode is an audio signal, it is decoded using an MP3 or AAC standard decoder. When the submodal is a tactile signal, interpolation reconstruction based on the sensory dead zone is used.