Narrowband transmission method and system based on semantic compression and audio and video joint perception

By using semantic compression and joint audio-video perception, semantic information of audio and video is extracted, and layered encapsulation and error protection are performed. This solves the problem of understandability and comprehensibility of audio and video transmission under narrowband conditions, and enables adaptive reconstruction of key information and reliable delivery of emergency instructions.

CN121963737APending Publication Date: 2026-05-01BEIJING IACTIVE NETWORK
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING IACTIVE NETWORK
Filing Date
2026-03-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Under narrowband conditions, existing technologies struggle to ensure the understandability and comprehensibility of audio and video transmission while simultaneously achieving adaptive reconstruction and reliable delivery of semantic information, especially in emergency command scenarios where the reliability of critical information is insufficient.

Method used

By using a method based on semantic compression and joint audio-video perception, audio semantic unit sequences and emergency command semantic information are extracted, scene semantic information is extracted from video signals, and lip-sync residual information is generated using cross-modal prediction. The data is then encapsulated into semantic data packets according to semantic importance, and error protection and scheduled transmission are performed. The receiving end reconstructs the audio and video, calculates semantic reconstruction confidence information, and the sending end adjusts field selection and error protection based on the confidence information.

Benefits of technology

Effectively eliminate audio and video redundancy under narrowband conditions, improve the reliable delivery probability of key semantics, ensure speech intelligibility and video intelligibility, and realize adaptive control of semantic reconstruction and timely and reliable delivery of emergency instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963737A_ABST
    Figure CN121963737A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital information transmission, and discloses a narrowband transmission method and system based on semantic compression and audio and video joint perception, and the method comprises the steps: obtaining a link state parameter, and determining a transmission budget; and extracting the audio semantic unit sequence and extracting the semantic information of the emergency instruction. And scene semantic information is extracted. And performing cross-modal prediction on the mouth shape parameters in the scene semantic information to obtain mouth shape residual error information. Information is divided into multiple layers of semantic layer information according to semantic importance and packaged into a semantic data packet. And performing error protection of different intensities on semantic layer information of different layers, and scheduling and sending according to the sending budget. And analyzing the semantic data packet, and calculating semantic reconstruction confidence information. And adjusting field selection and error protection of the semantic data packet according to the semantic reconstruction confidence information and the link state parameters. According to the invention, the audio and video availability under the narrow-band weak network condition is realized, the audio and video transmission efficiency is improved, and the stability of data transmission is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Narrowband transmission method and system based on semantic compression and joint audio / video sensing Technical Field

[0001] This invention relates to the field of digital information transmission technology, and more specifically, to a narrowband transmission method and system based on semantic compression and joint audio and video perception. Background Technology

[0002] With the increasing prevalence of network bandwidth constraints, packet loss, and latency fluctuations in scenarios such as mobile communication, emergency command, remote inspection, and vehicle networking, traditional pixel-centric video compression and transmission cannot simultaneously ensure understandable voice, understandable video, and low latency experience under narrowband conditions. Therefore, there is a gradual shift towards a semantic-centric transmission approach, which prioritizes the transmission of semantic information that is most critical to the understanding task and implements scheduling when link conditions change to improve availability under narrowband transmission.

[0003] In existing technologies, some patents have proposed video transmission schemes based on semantic communication. For example, invention patent CN115802060A extracts the semantic features of the current video frame and constructs a channel input sequence based on channel bandwidth cost to achieve semantic feature transmission. However, it mainly focuses on video semantic features and channel bandwidth allocation, lacking feedback adjustment for semantic reconstruction confidence at the receiving end. This makes it difficult to trigger deterministic field selection and resynchronization control when semantic reconstruction degrades. Furthermore, it does not establish cross-modal transmission between audio and video, making it difficult to further eliminate audio and video redundancy and improve narrowband efficiency.

[0004] Therefore, it is necessary to design a narrowband transmission method and system based on semantic compression and joint audio and video perception to solve the problems existing in the current technology. Summary of the Invention

[0005] In view of this, the present invention proposes a narrowband transmission method and system based on semantic compression and joint audio and video perception, aiming to solve the problems that key semantics are easily lost and difficult to adaptively reconstruct and recover in audio and video transmission under existing narrowband weak network conditions, resulting in incomprehensible voice, incomprehensible video and unreliable delivery of emergency commands.

[0006] This invention proposes a narrowband transmission method based on semantic compression and joint audio-video perception, comprising: obtaining link state parameters at the transmitting end and determining the transmission budget; extracting an audio semantic unit sequence and emergency instruction semantic information from the audio signal; extracting scene semantic information from the video signal; performing cross-modal prediction on lip-sync parameters in the scene semantic information based on the audio semantic unit sequence to obtain lip-sync residual information; dividing the audio semantic unit sequence, emergency instruction semantic information, scene semantic information, and lip-sync residual information into multiple semantic layers according to semantic importance and encapsulating them into semantic data packets; applying different levels of error protection to the semantic layers and scheduling transmission according to the transmission budget; parsing the semantic data packets at the receiving end, recovering the audio, and reconstructing the video based on the scene semantic information and lip-sync residual information, calculating semantic reconstruction confidence information; adjusting the field selection and error protection of the semantic data packets at the transmitting end according to the semantic reconstruction confidence information and the link state parameters, and sending semantic keyframe data packets for semantic resynchronization when the confidence decreases; and sending semantic keyframe data packets and enhancing error protection when emergency instruction semantic information exists.

[0007] Furthermore, the link status parameters include available throughput, round-trip time, latency jitter, packet loss rate, and burst packet loss indicators; the transmission budget includes payload budget and error correction redundancy budget.

[0008] Furthermore, the error correction redundancy budget includes a first error correction redundancy budget for the first semantic layer information, a second error correction redundancy budget for the second semantic layer information, and a third error correction redundancy budget for the third semantic layer information; the error protection strength is determined by a preset correspondence between redundancy levels and error correction redundancy budgets; when the error correction redundancy budget is insufficient to meet the preset redundancy level, the transmitting end reduces the redundancy level in order from low-importance semantic layer information to high-importance semantic layer information until the error correction redundancy budget is met; when the error correction redundancy budget increases, the transmitting end increases the redundancy level in order from high-importance semantic layer information to low-importance semantic layer information until the preset redundancy level is reached or the error correction redundancy budget is exhausted.

[0009] Furthermore, the field selection includes incremental field sending and full field sending; the incremental field sending is the sending of changed fields relative to the most recent semantic keyframe data packet; the full field sending is the sending of the baseline fields carried by the semantic keyframe data packet.

[0010] Furthermore, the semantic reconstruction confidence information is determined based on the audio recovery success rate, the cross-frame consistency of scene semantic information, and the audio semantic unit sequence and lip-sync reconstruction results.

[0011] Furthermore, the confidence decline condition is that the semantic reconstruction confidence information is lower than a preset confidence threshold and continues for a preset number of scheduling cycles; when the confidence decline condition is met, the sending end switches to full field transmission and increases the error protection strength of highly important semantic layer information.

[0012] Furthermore, the semantic resynchronization includes: when the confidence decline condition is met, the sending end repeatedly sends the semantic keyframe data packet in multiple consecutive scheduling cycles until the semantic reconstruction confidence information recovers to the confidence recovery condition; the confidence recovery condition is that the semantic reconstruction confidence information is not lower than the preset confidence threshold and continues for a preset number of scheduling cycles; after the semantic reconstruction confidence information recovers to the confidence recovery condition, the sending end switches to sending the incremental field.

[0013] Furthermore, when the emergency instruction semantic information indicates an emergency state, the sending end sends the semantic key frame data packet within the current scheduling period, and implements a higher error protection strength than that for non-emergency states on the semantic layer information carrying the emergency instruction semantic information.

[0014] Furthermore, the generation of the lip shape residual information includes: the sending end obtaining predicted lip shape parameters based on the cross-modal prediction; comparing the predicted lip shape parameters with the lip shape parameters to obtain lip shape difference information; when the lip shape difference information meets the residual transmission condition, sending the lip shape difference information as the lip shape residual information in the semantic data packet; when the lip shape difference information does not meet the residual transmission condition, sending the lip shape parameters as alternative lip shape information in the semantic data packet.

[0015] Compared with existing technologies, the advantages of this invention are as follows: The transmitting end forms an executable transmission budget based on link state parameters, ensuring a unified capacity limit for subsequent encoding, protection, and scheduling, thus avoiding disordered congestion and jitter in bandwidth fluctuation and packet loss scenarios. On the content side, instead of focusing on pixel frames, it extracts audio semantic unit sequences and emergency instruction semantic information from the audio, extracts scene semantic information including lip-sync parameters from the video, and further utilizes the audio semantic unit sequences to perform cross-modal prediction of lip-sync parameters, transmitting only lip-sync residual information. This eliminates redundant expressions between audio and video at the transmission level, allowing limited bits to carry the information necessary for understanding under narrowband conditions. On the reliability side, the audio semantic unit sequences, emergency instruction semantic information, and scene semantic information... Lip-sync residual information is encapsulated into semantic data packets according to semantic importance, with different levels of error protection applied to different layers. Then, transmission is scheduled according to the transmission budget, ensuring that critical semantics have a higher delivery probability under high and sudden packet loss, while non-critical semantics can be downgraded to maintain overall continuity. At the receiving end, after recovering audio and reconstructing video using semantic data packets, semantic reconstruction confidence information is calculated and transmitted back. The sending end adjusts field selection and error protection accordingly, achieving closed-loop control. When confidence decreases, semantic keyframe data packets are triggered to achieve semantic resynchronization, quickly correcting accumulated errors from incremental transmission and restoring stable reconstruction. Simultaneously, when emergency command semantic information exists, semantic keyframe data packets are sent first with enhanced error protection, improving the reliable arrival and understandability of emergency commands and critical scenarios under extremely weak network conditions.

[0016] On the other hand, this application also provides a narrowband transmission system based on semantic compression and audio-video joint perception, used to apply the above-mentioned narrowband transmission method based on semantic compression and audio-video joint perception, including: an acquisition unit configured to obtain link state parameters from the transmitting end to determine the transmission budget; extract an audio semantic unit sequence and extract emergency instruction semantic information from the audio signal; extract scene semantic information from the video signal; a prediction unit configured to perform cross-modal prediction on lip-shape parameters in the scene semantic information based on the audio semantic unit sequence to obtain lip-shape residual information; and an encapsulation unit configured to encapsulate the audio semantic unit sequence, emergency instruction semantic information, scene semantic information, and lip-shape residual information. Semantic information is divided into multiple semantic layers based on semantic importance and encapsulated into semantic data packets. Different levels of error protection are applied to the semantic information at different layers, and transmission is scheduled according to the transmission budget. A parsing unit is configured to parse the semantic data packets at the receiving end, recover the audio, reconstruct the video based on the scene semantic information and lip-sync residual information, and calculate semantic reconstruction confidence information. A judgment unit is configured to allow the sending end to adjust the field selection and error protection of the semantic data packets based on the semantic reconstruction confidence information and the link state parameters; when confidence decreases, a semantic keyframe data packet is sent for semantic resynchronization; when emergency instruction semantic information exists, the semantic keyframe data packet is sent and error protection is enhanced.

[0017] It is understandable that the aforementioned narrowband transmission methods and systems based on semantic compression and joint audio-video perception have the same beneficial effects, and will not be elaborated further here. Attached Figure Description

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 is a flowchart of a narrowband transmission method based on semantic compression and audio / video joint sensing provided by an embodiment of the present invention; Figure 2 is a functional block diagram of a narrowband transmission system based on semantic compression and audio / video joint sensing provided by an embodiment of the present invention. Detailed Implementation

[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] In some embodiments of this application, referring to Figure 1, this application proposes a narrowband transmission method based on semantic compression and joint audio-video perception, including: S100: Determining the transmission budget based on link state parameters obtained at the transmitting end. Extracting audio semantic unit sequences and extracting emergency instruction semantic information from the audio signal. Extracting scene semantic information from the video signal.

[0021] S200: Based on the audio semantic unit sequence, perform cross-modal prediction of lip shape parameters in scene semantic information to obtain lip shape residual information.

[0022] S300: The audio semantic unit sequence, emergency instruction semantic information, scene semantic information, and lip-sync residual information are divided into multiple semantic layers according to semantic importance and encapsulated into semantic data packets. Different levels of error protection are applied to the semantic layer information of different layers, and transmission is scheduled according to the transmission budget.

[0023] S400: Based on the semantic data packets parsed at the receiving end, recover the audio and reconstruct the video based on scene semantic information and lip-sync residual information, and calculate the semantic reconstruction confidence information.

[0024] S500: The sending end adjusts the field selection and error protection of semantic data packets based on semantic reconstruction confidence information and link state parameters. When confidence decreases, it sends semantic key frame data packets for semantic resynchronization. When urgent instruction semantic information exists, it sends semantic key frame data packets and enhances error protection.

[0025] Specifically, this embodiment uses bidirectional audio and video communication in a narrowband weak network environment as an application scenario. The sending end and receiving end are deployed on a mobile terminal and a remote scheduling terminal, respectively, and they transmit data through a wireless network or public network link. In this embodiment, both the sending end and the receiving end maintain a unified time base and data sequence number for timing alignment and packet loss identification of semantic data packets. The semantic data packets adopt a "semantic frame header plus semantic payload" structure, where the semantic frame header at least includes a timestamp, sequence number, semantic layer identifier, and a set of field identifiers to support receiver reordering, incremental field concatenation, and semantic keyframe alignment. To avoid ambiguity in terminology and ensure consistency in the terminology of the claims, in this embodiment: scene semantic information is a structured description set of video content, including at least target category information, target location information, target state information, and lip shape parameters. Lip shape parameters are a set of parameters describing the mouth movement state, including at least one or more of the following: mouth opening and closing state, mouth keypoint displacement, or lip shape category. The audio semantic unit sequence is a discrete semantic representation sequence of audio content, preferably obtained by vector quantization mapping of acoustic embedding vectors or output from an edge-side speech recognition model. Emergency instruction semantic information represents whether an emergency instruction has occurred and its semantic category. Multi-layer semantic information includes at least a first semantic layer, a second semantic layer, and a third semantic layer, where the first semantic layer is of high importance, the second semantic layer is of medium importance, and the third semantic layer is of low importance. Lip shape residual information is the difference between the predicted lip shape parameters and the actual lip shape parameters. Alternative lip shape information is the lip shape parameters sent when the lip shape difference information does not meet the residual transmission condition. Semantic reconstruction confidence information is a scalar indicator characterizing the availability of semantic reconstruction at the receiver, preferably ranging from 0 to 1. The semantic keyframe data packet is a data packet containing a full set of fields used to establish or reset the receiver's semantic baseline, and its field set includes at least the baseline fields for scene semantic information and the baseline fields for lip shape parameters.

[0026] In step S100, the sending end obtains link state parameters and determines the transmission budget. The link state parameters are obtained by the sending end within a sliding time window, preferably 300 milliseconds to 2 seconds. The link state parameters include at least available throughput, round-trip time (RTD), latency jitter, packet loss rate, and burst packet loss metric. Available throughput is estimated jointly by the effective transmission volume of the sending end within the sliding time window and the acknowledgment information from the receiving end. RTD is calculated from timestamp probe packets. Latency jitter is calculated from the RTD difference between adjacent probe packets. Packet loss rate is calculated from the sequence number missing ratio. Burst packet loss metric is calculated from the length of consecutively lost data packets. The transmission budget includes a payload budget and an error correction redundancy budget. The payload budget limits the total semantic payload that can be transmitted within the current scheduling period, and the error correction redundancy budget limits the total redundancy introduced by error protection. The scheduling period is preferably 100 milliseconds to 500 milliseconds. Subsequently, the transmitting end processes the audio signal in frames with a fixed frame length, preferably 20 to 40 milliseconds. It extracts audio semantic unit sequences from each frame, which can be obtained from acoustic embedding vectors via vector quantization mapping or from discrete semantic tag sequences output by the edge speech recognition model. Simultaneously, it performs emergency command detection on the audio signal. Emergency command semantic information is used to characterize whether a preset emergency command vocabulary or emergency intent appears. The emergency command vocabulary can include commands such as "emergency," "stop," and "evacuate," or their synonyms. The transmitting end extracts frames from the video signal at a low fundamental sampling frequency and performs target detection and tracking to generate scene semantic information. The low fundamental sampling frequency is preferably 1 to 5 frames per second. When emergency command semantic information exists or a sudden change in target state is detected, the frame extraction frequency can be temporarily increased to update the scene semantic information, provided it does not exceed the transmission budget.

[0027] For example, if the estimated available throughput within a 1-second statistical window is 64 kilobits per second, and the scheduling period is 200 milliseconds, then the total transmission volume that can be carried in this scheduling period is approximately 12.8 kilobits, equivalent to 12,800 bits. The payload budget can be set to 9,600 bits, and the error correction redundancy budget can be set to 3,200 bits. Furthermore, the error correction redundancy budget can be allocated according to semantic importance as a first error correction redundancy budget of 1,920 bits, a second error correction redundancy budget of 960 bits, and a third error correction redundancy budget of 320 bits, so as to prioritize the protection of highly important semantic layer information under budget constraints.

[0028] In step S200, the transmitting end performs cross-modal prediction of lip-shape parameters based on the audio semantic unit sequence and generates lip-shape residual information. Specifically, the transmitting end inputs the audio semantic unit sequence into a cross-modal prediction model to output predicted lip-shape parameters. The cross-modal prediction model can be a lightweight neural network or a sequence mapping-based regression model, and is deployed at the transmitting end to meet real-time requirements. Subsequently, the predicted lip-shape parameters are compared with the lip-shape parameters extracted from the video signal to obtain lip-shape difference information. To reduce the transmission burden under narrowband conditions, this embodiment sets residual transmission conditions: when the lip-shape difference information is within a preset difference threshold range and continuously meets a preset number of frames, only the lip-shape difference information is transmitted as lip-shape residual information. When the lip-shape difference information exceeds the preset difference threshold range or continuously fails to meet the preset number of frames, the transmitting end transmits the lip-shape parameters as substitute lip-shape information and simultaneously increases the error protection strength of lip-shape related fields to resist lip-shape drift caused by packet loss. The preset difference threshold and preset frame number are configured by the system during deployment. Preferably, the preset frame number is 1 to 3 frames, so as to reduce the lip-sync field bitrate in steady state and switch to more robust lip-sync parameter transmission in time during sudden changes.

[0029] In step S300, the sending end performs layered encapsulation of semantic information and implements error protection and budget scheduling. Specifically, the sending end classifies audio semantic unit sequences and emergency instruction semantic information into high-importance semantic layer information, scene semantic information and lip-sync residual information into medium-importance semantic layer information, and optional refinement fields into low-importance semantic layer information. This layering reflects the "necessity for the understanding task" and "contribution to reconstruction quality." The sending end encapsulates each layer of semantic layer information into semantic data packets and applies different levels of error protection to different layers, with the error protection strength determined by the redundancy level. A preset correspondence exists between the redundancy level and the error correction redundancy budget: when the error correction redundancy budget is sufficient, a high redundancy level is used for high-importance semantic layer information, a medium redundancy level for medium-importance semantic layer information, and a low redundancy level for low-importance semantic layer information. When the error correction redundancy budget is insufficient to meet the current redundancy level, the sending end reduces the redundancy level from low-importance semantic layer information to high-importance semantic layer information until the budget is met, prioritizing the reliable delivery of key semantic information. Subsequently, the sending end schedules the transmission of semantic data packets according to the payload budget, prioritizing the transmission of high-importance semantic layer information, followed by medium-importance semantic layer information, and finally low-importance semantic layer information. When emergency instruction semantic information exists, the sending end either interrupts the transmission of semantic keyframe data packets or increases the transmission ratio of high-importance semantic layer information within the current scheduling cycle to improve transmission reliability and update timeliness in emergency situations.

[0030] In step S400, the receiving end parses the semantic data packets and reconstructs the audio and video, while simultaneously calculating semantic reconstruction confidence information. The receiving end first reorders and times-aligns the semantic data packets according to the semantic frame header, and then decodes high-importance semantic layer information based on the semantic layer identifier to recover the audio. Audio recovery includes generating a speech waveform based on the audio semantic unit sequence using a speech synthesis model, or generating speech text first and then synthesizing it into a speech waveform. Video reconstruction includes generating a synthetic video frame sequence or a playable presentation stream based on scene semantic information, generating a structured presentation based on target category and target location information, and updating the motion state of the lip-sync region based on lip-sync residual information or substitute lip-sync information. When the semantic keyframe data packet arrives, the semantic baseline is reset with its full set of fields. The receiving end calculates semantic reconstruction confidence information, which is determined based on at least the following factors: audio recovery success rate, cross-frame consistency of scene semantic information, and consistency between the audio semantic unit sequence and the lip-sync reconstruction result. The audio recovery success rate can be determined by a combination of missing audio semantic unit sequences and decoding availability. The cross-frame consistency of scene semantic information can be determined by the continuity of target tracking and the field missing rate. Consistency can be determined by the matching of the pronunciation rhythm and lip-sync update rhythm corresponding to the audio semantic unit sequence. The receiver sends the semantic reconstruction confidence information back to the transmitter as input for the transmitter's adaptive control. The feedback can be carried through a control channel multiplexed with the receiver's confirmation information.

[0031] In step S500, the transmitting end performs closed-loop adaptive control based on semantic reconstruction confidence information and link state parameters. The transmitting end sets the confidence decline condition to a threshold value below the semantic reconstruction confidence information for a preset number of scheduling cycles. When the confidence decline condition is met, the transmitting end switches to full-field transmission and sends semantic keyframe data packets for semantic resynchronization, while simultaneously increasing the error protection strength of highly important semantic layer information. Semantic keyframe data packets are repeatedly sent over multiple consecutive scheduling cycles until the semantic reconstruction confidence information meets the confidence recovery condition, at which point the transmission switches back to incremental field transmission. The confidence recovery condition is that the semantic reconstruction confidence information is not lower than the preset threshold value for a preset number of scheduling cycles. Furthermore, when the emergency instruction semantic information indicates an emergency state, the transmitting end sends semantic keyframe data packets within the current scheduling cycle and implements a higher error protection strength than in non-emergency states for the semantic layer information carrying the emergency instruction semantic information, ensuring reliable delivery of key instructions and key scenario semantics in narrowband weak networks.

[0032] Understandably, by forming an end-to-end availability guarantee chain through cross-modal lip-sync residual transmission, semantic hierarchical unequal error protection and budget scheduling, semantic reconstruction confidence closed-loop feedback, and semantic keyframe resynchronization, audio and video redundancy can be reduced in weak network and narrowband environments, and limited transmission resources can be prioritized for high-importance semantics, thereby improving voice and video intelligibility. Simultaneously, semantic keyframe resynchronization triggered by confidence decline suppresses the accumulation of errors in incremental transmission, and further enhances error protection and priority transmission when emergency commands occur, making the arrival probability and update timeliness of key semantics higher in emergency scenarios.

[0033] In this embodiment, the sending end uses a sliding time window to statistically analyze link state parameters and periodically updates the transmission budget, ensuring that semantic layered encapsulation, error protection, and scheduled transmission are all subject to a unified budget constraint. Link state parameters include available throughput, round-trip time (RTD), latency jitter, packet loss rate, and burst packet loss metric. Available throughput characterizes the upper limit of the effective payload that can be stably carried per unit time and can be estimated from the amount of effective payload sent by the sending end within the time window and the acknowledgment information from the receiving end. RTD characterizes the end-to-end interaction latency and can be calculated from timestamp probe packets and acknowledgment information. Latency jitter characterizes the degree of latency fluctuation and can be statistically obtained from the difference in RTD between adjacent probe packets. Packet loss rate characterizes the proportion of lost data packets and can be statistically obtained from the proportion of missing sequence numbers. Burst packet loss metric characterizes the severity of consecutive packet losses and can be statistically obtained from the length of consecutively missing data packet segments. The transmission budget includes a payload budget and an error correction redundancy budget. The payload budget is used to limit the total amount of transmittable load of semantic layer information in the current scheduling period, while the error correction redundancy budget is used to limit the total amount of redundancy introduced by error protection, thereby changing the strength control of error protection from "arbitrary configuration" to "deterministic configuration constrained by budget".

[0034] Furthermore, to achieve unequal error protection for multi-layer semantic information, this embodiment subdivides the error correction redundancy budget into a first error correction redundancy budget, a second error correction redundancy budget, and a third error correction redundancy budget, corresponding to the first, second, and third semantic layer information, respectively. To avoid ambiguity in the description of "error protection strength," this embodiment uses redundancy levels for discretization. Redundancy levels characterize the degree of redundancy introduced by error protection and correspond one-to-one with the error correction coding configuration. The transmitting end pre-stores a correspondence table between redundancy levels and error correction redundancy overhead. This table is either configured offline or factory-fixed, and the redundancy level is selected at runtime based on the error correction redundancy budget. Specifically, when the error correction redundancy budget meets the target redundancy level requirements for each semantic layer information, the transmitting end configures a higher redundancy level for the first semantic layer information than for the second semantic layer information, and a higher redundancy level for the second semantic layer information than for the third semantic layer information. When the error correction redundancy budget is insufficient to meet the target redundancy level requirement, the transmitter reduces the redundancy level in order from low-importance semantic layer information to high-importance semantic layer information. Specifically, it prioritizes reducing the redundancy level of the third semantic layer information, then the second semantic layer information, and finally the first semantic layer information, until the overall redundancy overhead does not exceed the error correction redundancy budget. Conversely, when the error correction redundancy budget increases, the transmitter increases the redundancy level in order from high-importance semantic layer information to low-importance semantic layer information. Specifically, it prioritizes restoring the redundancy level of the first semantic layer information, then the second semantic layer information, and finally the third semantic layer information, until the target redundancy level is reached or the error correction redundancy budget is exhausted. Through this sequential rule, the adjustment of error protection strength is deterministic, avoiding random preemption of redundancy resources or frequent oscillations under link fluctuations.

[0035] Understandably, by binding link state parameters to the transmission budget and further decomposing the error correction redundancy budget into each semantic layer, the determination and adjustment of error protection strength form a "budget-constrained hierarchical redundancy allocation mechanism." This ensures that, even in narrowband weak networks, with packet loss and sudden packet loss, the reliable delivery of highly important semantic layer information is always prioritized, reducing the degradation of audio and video availability caused by the loss of critical semantics. Based on the overall scheme, this embodiment further provides a configuration basis for semantic hierarchical error protection and a stable error correction resource supply mechanism for semantic confidence closed-loop adaptive control, making it easier for the overall scheme to maintain stable and controllable service quality during long-term operation in weak networks.

[0036] In this embodiment, the sending end adopts a field selection mechanism of "semantic keyframe baseline plus incremental update" to reduce narrowband transmission load and suppress the accumulation of semantic reconstruction errors. Specifically, the sending end maintains a field set and field identifier for each type of scene semantic information. The field set includes at least a target category field, a target location field, a target state field, and a lip-sync parameter field. The sending end also maintains a baseline state cache corresponding to the semantic keyframe data packet to record the baseline field values ​​carried by the most recent semantic keyframe data packet. When generating a semantic data packet, the sending end first compares the currently extracted scene semantic information with the baseline state cache field by field to determine whether the field has changed. When a field changes, it is marked as a changed field and forms an incremental field set. Then, incremental field transmission is performed, that is, only the incremental field set is encapsulated into the semantic data packet and sent. The determination of "field change" can use different rules depending on the field type: for the target location field, it can be based on whether the position offset exceeds a preset position offset threshold. For the target state field, it can be based on whether the state enumeration has changed or whether the state confidence has crossed a preset confidence change threshold. For lip-sync parameter fields, the transmission can be based on whether the lip-sync residual information meets the residual transmission conditions or whether alternative lip-sync information is enabled. Correspondingly, full field transmission is used to send semantic keyframe data packets, which involves completely encapsulating and sending the baseline fields from the field set to establish or reset the semantic baseline at the receiving end. After the semantic keyframe data packet transmission is completed and confirmed by the receiving end, the sending end updates the baseline state cache, ensuring that subsequent incremental field transmissions always use the "most recent semantic keyframe data packet" as a reference, avoiding incremental splicing errors caused by inconsistent reference states.

[0037] In this embodiment, the receiving end calculates semantic reconstruction confidence information during the process of parsing semantic data packets and reconstructing audio and video. This confidence information is used to characterize the availability of the current semantic reconstruction and drive the adaptive control of the sending end. The semantic reconstruction confidence information is determined based on at least the following factors: audio recovery success rate, cross-frame consistency of scene semantic information, and consistency between the audio semantic unit sequence and the lip-sync reconstruction result. Audio recovery success rate characterizes the intelligibility and continuity of the audio recovery, and can be determined based on the missing proportion of the audio semantic unit sequence, the length of consecutive missing segments, and the continuity of the decoded output. Cross-frame consistency of scene semantic information characterizes the stability of scene semantics, and can be determined based on whether target tracking is continuous, whether target identifiers switch frequently, and the missing rate of incremental fields in consecutive frames. Consistency characterizes the semantic consistency at the audio-visual synchronization level, and can be determined based on whether the pronunciation rhythm corresponding to the audio semantic unit sequence matches the update rhythm of the lip-sync reconstruction result, and whether the lip-sync parameter update time maintains a stable deviation range from the audio semantic boundary. To facilitate engineering implementation, the receiving end can map the above factors to confidence sub-indicators and perform weighted synthesis to obtain semantic reconstruction confidence information. This information is then transmitted back to the sending end in each scheduling cycle, allowing the sending end to adjust field selection and error protection based on confidence change trends. The thresholds and weights of each sub-indicator are deployment parameters that can be configured according to the business requirements of "voice intelligibility priority" or "video intelligibility priority".

[0038] Understandably, the mechanism of "establishing a semantic baseline by sending all fields and reducing continuous load by sending incremental fields" simplifies semantic data transmission in narrowband weak networks, reduces bandwidth consumption caused by invalid and repeated fields, and reduces ambiguity in incremental updates by comparing fields with semantic keyframes as a reference. Simultaneously, by incorporating audio recovery quality, scene semantic stability, and audio-visual consistency into a unified evaluation criterion through semantic reconstruction confidence information, the sending end can control based on the actual reconstruction effect at the receiving end rather than solely on link statistics, thus more promptly detecting implicit degradation such as "deterioration in reconstruction quality despite no obvious packet loss." This provides a quantifiable feedback indicator system for the semantic confidence closed-loop adaptive mechanism in the overall solution, giving semantic keyframe resynchronization and error protection adjustments clearer triggering criteria, thereby further improving the stability and controllability of the overall solution during long-term operation in weak networks.

[0039] In this embodiment, the sending end executes a "confidence-triggered semantic resynchronization and emergency priority protection" control strategy based on the semantic reconstruction confidence information returned by the receiving end and the link status parameters obtained by the sending end. This strategy aims to suppress the accumulation of errors in incremental field transmission and ensure reliable delivery of emergency commands under narrowband weak network conditions. The sending end uses a scheduling period as the control granularity, preferably between 100 milliseconds and 500 milliseconds. In each scheduling period, the sending end receives semantic reconstruction confidence information and maintains a continuous counter to determine confidence decline and confidence recovery conditions. A preset confidence threshold and a preset number of scheduling periods are deployment parameters. The preset confidence threshold distinguishes between "semantic reconstruction available" and "semantic reconstruction degraded," and the preset number of scheduling periods avoids false triggering caused by instantaneous fluctuations. Preferably, the preset number of scheduling periods is between 1 and 5 to balance trigger sensitivity and stability. To ensure clear triggering rules, the sending end adopts the following strategy when determining the confidence decline condition: when the semantic reconstruction confidence information is continuously lower than the preset confidence threshold and continuously reaches the preset number of scheduling periods, the confidence decline condition is deemed to be met. When the confidence information of semantic reconstruction is continuously not lower than the preset confidence threshold and continues to reach the preset scheduling cycle number, it is determined that the confidence recovery condition is met.

[0040] When the confidence decline condition is met, the sender immediately switches to full field transmission and enters the semantic resynchronization process. Full field transmission is achieved by sending semantic keyframe data packets. These packets carry a set of reference fields used to establish or reset the receiver's semantic baseline, including at least the reference fields for scene semantic information and lip-sync parameters. Simultaneously, the sender increases the error protection strength of high-importance semantic layer information to improve the arrival probability of key semantics under packet loss and sudden packet loss conditions. During semantic resynchronization, the sender repeatedly sends semantic keyframe data packets over multiple consecutive scheduling cycles. The number of repeated transmission cycles is constrained by the transmission budget and link state parameters, and an upper limit is preferably set to avoid prolonged occupation of the payload budget. After receiving the semantic keyframe data packets, the receiver resets the semantic baseline and recalculates the semantic reconstruction confidence information, which is then transmitted back. The sender continuously monitors the semantic reconstruction confidence information. When the confidence recovery condition is met, the sender exits the semantic resynchronization process and switches back to incremental field transmission, thereby restoring the updated changed fields with reference to the most recent semantic keyframe data packet, reducing subsequent transmission load and maintaining reconstruction continuity. To prevent oscillations caused by frequent switching, the sending end can set a minimum dwell period after exiting semantic resynchronization. During this minimum dwell period, incremental field transmission is prioritized unless the confidence decline condition is met again or an emergency state is detected.

[0041] In this embodiment, to meet the requirement of "priority arrival of critical instructions and key semantics" in emergency command scenarios, the sending end implements an emergency priority protection strategy when the emergency instruction semantic information indicates an emergency state. Specifically, once the emergency instruction semantic information is extracted and determined to be an emergency state, the sending end directly sends semantic key frame data packets within the current scheduling cycle to quickly establish or refresh the semantic baseline at the receiving end. Simultaneously, the semantic layer information carrying the emergency instruction semantic information is subject to a higher error protection strength than in non-emergency states. This increased error protection strength can be achieved by increasing the redundancy level or increasing the allocation ratio of error correction redundancy budget in highly important semantic layer information, thereby improving the reliable delivery probability of emergency instruction semantic information under bandwidth constraints and packet loss fluctuations. After the emergency state ends, the sending end can gradually restore the error protection configuration to that of the non-emergency state based on semantic reconstruction confidence information and link state parameters to avoid long-term occupation of redundant resources leading to a decline in the transmission quality of other semantic layer information.

[0042] Understandably, by using a semantic resynchronization mechanism triggered by semantic reconstruction confidence information, the system can promptly switch to full-field transmission and repeatedly transmit semantic keyframe data packets when errors accumulate due to incremental field transmission, scene semantics are lost, or audio-visual consistency degrades. This allows for a quicker recovery of the receiver's semantic baseline and a return to low-load incremental field transmission, achieving closed-loop control where degradation is detectable, mismatch is correctable, and recovery is determineable. Simultaneously, an emergency priority protection strategy enhances the error protection strength of semantic layer information carrying emergency instruction semantics in emergency situations and forces the transmission of semantic keyframe data packets in the current scheduling cycle. This ensures that key instructions and key scene semantics maintain a higher arrival probability and timely update even under extremely weak network conditions. This embodiment strengthens the stability and controllability of the overall solution in long-term weak network operation and emergency scenarios.

[0043] In this embodiment, the transmitting end employs a "cross-modal prediction plus residual adaptive transmission" generation and transmission mechanism for lip-sync related fields to reduce the transmission load of lip-sync information under narrowband conditions and maintain audio-visual consistency. The transmitting end acquires the audio semantic unit sequence and video-side lip-sync parameters within the same time alignment window. The lip-sync parameters are constituent fields of scene semantic information, preferably including at least one of mouth opening / closing state, mouth keypoint displacement, and lip-sync category. The transmitting end inputs the audio semantic unit sequence into a cross-modal prediction model to obtain predicted lip-sync parameters. The cross-modal prediction model can be an end-side lightweight sequence prediction model, whose input is a series of consecutive audio semantic units and their temporal position information, and whose output is predicted lip-sync parameters aligned with the video frame extraction time. Subsequently, the transmitting end compares the predicted lip-sync parameters with the actual lip-sync parameters to obtain lip-sync difference information. This lip-sync difference information characterizes the amount, category, or direction of difference between the predicted lip-sync parameters and the actual lip-sync parameters, and serves as the basis for determining whether to transmit the residual.

[0044] To ensure the "residual transmission condition" remains unambiguous, this embodiment limits the residual transmission condition to at least one of the following: First, the lip-shape difference information does not exceed a preset difference threshold and is continuously satisfied for a preset number of frames. Second, although the lip-shape difference information exceeds the preset difference threshold, it belongs to a preset tolerable difference type and is continuously satisfied for a preset number of frames. The preset difference threshold, preset number of frames, and preset tolerable difference type are deployment parameters. The preset number of frames is preferably 1 to 3, used to increase the applicability of residual transmission during the steady-state pronunciation phase and to promptly switch to alternative lip-shape information transmission during the sudden lip-shape change phase. Specifically, when the lip-shape difference information meets the residual transmission condition, the sending end encapsulates the lip-shape difference information as lip-shape residual information and sends it into the semantic data packet. At the receiving end, the lip-shape residual information is used to correct the predicted lip-shape parameters to obtain the reconstructed lip-shape parameters used for lip-shape presentation. Correspondingly, when the lip-shape difference information does not meet the residual transmission condition, the sending end does not send the lip-shape residual information but instead encapsulates the lip-shape parameters as alternative lip-shape information and sends them into the semantic data packet. At the receiving end, the alternative lip-sync information is used to directly update the lip-sync rendering, and upon arrival, it can be used to recalibrate subsequent prediction correction links, thereby preventing the continuous accumulation of errors. To further improve stability under extremely weak network conditions, the sending end can simultaneously increase the error protection strength of the semantic layer information containing lip-sync related fields when enabling the transmission of alternative lip-sync information, in order to reduce audio-visual mismatch caused by the loss of key lip-sync fields.

[0045] Understandably, by using cross-modal prediction driven by audio semantic unit sequences, lip-sync information is transmitted in residual form during most steady-state phases, significantly reducing the sustained bitrate overhead of the lip-sync field and allocating more limited transmission resources to urgent command semantic information and critical scene semantic information. Simultaneously, through an adaptive switching mechanism controlled by residual transmission conditions, alternative lip-sync information is promptly transmitted during phases of drastic lip-sync changes or increased prediction deviations, enabling rapid correction of the received lip-sync presentation and thus improving audio-visual consistency and lip-sync reconstruction stability. This embodiment enhances the overall solution's technical effectiveness in ensuring stable presentation of understandable speech and understandable video.

[0046] In summary, the sending end formulates an executable transmission budget based on link state parameters, ensuring a unified capacity limit for subsequent encoding, protection, and scheduling, thus avoiding disordered congestion and jitter under bandwidth fluctuations and packet loss scenarios. On the content side, instead of focusing on pixel frames, audio semantic unit sequences are extracted from audio, along with emergency command semantic information. Scene semantic information, including lip-sync parameters, is extracted from video. Furthermore, the audio semantic unit sequences are used to perform cross-modal prediction of lip-sync parameters, transmitting only lip-sync residual information. This eliminates redundant expressions between audio and video at the transmission layer, allowing limited bits to prioritize the transmission of understandable information under narrowband conditions. On the reliability side, audio semantic unit sequences, emergency command semantic information, scene semantic information, and lip-sync residual information are layered and encapsulated into semantic data packets according to semantic importance. Different levels of error protection are applied to different layers, and transmission is scheduled according to the transmission budget. This ensures that critical semantics have a higher delivery probability under high packet loss and burst packet loss conditions, while non-critical semantics can be degraded to ensure overall continuity. After recovering audio and reconstructing video using semantic data packets at the receiving end, semantic reconstruction confidence information is calculated and transmitted back. The sending end then adjusts field selection and error protection accordingly, achieving closed-loop control. When confidence decreases, semantic keyframe data packets are triggered to achieve semantic resynchronization, quickly correcting accumulated errors from incremental transmission and restoring stable reconstruction. Simultaneously, when emergency command semantic information exists, semantic keyframe data packets are sent first, and error protection is enhanced, improving the reliable arrival and understandability of emergency commands and critical scenarios under extremely weak network conditions.

[0047] Based on another preferred embodiment of the above, referring to Figure 2, this embodiment provides a narrowband transmission system based on semantic compression and audio / video joint perception, used to apply the above-mentioned narrowband transmission method based on semantic compression and audio / video joint perception, including: a data acquisition unit configured to determine a transmission budget based on link status parameters acquired from the transmitting end; extracting an audio semantic unit sequence and extracting emergency instruction semantic information from the audio signal; and extracting scene semantic information from the video signal.

[0048] The prediction unit is configured to perform cross-modal prediction of lip shape parameters in scene semantic information based on audio semantic unit sequences to obtain lip shape residual information.

[0049] The encapsulation unit is configured to divide the audio semantic unit sequence, emergency instruction semantic information, scene semantic information, and lip-sync residual information into multiple semantic layers according to semantic importance and encapsulate them into semantic data packets. Different levels of error protection are applied to the semantic layer information at different levels, and transmission is scheduled according to the transmission budget.

[0050] The parsing unit is configured to parse semantic data packets based on the receiving end, recover audio, reconstruct video based on scene semantic information and lip-sync residual information, and calculate semantic reconstruction confidence information.

[0051] The judgment unit is configured to allow the sender to adjust the field selection and error protection of semantic data packets based on semantic reconstruction confidence information and link state parameters. When confidence decreases, it sends semantic keyframe data packets for semantic resynchronization. When urgent instruction semantic information exists, it sends semantic keyframe data packets and enhances error protection.

[0052] It is understandable that the aforementioned narrowband transmission methods and systems based on semantic compression and joint audio-video perception have the same beneficial effects, and will not be elaborated further here.

[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A narrowband transmission method based on semantic compression and joint audio / video sensing, characterized in that, include: Based on the link status parameters obtained from the transmitting end, the transmission budget is determined; audio semantic unit sequences and emergency instruction semantic information are extracted from the audio signal; scene semantic information is extracted from the video signal. Based on the audio semantic unit sequence, cross-modal prediction is performed on the lip-shape parameters in the scene semantic information to obtain lip-shape residual information; the audio semantic unit sequence, emergency instruction semantic information, scene semantic information, and lip-shape residual information are divided into multiple semantic layers according to semantic importance and encapsulated into semantic data packets; different levels of error protection are applied to the semantic layer information of different layers, and transmission is scheduled according to the transmission budget; based on the receiver parsing the semantic data packets, the audio is recovered, and the video is reconstructed based on the scene semantic information and lip-shape residual information, and semantic reconstruction confidence information is calculated; the transmitter adjusts the field selection and error protection of the semantic data packets according to the semantic reconstruction confidence information and the link state parameters, and sends semantic keyframe data packets for semantic resynchronization when the confidence decreases; when emergency instruction semantic information exists, semantic keyframe data packets are sent and error protection is enhanced.

2. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 1, characterized in that, The link status parameters include available throughput, round-trip time, latency jitter, packet loss rate, and burst packet loss indicators; the transmission budget includes payload budget and error correction redundancy budget.

3. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 2, characterized in that, The error correction redundancy budget includes a first error correction redundancy budget for first semantic layer information, a second error correction redundancy budget for second semantic layer information, and a third error correction redundancy budget for third semantic layer information. The error protection strength is determined by a preset correspondence between redundancy levels and error correction redundancy budgets. When the error correction redundancy budget is insufficient to meet the preset redundancy level, the transmitting end reduces the redundancy level in order from low-importance semantic layer information to high-importance semantic layer information until the error correction redundancy budget is met. When the error correction redundancy budget increases, the transmitting end increases the redundancy level in order from high-importance semantic layer information to low-importance semantic layer information until the preset redundancy level is reached or the error correction redundancy budget is exhausted.

4. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 1, characterized in that, The field selection includes incremental field sending and full field sending; the incremental field sending is the sending of changed fields relative to the most recent semantic keyframe data packet; the full field sending is the sending of the baseline fields carried in the semantic keyframe data packet.

5. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 4, characterized in that, The semantic reconstruction confidence information is determined based on the audio recovery success rate, the cross-frame consistency of scene semantic information, and the audio semantic unit sequence and lip-sync reconstruction results.

6. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 5, characterized in that, The confidence decline condition is that the confidence information of the semantic reconstruction is lower than the preset confidence threshold and continues for a preset number of scheduling cycles; when the confidence decline condition is met, the sending end switches to full field transmission and increases the error protection strength of the high-importance semantic layer information.

7. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 6, characterized in that, The semantic resynchronization includes: when the confidence decline condition is met, the sending end repeatedly sends the semantic keyframe data packet in multiple consecutive scheduling cycles until the semantic reconstruction confidence information recovers to the confidence recovery condition; the confidence recovery condition is that the semantic reconstruction confidence information is not lower than the preset confidence threshold and continues for a preset number of scheduling cycles; after the semantic reconstruction confidence information recovers to the confidence recovery condition, the sending end switches to sending the incremental field.

8. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 6, characterized in that, When the emergency instruction semantic information indicates an emergency state, the sending end sends the semantic key frame data packet within the current scheduling period and implements a higher error protection strength than that for non-emergency states on the semantic layer information carrying the emergency instruction semantic information.

9. The narrowband transmission method based on semantic compression and joint audio / video perception according to claim 8, characterized in that, The generation of the lip shape residual information includes: the sending end obtaining predicted lip shape parameters based on the cross-modal prediction; comparing the predicted lip shape parameters with the lip shape parameters to obtain lip shape difference information; when the lip shape difference information meets the residual transmission condition, sending the lip shape difference information as the lip shape residual information in the semantic data packet; when the lip shape difference information does not meet the residual transmission condition, sending the lip shape parameters as alternative lip shape information in the semantic data packet.

10. A narrowband transmission system based on semantic compression and joint audio / video sensing, used to apply the narrowband transmission method based on semantic compression and joint audio / video sensing as described in any one of claims 1-9, characterized in that, include: The acquisition unit is configured to obtain link status parameters from the transmitting end to determine the transmission budget; extract audio semantic unit sequences and emergency instruction semantic information from the audio signal; and extract scene semantic information from the video signal. The prediction unit is configured to perform cross-modal prediction on lip-sync parameters in the scene semantic information based on the audio semantic unit sequences to obtain lip-sync residual information. The encapsulation unit is configured to divide the audio semantic unit sequences, emergency instruction semantic information, scene semantic information, and lip-sync residual information into multiple semantic layers according to semantic importance and encapsulate them into semantic data packets; apply different levels of error protection to the semantic layer information at different levels and schedule transmission according to the transmission budget. The parsing unit is configured to parse the semantic data packets based on the receiving end, recover the audio, reconstruct the video based on the scene semantic information and lip-sync residual information, and calculate semantic reconstruction confidence information. The judgment unit is configured to adjust the field selection and error protection of the semantic data packet based on the semantic reconstruction confidence information and the link state parameters. When the confidence decreases, the semantic key frame data packet is sent for semantic resynchronization. When the emergency instruction semantic information exists, the semantic key frame data packet is sent and error protection is enhanced.

Citation Information

Patent Citations

  • Semantic communication video transmission method and related equipment

    CN115802060A

  • Incremental retransmission method and device for monitoring image in electric power scene

    CN117082301A

  • Emergency talkback method with wide-band and narrow-band simultaneous voice access capability

    CN119676660A

  • Plug-and-play wireless high-definition audio and video transmission method and system

    CN121000923A

  • Mining intelligent voice and video interactive communication method and system

    CN121284187A