A semantically aware video compression, transmission method, apparatus, and electronic device at the edge.

By using edge-end video compression technology, semantic mask matrix and residual data are used to identify edge intersection states, implement semantic gradient splitting and adaptive quantization, solve the video coding pollution problem, and achieve efficient video transmission and reconstruction.

CN122069350BActive Publication Date: 2026-07-31HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +2
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-04-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, topological mismatch between pixel-level semantic contours and fixed-size coding block grids in edge-end video compression leads to coding pollution and reduced target edge sharpness, failing to effectively eliminate prediction residual redundancy, especially at high-confidence semantic boundaries where coding pollution is severe.

Method used

By acquiring the binarized semantic mask matrix and transform residual data, the edge crossover state of the largest coding tree unit is identified, semantic gradient recursive splitting is implemented to generate a set of sub-coding units, and coupling priority scores are calculated. The quantization parameters are adaptively adjusted to achieve adaptive mapping and stream priority control of the video encoder.

Benefits of technology

Eliminating coding contamination at the physical structure level preserves the sharpness of the foreground contours in the reconstructed video, reduces the ineffective computational burden on edge processors, and ensures low-latency video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069350B_ABST
    Figure CN122069350B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image communication technology and discloses a video compression and transmission method, apparatus, and electronic device with edge semantic awareness. The method includes: acquiring a video frame to be transmitted, a semantic mask matrix, semantic confidence data, and transform residual data; physically aligning the semantic mask matrix with the maximum coding tree unit grid of the video encoder; identifying target grids in an edge-intersection state by calculating the pixel spatial variance metric within the grid; adaptively partitioning the target grid into sub-coding units based on semantic gradients to generate a set of sub-coding units matching the object contour; and calculating a coupling priority score by combining the transform residual data and semantic confidence data to guide the bitrate control module in implementing adaptive mapping of quantization parameters. This invention locks the foreground contour sharpness through a semantic edge pre-partitioning mechanism and eliminates coding prediction redundancy using priority scoring, thereby achieving a correlation between local spatiotemporal physical entropy change and semantic importance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image communication technology, and in particular relates to a video compression and transmission method, apparatus and electronic device with edge semantic awareness. Background Technology

[0002] Currently, in edge video compression applications, processing devices typically use convolutional neural networks to identify specific targets in video frames and adjust the quantization parameters of the encoder in different spatial regions based on the identification results, thereby preserving key visual features under limited bandwidth conditions. As video acquisition resolution evolves towards ultra-high definition, the topological mismatch between pixel-level semantic edges and hard-coded block grids has become an objective obstacle restricting transmission performance. Since the basic processing unit of mainstream video coding standards is a fixed-size square grid, while the semantic contour of the target has irregular geometric features, physical overlap deviation occurs when the two are aligned in the spatial coordinate system. This mismatch leads to coding pollution in the boundary grid area, which not only causes bit resources to flow to useless background pixels, but also causes high-frequency components of the target edge to be lost in the quantization process, thereby reducing the contour sharpness of the reconstructed video.

[0003] Existing image communication architectures suffer from underlying grid physical limitations and inadequate control methods. For example, Chinese invention patent CN117793352B discloses a video compression method, device, and storage medium based on semantic understanding. Compared to this patent, which relies on extracting semantic tags to calculate content complexity and rate of change to adjust quantization parameters, the macroscopic semantic feature-driven quantization adjustment software mechanism implicitly depends on the assumption that the target semantic object and the basic coding block naturally fit each other. When facing foreground targets with highly irregular geometric distortions at the edge, the macroscopic software control logic struggles to penetrate the underlying rigid grid topology barrier. Essentially, it remains constrained by the allocation of bits within the traditional fixed square grid, inevitably causing physical-level coding pollution at high-confidence semantic boundaries. The core presuppositions are fundamentally mismatched with the actual irregular contour boundary conditions, making it impossible to eliminate prediction residual redundancy.

[0004] Therefore, how to drive the coding grid to perform topology reconstruction through semantic features and synchronize the evaluation logic of the underlying coding state to achieve deep coupling of video data in physical topology and spatiotemporal energy envelope has become the technical problem to be solved by this invention. Summary of the Invention

[0005] This invention provides a semantically aware transmission method at the edge, comprising the following steps: The system acquires the video frame to be transmitted and the corresponding binary semantic mask matrix, and extracts the semantic confidence data corresponding to the binary semantic mask matrix in real time from the lightweight inference model deployed on the edge side. At the same time, it synchronously extracts the transform residual data representing the prediction error density from the hardware register inside the video encoder. The binary semantic mask matrix is ​​physically aligned with the preset maximum coding tree cell grid in the video encoder in the spatial coordinate system; Traverse each maximum coding tree unit in the current video frame, calculate the spatial variance metric of the corresponding mask pixels within the grid area of ​​each maximum coding tree unit, and identify whether the maximum coding tree unit is in an edge crossing state that crosses the physical boundary between the foreground and the background based on the spatial variance metric. For each maximum coding tree unit in the edge crossing state, a semantic gradient-based sub-coding unit recursive split is implemented. The maximum coding tree unit in the edge crossing state is adaptively divided into a set of sub-coding units that match the object contour, so as to limit the overflow of prediction bits to the background pixel area and lock the edge sharpness of the foreground contour at the underlying bitstream syntax structure level. The weighted product of the transformation residual data and the semantic confidence data is calculated to generate a coupled priority score that characterizes the local spatiotemporal physical entropy change, thereby establishing the relationship between motion complexity and semantic importance. The quantization parameter correction amount of the video encoder bitrate control module is calculated based on the coupling priority score, and the bitrate control module is guided to perform adaptive mapping of quantization parameters on the sub-coding unit set to generate video bitstream.

[0006] Preferably, the step of dividing each maximum coding tree unit in the edge-crossing state into sub-coding units further includes: monitoring the temporal stability data of the corresponding maximum coding tree unit during the coding prediction process; when the temporal stability data is higher than a preset stability threshold and the transform residual data is lower than a preset noise threshold, blocking the transform path in the sub-coding unit set through controller instructions, and guiding the video encoder to start the skip mode to reduce the invalid computing load of the edge-side hardware.

[0007] Preferably, the step of calculating the spatial variance metric of the corresponding mask pixels within the grid area of ​​each maximum coding tree unit further includes: if the spatial variance metric of a certain maximum coding tree unit is identified as 0 and the mean value of the corresponding mask pixels is a background attribute, then the maximum coding tree unit is marked as a static background block; grid dormancy processing is implemented on the static background block, the motion vector search path of the static background block is forcibly blocked and only the DC component is retained.

[0008] Preferably, the video frames to be transmitted are extracted by a lightweight transformer model deployed at the edge; the binarized semantic mask matrix is ​​generated by the attention map transformation of the lightweight transformer model; the semantic confidence data is determined by the maximum probability distribution output of the last layer classifier of the lightweight transformer model; the lightweight transformer model is obtained by parameter pruning of a preset deep learning network.

[0009] Preferably, coupling priority scoring satisfy: Where S is the coupling priority score, and α is the preset spatial domain weight coefficient. The data represents semantic confidence scores, and β is a preset time-domain weighting coefficient. The steps for adaptively mapping the quantization parameters of the sub-coding unit set, which are derived from the transform residual data extracted from the hardware registers inside the video encoder and where the sum of α and β is 1, specifically include: establishing a nonlinear mapping function that couples the priority score with the quantization step size of the video encoder; when an object contour is identified to contain a high-confidence foreground and unpredictable deformation, the quantization parameters of the sub-coding unit set are reduced according to the nonlinear mapping function, and an enhanced coding action for high-frequency textures is triggered, so that the allocation of coding resources is tilted towards the object contour.

[0010] Preferably, the step of implementing adaptive mapping of quantization parameters for the sub-coding unit set further includes: real-time monitoring of available bandwidth data of the physical channel on the edge side; using the available bandwidth data as the global constraint boundary of the bitrate control module; within the global constraint boundary, adjusting the quantization parameters of each sub-coding unit set to ensure that the bit allocation ratio of the corresponding foreground region in the sub-coding unit set is always higher than the bit allocation ratio of the corresponding background region, and the transmission of the video stream is achieved through a stream priority control protocol; marking data packets containing object outlines in the video stream as high-priority streams and marking data packets containing background regions as low-priority streams; and, in the case of network congestion, actively discarding data packets of low-priority streams through the gateway device to ensure that the end-to-end transmission latency of high-priority streams is less than 50ms.

[0011] Preferably, the method further includes a cloud feedback step: the receiving end receives the video bitstream and reconstructs the video frame; calculates the edge detection operator response value of the reconstructed video frame and sends the edge detection operator response value back to the edge side; the edge side dynamically adjusts the recursion depth of the largest coding tree unit in the edge crossing state to implement sub-coding unit partitioning based on the received edge detection operator response value, and the step of implementing adaptive mapping of quantization parameters for the sub-coding unit set is implemented by a lookup table method: multiple sets of quantization parameter offset tables are pre-stored; according to the numerical range to which the coupling priority score belongs, a matching quantization parameter offset is selected from multiple sets of quantization parameter offset tables; the original quantization parameters of the sub-coding unit set are corrected using the matching quantization parameter offset, and the corrected quantization parameters for entropy coding are finally generated.

[0012] An edge-end semantic-aware video compression method is provided to implement an edge-end semantic-aware transmission method.

[0013] An edge semantic awareness device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement an edge semantic awareness transmission method.

[0014] An edge-end semantic-aware electronic device and an edge-end semantic-aware transmission method are applied in the electronic device for video bitstream transmission in security monitoring systems, video conferencing terminals, or intelligent traffic monitoring systems.

[0015] Compared with existing technologies, the present invention, a semantically aware video compression and transmission method, apparatus, and electronic device at the edge, has the following advantages: 1. In edge semantic perception, the geometric gradient variance of the semantic mask is used to forcibly intervene in the coding tree unit partitioning depth of the video encoder, so that the physical topology of the video grid and the real semantic boundary of the target are spatially decoupled. This solves the problem of physical dimension misalignment between pixel-level semantic contours and rigid coding block grids. This grid pre-tearing mechanism, which is directly driven by semantic features, short-circuites the original rate-distortion optimization search path of the encoder, so that the target edge is precisely crushed into extremely small coding units. This eliminates the coding pollution phenomenon of target edge bits overflowing into the background area at the physical structure level, and maintains the edge sharpness of the foreground contour in the reconstructed video.

[0016] 2. By synchronously extracting the absolute transformation error and physical quantities native to the video encoder hardware and multiplicatively coupling them with semantic confidence at the register level, this invention constructs a priority evaluation system characterizing the spatiotemporal energy envelope. This eliminates the dual computational redundancy between external motion estimation and internal coding prediction in edge devices, enabling the system to forcibly block the residual transformation and quantization path through hardware-level instructions when a high-confidence target is identified as physically stationary or in a predictable translational state. This allows the core computing power to accurately converge to the local spatiotemporal region that truly triggers physical entropy change, reducing the ineffective computational burden on the edge processor and enhancing the stability of power consumption control.

[0017] 3. By using a coupled priority score generated jointly by semantic space confidence and temporal prediction residuals, the bitrate control module of the video encoder is dynamically guided to achieve adaptive attenuation of the video bitstream based on the physical prediction difficulty. When a target area is identified to contain a high-confidence foreground and unpredictable deformation, a high-sensitivity coding action is triggered to capture high-frequency textures. When the target is identified to be in a stable motion condition, the encoder is forced to execute a skip mode. This mechanism works in deep collaboration with the stream priority control of the transmission protocol to effectively suppress the risk of transmission buffer overflow caused by the surge of background redundant data in sudden dynamic scenes, ensuring low-latency response of end-to-end video transmission in a restricted network environment. Attached Figure Description

[0018] Figure 1 This is the overall flowchart of the edge-end semantic-aware video compression and transmission method of the present invention; Figure 2 This is the logic diagram of adaptive mapping of quantization parameters for coupling priority and code stream generation in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0020] It should be noted that all directional and positional terms used in this invention, such as: up, down, left, right, front, back, vertical, horizontal, inner, outer, top, bottom, transverse, longitudinal, center, etc., are only used to explain the relative positional relationship and connection between components in a specific state (as shown in the accompanying drawings). They are only for the convenience of describing this invention and do not require that this invention be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention. In addition, the descriptions of "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated.

[0021] In the description of this invention, unless otherwise explicitly specified and limited, the terms installation, connection, and linking should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections; they can refer to direct connections or indirect connections through an intermediate medium; they can refer to the internal connection of two components. For those skilled in the art, the specific meaning of the above terms in this invention can be understood in conjunction with the specific circumstances.

[0022] In the description of this specification, references to the terms "an embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example, and the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0023] A semantically aware transmission method at the edge includes the following steps: The system acquires the video frame to be transmitted and the corresponding binary semantic mask matrix, and extracts the semantic confidence data corresponding to the binary semantic mask matrix in real time from the lightweight inference model deployed on the edge side. At the same time, it synchronously extracts the transform residual data representing the prediction error density from the hardware register inside the video encoder. The binary semantic mask matrix is ​​physically aligned with the preset maximum coding tree cell grid in the video encoder in the spatial coordinate system; Traverse each maximum coding tree unit in the current video frame, calculate the spatial variance metric of the corresponding mask pixels within the grid area of ​​each maximum coding tree unit, and identify whether the maximum coding tree unit is in an edge crossing state that crosses the physical boundary between the foreground and the background based on the spatial variance metric. For each maximum coding tree unit in the edge crossing state, a semantic gradient-based sub-coding unit recursive split is implemented. The maximum coding tree unit in the edge crossing state is adaptively divided into a set of sub-coding units that match the object contour, so as to limit the overflow of prediction bits to the background pixel area and lock the edge sharpness of the foreground contour at the underlying bitstream syntax structure level. The weighted product of the transformation residual data and the semantic confidence data is calculated to generate a coupled priority score that characterizes the local spatiotemporal physical entropy change, thereby establishing the relationship between motion complexity and semantic importance. The quantization parameter correction amount of the video encoder bitrate control module is calculated based on the coupling priority score, and the bitrate control module is guided to perform adaptive mapping of quantization parameters on the sub-coding unit set to generate video bitstream.

[0024] Preferably, the step of dividing each maximum coding tree unit in the edge-crossing state into sub-coding units further includes: monitoring the temporal stability data of the corresponding maximum coding tree unit during the coding prediction process; when the temporal stability data is higher than a preset stability threshold and the transform residual data is lower than a preset noise threshold, blocking the transform path in the sub-coding unit set through controller instructions, and guiding the video encoder to start the skip mode to reduce the invalid computing load of the edge-side hardware.

[0025] Preferably, the step of calculating the spatial variance metric of the corresponding mask pixels within the grid area of ​​each maximum coding tree unit further includes: if the spatial variance metric of a certain maximum coding tree unit is identified as 0 and the mean value of the corresponding mask pixels is a background attribute, then the maximum coding tree unit is marked as a static background block; grid dormancy processing is implemented on the static background block, the motion vector search path of the static background block is forcibly blocked and only the DC component is retained.

[0026] Preferably, the video frames to be transmitted are extracted by a lightweight transformer model deployed at the edge; the binarized semantic mask matrix is ​​generated by the attention map transformation of the lightweight transformer model; the semantic confidence data is determined by the maximum probability distribution output of the last layer classifier of the lightweight transformer model; the lightweight transformer model is obtained by parameter pruning of a preset deep learning network.

[0027] Preferably, the coupling priority score S satisfies: Where S is the coupling priority score, and α is the preset spatial domain weight coefficient. The data represents semantic confidence scores, and β is a preset time-domain weighting coefficient. The steps for adaptively mapping the quantization parameters of the sub-coding unit set, which are derived from the transform residual data extracted from the hardware registers inside the video encoder and where the sum of α and β is 1, specifically include: establishing a nonlinear mapping function that couples the priority score with the quantization step size of the video encoder; when an object contour is identified to contain a high-confidence foreground and unpredictable deformation, the quantization parameters of the sub-coding unit set are reduced according to the nonlinear mapping function, and an enhanced coding action for high-frequency textures is triggered, so that the allocation of coding resources is tilted towards the object contour.

[0028] Preferably, the step of implementing adaptive mapping of quantization parameters for the sub-coding unit set further includes: real-time monitoring of available bandwidth data of the physical channel on the edge side; using the available bandwidth data as the global constraint boundary of the bitrate control module; within the global constraint boundary, adjusting the quantization parameters of each sub-coding unit set to ensure that the bit allocation ratio of the corresponding foreground region in the sub-coding unit set is always higher than the bit allocation ratio of the corresponding background region, and the transmission of the video stream is achieved through a stream priority control protocol; marking data packets containing object outlines in the video stream as high-priority streams and marking data packets containing background regions as low-priority streams; and, in the case of network congestion, actively discarding data packets of low-priority streams through the gateway device to ensure that the end-to-end transmission latency of high-priority streams is less than 50ms.

[0029] Preferably, the method further includes a cloud feedback step: the receiving end receives the video bitstream and reconstructs the video frame; calculates the edge detection operator response value of the reconstructed video frame and sends the edge detection operator response value back to the edge side; the edge side dynamically adjusts the recursion depth of the largest coding tree unit in the edge crossing state to implement sub-coding unit partitioning based on the received edge detection operator response value, and the step of implementing adaptive mapping of quantization parameters for the sub-coding unit set is implemented by a lookup table method: multiple sets of quantization parameter offset tables are pre-stored; according to the numerical range to which the coupling priority score belongs, a matching quantization parameter offset is selected from multiple sets of quantization parameter offset tables; the original quantization parameters of the sub-coding unit set are corrected using the matching quantization parameter offset, and the corrected quantization parameters for entropy coding are finally generated.

[0030] An edge-end semantic-aware video compression method is provided to implement an edge-end semantic-aware transmission method.

[0031] An edge semantic awareness device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement an edge semantic awareness transmission method.

[0032] An edge-end semantic-aware electronic device and an edge-end semantic-aware transmission method are applied in the electronic device for video bitstream transmission in security monitoring systems, video conferencing terminals, or intelligent traffic monitoring systems.

[0033] Example 1: In the current intelligent traffic monitoring system, under the condition of edge-end compressed transmission of all-weather video streams at complex intersections, the roadside electronic equipment continuously acquires high-resolution video frames to be transmitted. Limited by the available bandwidth of the physical channel, the conventional image communication mechanism based on fixed square maximum coding tree unit grids generates physical overlap deviations in the spatial coordinate system when processing the semantic contours of irregular geometric features. Topological mismatch leads to coding pollution in the boundary grid area, causing bit resources to overflow into the background pixel area. At the same time, it causes the high-frequency components of vehicle contour edges to be lost in the quantization process, increasing the probability of congestion in the end-to-end transmission buffer. The edge-end semantic perception device acquires the video frames to be transmitted and their corresponding binary semantic mask matrices, and extracts the semantic confidence data corresponding to the binary semantic mask matrices in real time from the lightweight inference model deployed on the edge side. It also synchronously extracts transform residual data characterizing the prediction error density from the hardware registers inside the video encoder. The device physically aligns the binarized semantic mask matrix with the preset maximum coding tree unit grid within the video encoder in spatial coordinate system. The device traverses each maximum coding tree unit in the current video frame, calculates the spatial variance metric of the corresponding mask pixels within each maximum coding tree unit grid area, and identifies the edge crossing state of the maximum coding tree unit that crosses the physical boundary between the foreground and background based on the spatial variance metric. For each maximum coding tree unit in the edge crossing state, the device recursively splits the sub-coding units according to the semantic gradient, dividing the maximum coding tree unit in the edge crossing state into a set of sub-coding units that match the contour of the object. The semantic edge pre-division mechanism replaces the encoder's native rate distortion optimization search path.

[0034] To suppress encoding pollution at the physical topology level by preventing target boundary bits from overflowing into the background region, edge electronic devices employ a heterogeneous system-on-a-chip architecture. The video encoder and lightweight inference model interact with each other via a shared memory-mapped bus to extract transform residual data. At the end of each maximum coding tree unit pipeline processing cycle, the device reads the absolute transform error and register address values ​​from the motion estimation hardware accelerator. While reading the hardware accelerator data, the device opens a circular buffer with a storage depth of 15 frames in the on-chip static random access memory, extracts the forward motion vector magnitude of the maximum coding tree unit corresponding to each frame in the buffer, and calculates the arithmetic mean of these 15 magnitude data as temporal stability data through shift accumulation operation. When the logic judgment unit finds that the temporal stability data is less than or equal to a preset stability threshold of 2 pixels precision, and the extracted transform residual data is lower than the preset noise floor baseline, when the controller issues a command to block the transform path in the sub-coding unit set, before the entropy coding stage starts, the device writes a preset skip mode flag bit to the coding tree unit configuration register through the standard video driver interface, directly truncates the clock signals of the discrete cosine transform calculation module and the quantization module in the corresponding spatial coordinates at the hardware level, so that the video encoder retains only the DC component and the forward predicted motion vector in the sub-coding unit set area, and cuts off the residual data generation path in the physical circuit.

[0035] After obtaining the set of sub-encoding units of the matched object contour, the device calculates the transform residual data. With semantic confidence data The weighted product generates a coupling priority score S representing the local spatiotemporal physical entropy change, and the specific mathematical relation is set as follows: Where S is the coupling priority score, and α is the preset spatial domain weight coefficient. The data represents semantic confidence scores, and β is a preset time-domain weighting coefficient. To transform the residual data, and assuming the sum of α and β is 1, based on the coupling priority score S, the device calculates the quantization parameter correction amount of the video encoder rate control module. This guides the rate control module to adaptively map the quantization parameters of the sub-coding unit set. When an object contour is identified as containing a high-confidence foreground and exhibiting deformation, the device reduces the quantization parameters of the sub-coding unit set according to the established nonlinear mapping function, enhancing high-frequency texture coding and guiding coding resources towards the object contour. If the spatial variance metric value corresponding to a maximum coding tree unit is identified as 0, and the mean value of the corresponding mask pixels represents a background attribute, the device marks this maximum coding tree unit as a static background block. The device then puts the grid of the static background block into hibernation, blocks the motion vector search path of the static background block, and retains the DC component, constructing a motion complex... The correlation index between noise and semantic importance is used. When the video stream enters the stream priority control protocol scheduling stage, the device monitors the available bandwidth data of the physical channel on the edge side in real time and sets the available bandwidth data as the global constraint boundary of the bitrate control module. Within the global constraint boundary, the device adjusts the quantization parameters of each sub-coding unit set so that the bit allocation ratio of the corresponding foreground region in the sub-coding unit set is greater than the bit allocation ratio of the corresponding background region. The device marks data packets containing object outlines in the video stream as high-priority streams and data packets containing background regions as low-priority streams. Under network congestion, the gateway device discards data packets of low-priority streams and locks the edge sharpness of the foreground outline at the underlying bitstream syntax structure level to maintain the end-to-end transmission latency of high-priority streams below 50ms.

[0036] Example 2: This example addresses the bandwidth fluctuations and loss of high-frequency edge features in the edge compression transmission of all-weather video streams at complex intersections. An edge computing simulation test platform is constructed to verify the semantic perception and encoding compression mechanism of the target object. This test platform accesses an open-source dataset of urban traffic monitoring, simulating a wireless communication channel with dynamic fluctuations from 10Mbps to 50Mbps. Gaussian white noise with a signal-to-noise ratio of 20dB is superimposed on the original video stream to reproduce sensor acquisition disturbances under physical electromagnetic conditions. The controlled parameters set for the experiment are the spatial domain weighting coefficient α and the temporal domain weighting coefficient β. The core engineering consideration for setting these parameters is to balance the geometric fidelity of the semantic contour with the smoothness of the video temporal prediction. When the speed of the traffic target in the video frame exceeds a preset speed threshold, to avoid a surge in prediction residuals caused by rapid displacement, the temporal domain weighting coefficient β tends towards the upper limit of its range. Under normal uniform speed conditions, α is set to 0.6 and... The parameter combination is set to 0.4 as the normal median.

[0037] The experimental group adopted the complete technical solution of this invention; control group one adopted the fixed square maximum coding tree unit grid partitioning mechanism of the conventional high-efficiency video coding standard; control group two removed the semantic edge pre-partitioning mechanism based on the experimental group, and only retained the adaptive mapping of quantization parameters; control group three set out-of-range weight parameters, i.e. α is 0.9 and β is 0.1. The test platform input a noisy traffic video stream with a resolution of 1920×1080 to each group. The device acquired the video frames to be transmitted and their corresponding binary semantic mask matrices, calculated the spatial variance metric of each maximum coding tree unit, and when processing the 150th frame of video, the device measured the spatial variance metric of a certain maximum coding tree unit at the edge of the target vehicle to be 412.5. It identified the edge intersection state and triggered semantic gradient recursive splitting, generating a set of sub-coding units matching the object contour. After acquiring the original data containing noise, the device extracted the transform residual data for this set of sub-coding units. The semantic confidence score is 125.6. It is 0.85, based on the mathematical relationship. The coupling priority score S was calculated to be 50.75. Based on the established nonlinear mapping function, the device reduced the quantization parameter of the high-confidence foreground from the initial 32 to 22, so that the quantization resources were tilted towards the foreground. At the same time, the dormant spatial variance metric was 0 and the mean was represented by the grid of background attributes, and the compressed video stream was output.

[0038] Under congested network conditions with available bandwidth fluctuating to 15Mbps, decoded video data from each receiving end was extracted and key performance indicators were compared. Control group 1, without processing by the method of this invention, suffered from coding pollution due to topology mismatch, resulting in a peak signal-to-noise ratio (SNR) of the foreground vehicle dropping to 28.4dB and an end-to-end transmission delay climbing to 124ms. Control group 2 had a SNR of 31.2dB, while the experimental group achieved a SNR of 36.8dB and a stable transmission delay of 42ms. The comparison between the experimental group and control group 2 confirmed that the technical gain generated by the cascade mechanism of pre-topology splitting and subsequent quantization parameter mapping is greater than the sum of the effects of a single module, exhibiting a synergistic effect. As the available bandwidth gradually decreased from 15Mbps to 5Mbps, the transmission delay of the experimental group... The latency remained between 45ms and 49ms. In the out-of-range control group 3, the transmission latency suddenly increased to 185ms when the bandwidth dropped to 8Mbps. This nonlinear inflection point confirmed that the excessively high spatial domain weight caused the temporal domain prediction to fail and led to a surge in bitstream. The optimal working window of α=0.6 and β=0.4 was confirmed. The above quantitative comparison results show that under the physical conditions of superimposed Gaussian white noise and dynamic bandwidth constraints, the coding framework jointly driven by the semantic edge pre-division mechanism and the coupling priority scoring suppressed the bit overflow of static background blocks and guided network resources to converge to the foreground high-frequency texture region. This scheme reshapes the topology of the bottom maximum coding tree unit, cuts off the interference path of background noise on object contour coding, and establishes a strong correlation model between local spatiotemporal physical entropy change and semantic importance.

[0039] Example 3: In the case of a smart traffic monitoring system processing high-resolution video frames to be transmitted at the roadside in a low-light environment at night, the physical boundary of the vehicle outline merges with the dark background. Conventional image communication devices lack deterministic geometric segmentation criteria and quantization step size adjustment benchmarks, causing bit allocation ambiguity when the maximum coding tree unit crosses the target boundary, leading to disordered consumption of coding resources and increasing the congestion probability of the physical channel on the edge side. The edge-side semantic perception device acquires the video frame to be transmitted and the corresponding binary semantic mask matrix, and maps it to the maximum coding tree unit grid with a size of 64 by 64 pixels inside the video encoder. To determine the recognition criteria for the edge crossing state, the device traverses 4096 masks within the maximum coding tree unit. For each mask pixel, the arithmetic mean of all mask pixel values ​​is calculated, and the squares of the differences between each mask pixel value and the arithmetic mean are accumulated. This sum is divided by the total number of pixels to obtain a spatial variance metric. When the spatial variance metric is greater than a preset variance activation threshold, the device determines that the maximum coding tree unit is in an edge-crossing state. For this state, the device splits the maximum coding tree unit into sub-coding units based on semantic gradients. Specifically, the device calculates the differences between adjacent pixels in the binarized semantic mask matrix to generate a semantic gradient matrix. Along the extreme value lines of the semantic gradient matrix, the maximum coding tree unit in the edge-crossing state is split into four 32x32 pixel primary sub-blocks. The device recalculates the spatial variance metric of each primary sub-block. If the spatial variance of a certain primary sub-block is greater than a preset variance activation threshold, the device determines that the maximum coding tree unit is in an edge-crossing state. If the variance metric remains greater than zero, it continues to be split into secondary sub-blocks of 16 x 16 pixels or 8 x 8 pixels until the spatial variance metric of all terminal sub-blocks converges to zero. Terminal sub-blocks with a spatial variance metric of zero constitute pure foreground or background sub-coding units, which are combined to form a set of sub-coding units that match the object contour. This transforms the abstract semantic boundary into a deterministic quadtree physical topology that the video encoder can process. When performing sub-coding unit partitioning on the largest coding tree unit at the edge intersection, the device executes a coordinate mapping forced quadtree splitting algorithm to extract the two-dimensional spatial coordinates of all mask pixels with extreme values ​​greater than zero from the semantic gradient matrix in the binary semantic mask matrix, which are then combined to form a contour coordinate set. When traversing the current largest coding tree unit, the geometric center coordinates of the unit and the current grid side length are extracted, and each coordinate point in the contour coordinate set is compared. If there is a target coordinate point that falls within the spatial physical area limited by the geometric center coordinates and the current grid side length, the quadtree splitting flag of the current largest coding tree unit is forcibly set to the active state, driving the video encoder to divide into four primary sub-blocks with half the side length. The spatial coordinate comparison and splitting flag setting steps are recursively performed on each primary sub-block until the physical area of ​​the generated terminal sub-block no longer contains any coordinate points in the contour coordinate set, or the side length of the terminal sub-block reaches the preset 8 x 8 pixel bottom hardware limit size. The terminal sub-blocks with zero spatial variance metric value are then combined to form a sub-coding unit set.

[0040] After obtaining the set of sub-coding units, the device extracts semantic confidence data from the lightweight inference model. Extract transform residual data from hardware registers The coupling priority score S is calculated based on mathematical relationships. To eliminate the blindness of quantization parameter setting, the device establishes a nonlinear mapping function between the coupling priority score S and the quantization parameter correction amount ΔQ. The specific mathematical relationship is set as ΔQ=γ. ln(S+1), where ΔQ is the quantization parameter correction amount, γ is the preset quantization gain coefficient, and S is the coupling priority score. According to this deterministic operator, when the object outline contains a high-confidence foreground, the device reduces the quantization parameter of the corresponding foreground sub-coding unit according to the calculated negative vectorization parameter correction amount ΔQ, driving the encoder to inject more bitstream into the region. At the same time, the device extracts the motion vector of the static background block, calculates the magnitude and direction dispersion of the motion vector to generate a divergence value. When it is determined that the divergence value is lower than the preset environmental disturbance threshold, the encoding mode of the corresponding grid is locked to the skip mode and only the DC component is retained. The above edge-aware compression mechanism reshapes the underlying encoding topology of the video frame through the quantized spatial variance measure and the deterministic recursive splitting algorithm. It uses a logarithmic nonlinear mapping function to directly convert the spatiotemporal physical entropy change into the quantization parameter control boundary, cuts off the path of invalid background area encroaching on physical bandwidth at the underlying data stream level, and maintains the edge sharpness and transmission timeliness of the high-priority stream under the limited channel.

[0041] Example 4: When the intelligent traffic monitoring system faces the initial physical deployment of edge-side equipment at a new monitored intersection, the edge-end semantic perception device initiates an offline benchmark calibration procedure for a specific scenario. The device controls the image sensor to continuously acquire a local standard test video sequence of a preset duration in an uncompressed state. It extracts the static noise floor parameters of the environmental background area and the average motion vector amplitude of the traffic target area from the local standard test video sequence. The device calculates the ratio of the total number of dynamic foreground pixels in the local scene to the total number of pixels in the video frame, generating a scene dynamic activity index. Based on the inverse proportional mapping relationship between the scene dynamic activity index and the spatial domain weight coefficient α and the temporal domain weight coefficient β, the device calculates the initial calibration value. When the scene dynamic activity index is greater than the preset activity scale, the device decreases the value of α and increases the value of β proportionally to offset the prediction residual increment caused by violent motion. Finally, the device writes the calibrated α and β into the hardware memory of the video encoder bitrate control module. When calculating the coupling priority score S, the device synchronously extracts the transform residual data from the hardware register. Before participating in the weighted product operation, the maximum and minimum values ​​are normalized based on the maximum full-load bit width extreme value of the hardware register. The values ​​of the weighted product spatial domain weight coefficient α and temporal domain weight coefficient β are constrained by the real-time output of the device's dynamic activity calibration procedure for the operating scene. The device generates a scene activity dispersion measure by statistically analyzing the ratio of the total number of foreground pixels with motion vector magnitude greater than zero within a past 25-frame video sliding time window to the total number of pixels in the video frames to be transmitted. When the value is lower than the threshold of 0.3, the parameter combination of α = 0.7 and β = 0.3 is written to the underlying register of the video encoder bitrate control module. When the value is greater than or equal to 0.3, the parameter combination is updated to α = 0.4 and β = 0.6, so that the weight coefficient values ​​are bound to the objectively measurable physical motion load.

[0042] After obtaining the above calibration parameters, the device conducts an on-site closed-loop calibration process for the environmental disturbance threshold. During the idle period when there are no target vehicles passing by, the device continuously acquires pure background spatial images, calculates the mean of the spatial variance of the mask pixels of the largest coding tree unit corresponding to all pure background spatial images, and simultaneously extracts the inherent thermal noise variance of the optical sensor. The device establishes the product of the mean of the spatial variance of the mask pixels and the inherent thermal noise variance as the environmental disturbance threshold applicable to the current specific physical channel and illumination conditions. This calibration process fills the system's judgment benchmark database based on the objective observation data of the local physical environment, so that the logical fulcrum of the subsequent grid dormancy judgment is anchored to the determined measurement facts, and the underlying data flow path of the device enters a steady-state operation state with clear numerical boundaries.

[0043] Example 5: After the intelligent traffic monitoring system completes physical deployment at the edge nodes, an offline optimization procedure is initiated for the spatial domain weighting coefficient α and the temporal domain weighting coefficient β. The edge semantic perception device uses maximizing the peak signal-to-noise ratio of the foreground contour and constraining the end-to-end transmission latency to be less than 50ms as the multi-dimensional optimization objective. A standard test video set containing 10,000 frames of dynamic traffic images is imported into the system. The device divides the standard test video set into a stable state subset consisting of targets traveling at constant speeds and a sudden change state subset consisting of targets accelerating, decelerating, and changing lanes. The device sets the test step size for parameter traversal to 0.1, driving α to increase from 0.1 to 0.9. The synchronous adjustment β is decreased from 0.9 to 0.1. While traversing each set of values ​​in the parameter matrix, the device calculates the peak signal-to-noise ratio (PSNR) and end-to-end transmission delay for the two subsets. Measurement data shows that when processing the stationary subset, the PSNR of the foreground contour monotonically increases with increasing α. However, after α exceeds 0.6, the quantization bitstream in the background region surges, causing the end-to-end transmission delay to exceed the 50ms constraint. Simultaneously, when processing the abrupt state subset, if β is less than 0.4, the prediction residual caused by large displacement cannot be compensated, resulting in frame continuity breaks. Based on the aforementioned objective observation limits, the device extracts the intersection coordinates that satisfy all hard physical constraints, setting α to 0.6 and... The value of 0.4 is established as the normal optimization output value and written into the underlying register of the video encoder.

[0044] To address the evolving physical environment, including lighting and weather conditions, a periodic environmental disturbance reference model is established to calibrate the environmental disturbance threshold. Every 24 hours, the edge-end semantic perception device wakes up the fine-tuning module. Under preset low-light or rain / fog conditions, it extracts all grids marked as static background blocks within the last hour. The device calculates the arithmetic mean of the motion vector magnitudes of each static background block to generate an environmental disturbance reference value. When the motion vector divergence of a background grid in the current video frame exceeds a tolerance limit of 1.5 times the environmental disturbance reference value, the device determines that the divergence change originates from a sudden environmental disturbance. The device then unlocks the skip mode of that grid, guiding the video encoder to initiate a full motion vector search to update the reference frame content. Based on the recalculated spatial variance metric, the device determines whether to restore the grid to its dormant state. This online fault-tolerant mechanism ensures that the environmental disturbance judgment criteria in the system's underlying data link always align with the drift trajectory of objective physical environment parameters.

[0045] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit of this application and the scope of protection of this invention, and all of these forms are within the protection scope of this application.

Claims

1. A method for edge-end semantic-aware transmission, characterized in that, Includes the following steps: The system acquires the video frame to be transmitted and the corresponding binary semantic mask matrix, and extracts the semantic confidence data corresponding to the binary semantic mask matrix in real time from the lightweight inference model deployed on the edge side. At the same time, it synchronously extracts the transform residual data representing the prediction error density from the hardware register inside the video encoder. The binary semantic mask matrix is ​​physically aligned with the preset maximum coding tree cell grid in the video encoder in the spatial coordinate system; Traverse each maximum coding tree unit in the current video frame, calculate the spatial variance metric of the corresponding mask pixels within the grid area of ​​each maximum coding tree unit, and identify whether the maximum coding tree unit is in an edge crossing state that crosses the physical boundary between the foreground and the background based on the spatial variance metric. For each maximum coding tree unit in the edge crossing state, a semantic gradient-based sub-coding unit recursive split is implemented. The maximum coding tree unit in the edge crossing state is adaptively divided into a set of sub-coding units that match the object contour, so as to limit the overflow of prediction bits to the background pixel area and lock the edge sharpness of the foreground contour at the underlying bitstream syntax structure level. The weighted product of the transformation residual data and the semantic confidence data is calculated to generate a coupled priority score that characterizes the local spatiotemporal physical entropy change, thereby establishing the relationship between motion complexity and semantic importance. The quantization parameter correction amount of the video encoder bitrate control module is calculated based on the coupling priority score, and the bitrate control module is guided to perform adaptive mapping of quantization parameters on the sub-coding unit set to generate video bitstream.

2. The edge-end semantic-aware transmission method of claim 1, wherein, The step of dividing each maximum coding tree unit in the edge-crossing state into sub-coding units further includes: monitoring the temporal stability data of the corresponding maximum coding tree unit during the coding prediction process; when the temporal stability data is higher than the preset stability threshold and the transform residual data is lower than the preset noise threshold, blocking the transform path in the sub-coding unit set through controller instructions, and guiding the video encoder to start the skip mode to reduce the invalid computing load of the edge hardware.

3. The method of claim 1, wherein, The step of calculating the spatial variance metric of the corresponding mask pixels within the grid region of each maximum coding tree unit further includes: if the spatial variance metric of a certain maximum coding tree unit is identified as 0 and the mean value of the corresponding mask pixels is a background attribute, then the maximum coding tree unit is marked as a static background block; grid dormancy processing is implemented on the static background block, the motion vector search path of the static background block is forcibly blocked and only the DC component is retained.

4. The edge-end semantic-aware transmission method of claim 1, wherein, The video frames to be transmitted are extracted by a lightweight transformer model deployed at the edge; the binary semantic mask matrix is ​​generated by the attention map transformation of the lightweight transformer model; the semantic confidence data is determined by the maximum probability distribution output of the last layer classifier of the lightweight transformer model; the lightweight transformer model is obtained by parameter pruning of a preset deep learning network.

5. The edge-end semantic-aware transmission method according to claim 1, characterized in that, Coupling priority score S satisfies: Where S is the coupling priority score, and α is the preset spatial domain weight coefficient. The data represents semantic confidence scores, and β is a preset time-domain weighting coefficient. The steps for adaptively mapping the quantization parameters of the sub-coding unit set, which are derived from the transform residual data extracted from the hardware registers inside the video encoder and where the sum of α and β is 1, specifically include: establishing a nonlinear mapping function that couples the priority score with the quantization step size of the video encoder; when an object contour is identified to contain a high-confidence foreground and unpredictable deformation, the quantization parameters of the sub-coding unit set are reduced according to the nonlinear mapping function, and an enhanced coding action for high-frequency textures is triggered, so that the allocation of coding resources is tilted towards the object contour.

6. The edge-end semantic-aware transmission method according to claim 1, characterized in that, The steps for implementing adaptive mapping of quantization parameters for the sub-coding unit set also include: real-time monitoring of available bandwidth data of the physical channel at the edge; using the available bandwidth data as the global constraint boundary of the bitrate control module; within the global constraint boundary, adjusting the quantization parameters of each sub-coding unit set to ensure that the bit allocation ratio of the corresponding foreground region in the sub-coding unit set is always higher than the bit allocation ratio of the corresponding background region, and the transmission of the video stream is achieved through a stream priority control protocol; marking data packets containing object outlines in the video stream as high-priority streams and marking data packets containing background regions as low-priority streams; and, in the case of network congestion, actively discarding data packets of low-priority streams through the gateway device to ensure that the end-to-end transmission latency of high-priority streams is less than 50ms.

7. The edge-end semantic-aware transmission method according to claim 1, characterized in that, It also includes the cloud feedback step: the receiving end receives the video stream and reconstructs the video frame; The edge detection operator response value of the reconstructed video frame is calculated and sent back to the edge side. The edge side dynamically adjusts the recursion depth of the sub-coding unit division of the largest coding tree unit in the edge crossing state according to the received edge detection operator response value. The step of adaptive mapping of quantization parameters on the sub-coding unit set is implemented by the lookup table method: multiple sets of quantization parameter offset tables are stored in advance. Based on the numerical range to which the coupling priority score belongs, select the matching quantization parameter offset from multiple quantization parameter offset tables; The original quantization parameters of the sub-coding unit set are corrected using the matched quantization parameter offset, and the corrected quantization parameters are finally generated for entropy coding.

8. A semantically aware video compression method at the edge, characterized in that, This includes the steps of the edge-end semantic-aware transmission method as described in claim 1.

9. A device for edge-end semantic awareness, characterized in that: It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the edge semantic awareness transmission method as described in claim 1.

10. An edge-end semantic-aware electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the edge semantic awareness transmission method as described in claim 1; the electronic device is used for video bitstream transmission in security monitoring systems, video conferencing terminals, or intelligent traffic monitoring systems.