A joint encoding compression method and system for adjacent multiple camera videos

CN122802696APending Publication Date: 2026-09-22HANGZHOU DIANZI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610833413.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]为解决上述低码率下视频失真的技术难题,本发明提出了一种面向邻近多个摄像头视频的联合编码压缩方法和系统,通过对多摄像头全局语义去重,结合跨模态特征编码和前-背景融合视频重建方案,有效消除多摄像头视频间的语义冗余,从而保障低码率下的高质量监控视频重建

Benefits of technology

1.区别于传统视频编码方法在单摄像头内进行冗余消除的技术方案,本发明采用脉冲驱动的语义去重视频场景分析方法的技术方案,通过构建主语义库对多摄像头间的重复前景进行全局重识别与去重以及背景关键帧提取,使得语义冗余数据大幅减少,从而有效降低多摄像头视频编码和重建所需的数据量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802696A_ABST
    Figure CN122802696A_ABST
Patent Text Reader

Abstract

A joint encoding compression method and system for adjacent multiple camera videos, the method comprising: collecting a sequence of original video frames of multiple cameras; using a pulse neural network to perform semantic de-duplication on the sequence of original video frames to generate a foreground main semantic library; extracting and encoding foreground motion features and spatial features; extracting foreground double-antagonistic semantic feature maps and background double-antagonistic semantic feature maps from foreground key frames and background key frames in the foreground main semantic library; performing frequency domain encoding and decoding on the foreground double-antagonistic semantic feature maps and the background double-antagonistic semantic feature maps; integrating the encoded foreground motion features, the encoded spatial features, the foreground double-antagonistic semantic feature maps and the background double-antagonistic semantic feature maps into a cross-modal feature set; reconstructing foreground and background images based on the pulse neural network; generating a complete foreground frame sequence based on a pose transfer network; performing cross-modal feature fusion to reconstruct a video, and outputting a reconstructed complete video frame sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surveillance video compression, and in particular to a joint encoding compression method and system for videos from multiple adjacent cameras. Background Technology

[0002] Video compression technology for surveillance scenarios is receiving increasing attention. To achieve effective video data compression, the industry has recently needed a method to remove redundancy from surveillance videos, thereby saving video transmission bitrate and reducing the amount of video storage data.

[0003] The essence of image and video compression is to use algorithms to eliminate various redundant information in image and video signals, such as spatial repetition, temporal redundancy, visually unnecessary elements, and additional information at the coding level. In the field of surveillance video compression, some preliminary research has emerged in recent years. This research can be broadly divided into two categories: traditional video coding and deep learning-based video coding. Traditional video coding typically refers to a series of hybrid coding frameworks based on a hybrid block structure. For example, the widely used H.264 / AVC video coding standard includes modules such as intra-frame prediction, inter-frame prediction, transform and quantization, and entropy coding, achieving high compression ratios, high image quality, and strong network adaptability. When applying deep learning technology to video compression frameworks, deep learning-based video compression can be divided into two categories: deep learning-based hybrid video compression and deep learning-based end-to-end video compression.

[0004] Most existing video coding methods are traditional video coding based on hybrid block structures and deep learning-based video coding. While these methods have achieved good video compression performance, they still have the following problems: 1) Existing video coding schemes are all pixel-level coding, and do not fully utilize the higher-level feature of semantic information within the image to achieve more efficient coding compression; 2) Existing deep learning-based video coding methods are all based on artificial neural networks, which have the problem of high energy consumption; 3) Traditional video coding processes based on QP operation quantization are prone to block artifacts in low bitrate transmission scenarios, resulting in a decrease in video quality. Summary of the Invention

[0005] To address the technical challenge of video distortion at low bitrates, this invention proposes a joint coding compression method and system for videos from multiple neighboring cameras. By performing global semantic deduplication on multiple cameras and combining cross-modal feature coding and foreground-background fusion video reconstruction schemes, semantic redundancy between multiple camera videos is effectively eliminated, thereby ensuring high-quality surveillance video reconstruction at low bitrates.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, a joint coding compression method for videos from multiple adjacent cameras includes the following steps: S1. Acquire raw video frame sequences from multiple cameras; use a spiking neural network to perform semantic deduplication on the raw video frame sequences and generate a foreground main semantic library; S2. Extract and encode foreground motion features and spatial features based on foreground frame sequences; extract foreground dual-antagonistic semantic feature maps and background dual-antagonistic semantic feature maps from foreground keyframes and background keyframes in the foreground main semantic library based on a dual-antagonistic mechanism; perform frequency domain encoding and decoding on the foreground dual-antagonistic semantic feature maps and background dual-antagonistic semantic feature maps; integrate the encoded foreground motion features, encoded spatial features, foreground dual-antagonistic semantic feature maps, and background dual-antagonistic semantic feature maps and output them as a cross-modal feature set; S3. Based on the cross-modal feature set, reconstruct the foreground and background images using a spiking neural network; generate a complete foreground frame sequence using a pose transfer network; perform cross-modal feature fusion to achieve video reconstruction, and output the reconstructed complete video frame sequence.

[0007] Preferably, S1 includes: S11. Segment the original video frame sequence based on a spiking neural network to generate a video scene segment set; S12. Slice the video frames and filter the foreground. Based on the filtering results, separate the foreground and extract the background to generate a complete sequence of background keyframes and foreground frames. S13. Extract multimodal features from the foreground frame sequence and perform feature fusion. Construct a pulse matrix sequence based on the generated fused features. Input the pulse matrix sequence into the constructed 4-layer pulse neural network to generate saliency scores for video frames. Extract foreground keyframes based on the saliency scores. S14. Obtain the set of foreground keyframes for each camera in the current time period; aggregate the foreground objects of all cameras; use the metric learning algorithm to calculate the similarity distance between any two foreground feature vectors, cluster the foreground objects and remove duplicate objects; based on the deduplication results, construct a non-duplicate foreground main semantic library.

[0008] Preferably, S11 includes: S111. Construct a 6-dimensional illumination feature vector for each video frame based on the original video frame sequence, including statistical features, distribution features, and structural features; calculate the firing probability based on the 6-dimensional illumination feature vector; copy the firing probability along the time axis as the benchmark firing probability for each micro-time step within the time window; generate a random variable that follows a uniform distribution from 0 to 1, compare the random variable with the firing probability, and generate an encoded pulse sequence. S112. Input the generated coded pulse sequence into a three-layer spiking neural network to integrate spatiotemporal features and construct the feature state vector of each frame. S113. Calculate the semantic similarity between adjacent video frames using feature state vectors; Based on semantic similarity, a set of scene change points is defined; the original video frame sequence is segmented using scene change points as boundaries to obtain preliminary video sequence segments; a merging mechanism based on segment statistical characteristics is used to merge the segments; after iterative merging, a set of video scene segments is obtained.

[0009] Preferably, S12 includes: S121. Input each video frame into the SparseFormer architecture to generate a set of moving foreground marker RoIs; S122. Divide the video frame image into slices to generate a set of mesh slices at different scales; use the set of moving foreground marked RoIs as a priori guide; for each mesh slice, traverse the set of moving foreground marked RoIs and calculate the spatial intersection area between the slice region and each marked RoI; determine whether the mesh slice is a foreground mesh slice containing the target based on the spatial intersection area and retain it. S123. After preprocessing the foreground mesh slices, input them into the pulse target detection backbone network to extract the deep spatial and texture features of the slices; the pulse detection head receives the membrane potential state of the neurons in the output layer of the backbone network at the last time step as a continuous value feature and inputs it into the detector to generate the local coordinate system detection box and the corresponding foreground category confidence score inside each slice, thereby generating the foreground frame sequence. S124. Construct an initial background image using statistical filtering methods by utilizing temporal redundancy information within video segments; generate a candidate background pixel set based on the foreground slice selection results. S125. Based on the candidate background pixel set, a deep learning model is used to generate and complete semantics for regions where effective background information cannot be obtained, and a missing region mask matrix is ​​defined. The initial background image and the missing region mask matrix are input into a pre-trained deep generative inpainting network to output complete background keyframes.

[0010] Preferably, S13 includes: S131. Multimodal features include temporal difference features and spatial structure features; each pixel intensity value in the fused features is mapped to a pulse firing probability to generate a pulse matrix sequence; the pulse matrix sequence is input into a spiking neural network to generate a saliency score for each video frame; S132. Based on the adaptive threshold strategy and saliency score, calculate the adaptive threshold; combine the adaptive threshold with the minimum time interval constraint, define the keyframe decision function, and mark the foreground frame sequence based on the keyframe decision function to generate foreground keyframes.

[0011] Preferably, S2 includes: S21. Input the foreground keyframes from the foreground main semantic library into the OpenPose network to extract the skeleton model of 18 key points as foreground motion features, and simultaneously extract the occupancy set of foreground pixels as spatial features. S22. After mapping the RGB images of the foreground keyframes and background keyframes in the foreground main semantic library to the antagonistic color space, apply the Laplacian operator to each channel to extract the double antagonistic feature map, and generate the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map. S23. Perform discrete cosine transform and quantization, zigzag scanning and entropy coding on the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map; perform inverse entropy coding at the decoding end, and output the recovered frequency domain coefficients and double antagonistic semantic feature map.

[0012] Preferably, S3 includes: S31. The multi-channel edge information of the recovered dual-antagonistic semantic feature map is pulse set encoding, the surface features are filled iteratively through the discrete diffusion dynamics equation, and the three channels are fused by perceptual weights and then mapped back to the RGB space through inverse transformation to generate the reconstructed color foreground keyframe sequence and the reconstructed background keyframe. S32. Perform multi-source feature joint encoding and initialization on the reconstructed color foreground keyframe sequence to generate initial image encoding features; input the initial image encoding features and motion features into the pose attention transfer architecture to generate a complete foreground frame sequence; S33. Fuse the complete foreground frame sequence and the reconstructed key background frames to output the reconstructed complete video frame sequence.

[0013] Preferably, S22 includes: The RGB images of foreground and background keyframes from the foreground main semantic library are mapped to an antagonistic color space using a linear transformation matrix, decomposing them into three independent channels: red-green, blue-yellow, and luminance. A Laplacian operator is then applied to each of the three channels. Perform second-order derivative filtering to obtain the double-antagonistic feature map of each channel.

[0014] Preferably, S31 includes: For each pixel location, three independent sets of spiking neurons are configured for each channel, with each set containing several LIF neurons. A frequency coding strategy is adopted so that the firing rate of neurons in each set is proportional to the intensity of the corresponding input edge. In the spiking neural network architecture, each neuron interacts with its neighboring neurons through horizontal connections. Through iterative updates over multiple time steps, surface feature filling is completed, resulting in reconstructed surface feature maps for each channel. Based on adjustable perceptual weights and surface feature maps, a fused antagonistic channel vector is generated. The fused antagonistic channel vector is then mapped back to the RGB color space.

[0015] Secondly, a joint coding compression system for videos from multiple adjacent cameras includes: The data acquisition and semantic deduplication module is used to acquire raw video frame sequences from multiple cameras, and to perform semantic deduplication on the raw video frame sequences using a spiking neural network to generate a foreground main semantic library. The feature extraction and cross-modal coding module is used for: extracting and encoding foreground motion features and spatial features based on foreground frame sequences; extracting foreground and background double-antagonistic semantic feature maps from foreground keyframes and background keyframes in the foreground main semantic library based on a double-antagonistic mechanism; performing frequency domain encoding and decoding on the foreground and background double-antagonistic semantic feature maps; and integrating the encoded foreground motion features, encoded spatial features, foreground double-antagonistic semantic feature maps, and background double-antagonistic semantic feature maps into a cross-modal feature set. The video reconstruction module is used for: reconstructing foreground and background images based on a spiking neural network according to a cross-modal feature set; generating a complete foreground frame sequence based on a pose transfer network; performing cross-modal feature fusion to achieve video reconstruction and outputting the reconstructed complete video frame sequence; The aforementioned joint coding compression system for videos from multiple adjacent cameras is used to implement the joint coding compression method and steps for videos from multiple adjacent cameras as described in the first aspect.

[0016] Compared with the prior art, the beneficial effects of the present invention are reflected in: 1. Unlike traditional video coding methods that eliminate redundancy within a single camera, this invention employs a pulse-driven semantic deduplication video scene analysis method. By constructing a main semantic library, it performs global re-identification and deduplication of repetitive foregrounds across multiple cameras and extracts background keyframes, thereby significantly reducing semantic redundancy data and effectively reducing the amount of data required for multi-camera video coding and reconstruction.

[0017] 2. Unlike traditional video decoding processes that rely on high-compression-rate quantization leading to block artifacts, this invention employs a cross-modal feature encoding based on spiking neural networks and a foreground / background fusion video reconstruction scheme based on pose transfer networks. By utilizing the complete multimodal features after decoding for prediction, generation, and fusion, the system can maintain the integrity of semantic information even under low bitrate transmission conditions, effectively eliminating block artifacts and significantly improving the visual perception quality of the reconstructed video. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the first part of the method framework of Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the second part of the method framework of Embodiment 1 of the present invention; Figure 3This is a schematic diagram of step S1 of Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of step S2 in Embodiment 1 of the present invention; Figure 5 This is a flowchart illustrating step S3 of embodiment 1 of the present invention; Figure 6 This is a performance comparison diagram of Embodiment 1 of the present invention. Detailed Implementation

[0019] To make the technical means, inventive features, objectives, and effects of the invention readily understandable, the invention is further described below with reference to specific illustrations. However, the invention is not limited to the embodiments described below.

[0020] It should be noted that the structures, proportions, sizes, etc., illustrated in the accompanying drawings of this specification are only used to complement the content disclosed in the specification for those skilled in the art to understand and read, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0021] Existing surveillance video coding methods have undergone multiple iterations and upgrades, resulting in significant performance improvements. However, previous methods have often been limited to redundancy elimination within a single camera frame, failing to explore the semantic connections between multiple camera frames in the same scene for redundancy elimination. Therefore, this invention proposes to reduce the transmission bitrate and storage data volume of multiple camera videos by jointly encoding and compressing multiple video data to eliminate semantic redundancy between multiple cameras.

[0022] Therefore, the technical problem this invention aims to solve is the challenge of severe video reconstruction distortion and blockiness in low-bitrate transmission scenarios when processing multi-camera video data in existing monitoring systems. Specifically, within a monitoring scenario, multiple cameras often capture the same foreground moving objects (i.e., semantic redundancy across cameras). Traditional video coding is mostly limited to pixel-level redundancy elimination within a single camera frame, failing to perform global semantic feature separation and redundancy elimination across multiple cameras. This results in a large amount of invalid redundant data consuming limited transmission resources under bandwidth-constrained low-bitrate transmission, forcing the encoder to adopt extremely high compression rates (such as strong quantization processes based on QP operations), ultimately leading to severe blockiness and quality degradation in the decoded monitoring video.

[0023] The specific steps of the method of the present invention include: Pulse-driven semantic deduplication video scene analysis: This invention proposes, as follows Figure 2This paper presents a pulse-driven semantic deduplication video scene analysis method. The method first performs video segmentation based on factors such as illumination, then separates the foreground and background, reconstructs the background frames, and extracts keyframes from the foreground frames. To eliminate semantic redundancy between multi-camera images, the semantic features extracted from multiple cameras need to be deduplicated. Based on foreground and background separation and keyframe extraction, this method first re-identifies the foreground keyframes extracted in scene analysis, retaining only one of the foreground elements that appear repeatedly across multiple cameras. A main semantic library is constructed, and these processed, non-repeating foreground keyframes are integrated into the main semantic library.

[0024] Cross-modal feature encoding based on spiking neural networks: This invention proposes, as follows Figure 3 The method presented is a cross-modal feature encoding approach based on spiking neural networks. The method first extracts video features from multiple modalities: background semantic features, foreground semantic features, foreground motion features, and foreground spatial features obtained after scene analysis. Then, an end-to-end encoding method based on spiking neural network image filling is used to encode the foreground and background semantic features of keyframes, and entropy encoding is performed on the motion and spatial features of the foreground keyframes.

[0025] Foreground / background fusion video reconstruction based on pose transfer networks: This invention proposes, as follows... Figure 4 The method shown is a foreground-background fusion video reconstruction method based on pose transfer network. At the decoding end, the received multimodal features are used to generate a complete foreground sequence through pose transfer network and then fused with the background to reconstruct a complete video frame sequence.

[0026] Example 1: like Figure 1 The method shown is a joint coding compression method for videos from multiple adjacent cameras, comprising the following steps: S1. Acquire raw video frame sequences from multiple cameras; use a spiking neural network to perform semantic deduplication on the raw video frame sequences, generating a foreground main semantic library; such as... Figure 2 As shown, it specifically includes: S11. Segment the original video frame sequence based on a spiking neural network to generate a video scene segment set; S111. Multidimensional feature pulsed representation integrating illumination semantics: To reduce computational redundancy and accelerate subsequent analysis, the original video frame sequence is first downsampled. For each downsampled frame, a 6-dimensional illumination feature vector is constructed. It covers statistical features, distribution features, and structural features, specifically including: average brightness features reflecting the overall brightness of the scene, brightness variance features reflecting the discreteness of the scene's illumination distribution, two-dimensional histogram features reflecting the statistical distribution law of brightness, contrast features reflecting the clarity of the scene's texture, and gradient features reflecting the edge intensity of illumination changes.

[0027] To adapt to the input requirements of spiking neural networks, the continuous-valued feature vectors need to be converted into discrete pulse sequences. For the ... Feature vector of frame image Each of its dimensions Each feature needs to be encoded independently. First, the system will encode the original feature values. Mapped to Within the probability interval. Define the features. Preset maximum value and minimum value Calculate the normalized distribution probability To prevent numerical overflow caused by noise, a truncation function is introduced, calculated as follows: in, Indicates the first The distribution probability after normalization of the dimensional features; Indicates the first The first frame of the image 3D input feature values; and They represent the first The maximum and minimum values ​​preset for the dimensional feature; This is a truncation function used to restrict a value to a certain range. Within the range, prevent overflow.

[0028] Unlike artificial neural networks that perform a single forward propagation, SNNs require processing information over a time period. Therefore, this invention introduces a time window. The distribution probability calculated above Copy along the timeline as each micro-time step within that time window. The baseline firing probability. That is, for any time within the processing period of the current frame. Distribution probability Equal to .

[0029] After obtaining the firing probability, a binary pulse is generated using a random process. For each time step... and each feature dimension Generate a random variable that follows a uniform distribution from 0 to 1. Follows distribution Subsequently, this random variable was compared with the distribution probability. A comparison is made to determine whether a pulse should be issued at that moment. Pulse signal. The generation follows Bernoulli's trial rule: in, Indicates at time step Time Pulse signals generated by dimensional features; This represents a random variable that follows a uniform distribution from 0 to 1. This represents the probability of this feature being emitted.

[0030] S112. Semantic State Quantization Based on Spatiotemporal Feature Integration of SNN: The generated coded pulse sequence is input into a three-layer spiking neural network (SNN) for spatiotemporal feature integration. The network output layer generates a pulse sequence. To quantify the feature states between frames, the output neuron is analyzed within the time window. The pulse firing situation within the range is statistically analyzed to construct the first... Frame feature state vector in, For the first The spatiotemporal feature state vector of a frame; The length of the time window; This represents the number of neurons in the output layer of the SNN. Indicates the first A neuron at time 1 The pulse output state (0 or 1).

[0031] S113. Adaptive segmentation combining cosine similarity and spatiotemporal consistency: Semantic similarity between adjacent video frames is calculated using feature state vectors. The cosine similarity formula is used to calculate the... Frame and the Frame similarity score .

[0032] Define a set of scene change points based on similarity scores. : in Represents the set of points where the scene changes. The total duration of each segment. A similarity score is given to adjacent frames. This is the similarity threshold used to determine whether a sudden change has occurred in the scene. The resulting set of change points is... The video sequence is segmented based on the points of change, resulting in preliminary video sequence segmentation. ,in .

[0033] To address the potential oversegmentation in the initial segmentation, a merging mechanism based on segmentation statistical characteristics is introduced. The merging operation is performed on the two segments when the following joint criterion is met: in, and This indicates two adjacent preliminary segments. and These represent the average brightness within the corresponding segments. and These represent the luminance variance within the corresponding segments. The average brightness difference threshold. The threshold for brightness variance stability is defined as follows: When the mean brightness of two segments is close and their internal illumination fluctuations are small, they are considered a continuation of the same scene. After iterative merging, the final set of video scene segments is obtained. .

[0034] S12. Slice the video frames and filter the foreground. Based on the filtering results, separate the foreground and extract the background to generate a complete sequence of background keyframes and foreground frames. S121. Global Motion Awareness and Coarse Foreground RoI Localization Based on SparseFormer: For the current video frame, the SparseFormer architecture is introduced to extract temporal motion priors. SparseFormer abandons dense pixel-level traversal and has extremely limited initialization. There are 3 potential labels, each containing an embedding vector. and an explicit region of interest descriptor When processing frame difference features, the network generates adaptive sampling points through a focus transformer and iteratively adjusts the labeled RoIs based on the sampled features. After multiple rounds of adjustment, this... The labeled RoIs automatically converge and focus on significant foreground regions that are changing within the image. The network ultimately outputs a set of labeled RoIs that roughly locate the moving foreground. .

[0035] S122. Multi-scale slice selection based on foreground RoI and spatial intersection: For high-resolution input images, multi-level regular grids are used to divide the images into slices, generating grid slices of different scales. The previously generated set of motion foreground tags RoIs will be used. As a priori guide. For each mesh slice traverse the set Calculate the spatial intersection area between the sliced ​​region and each marked RoI. The slice activation function is defined as follows: in, Indicates the first A grid slice of image; Indicates the first Whether each grid slice is activated and retained; Represents the foreground label set The first in One marked area; This represents the spatial intersection area between the grid slice and the marked region. If a slice satisfies... If a slice has at least one spatial overlap with a foreground focus area marked by a SparseFormer, it is determined to be a foreground mesh slice and retained; otherwise, if there is no overlap, it is determined to be a pure static background and discarded directly.

[0036] S123. Pulse-Driven Foreground Extraction and Classification: After preprocessing the foreground mesh slices, they are input into the pulse target detection backbone network to extract the deep spatial and texture features of the slices. Since a large amount of static background has been filtered out from the slices, the backbone network can highly focus on the feature extraction of local foregrounds. The pulse detection head receives the membrane potential state of the neurons in the output layer of the backbone network at the last time step as a continuous value feature and inputs it into the detector to generate anchor boxes of different sizes, thereby outputting the local coordinate system detection box inside each slice. The corresponding category confidence scores are used for foreground extraction and foreground discrimination.

[0037] S124. Initial Background Reconstruction Based on Temporal Masking and Median Filtering: For each video segment frame sequence obtained through scene analysis, foreground-background separation processing is first performed to generate the corresponding binary foreground mask sequence. .

[0038] The background semantics of surveillance videos gradually change over long periods (especially lighting semantics), but these changes are usually small within short periods. This creates temporal redundancy. Therefore, background keyframe extraction (each background keyframe corresponds to a video segment) can be performed on the segmented video after scene analysis to eliminate temporal redundancy between segments. Within a video segment, background parts occluded by the foreground will reappear at other times, creating intra-segment temporal redundancy. This stage utilizes the temporal redundancy information within the video segments to construct an initial background image using statistical filtering methods. .

[0039] For each spatial location of the image plane Traverse the time dimension Collect all pixel values ​​that are not occluded by the foreground to construct a candidate background pixel set. : in, Spatial coordinates The set of valid candidate background pixels at the location; For the first Frame in coordinates Pixel value at; For the first The binary foreground mask value corresponding to the frame (0 indicates no occlusion). A subset is extracted when there are too many samples.

[0040] S125. Background semantic repair based on deep generative networks: For regions where effective background information cannot be obtained due to long-term occlusion in the first stage (i.e., NaN regions), a deep learning model is used for semantic generation and completion.

[0041] Define the missing region mask matrix based on the validity of the initial background. ,in For indicator variables: in, This is an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise; This represents the spatial coordinates collected in the preceding steps. The set of valid candidate background pixels at the location The number of elements (i.e., the number of unoccluded pixel samples); This is a preset threshold for the number of valid background pixels. When the number of valid pixel samples at a certain location... Less than the threshold When the location is determined to lack sufficient background information for statistical filtering, it is considered a region with missing background.

[0042] Construct a pre-trained DeepFill V2 deep generative repair network The network parameters are It employs an encoder-decoder architecture incorporating an attention mechanism to maintain the semantic consistency of the generated content. The initial background... With missing mask The network takes the inputs together and outputs the final complete background keyframes. : Through this step, the model uses known background texture features to infer and fill in missing areas, thereby obtaining a complete background image without holes.

[0043] S13. Extract multimodal features from the foreground frame sequence and perform feature fusion. Construct a pulse matrix sequence based on the generated fused features. Input the pulse matrix sequence into the constructed 4-layer pulse neural network to generate saliency scores for video frames. Extract foreground keyframes based on the saliency scores. S131. Saliency Evaluation Combining Multimodal Features and Spiking Neural Networks: To capture the dynamic changes and morphological details of foreground objects, the foreground frame sequences extracted in S12 are first preprocessed by size normalization and grayscale conversion. Multimodal features of the preprocessed frame sequences are then extracted, including temporal difference features obtained by calculating the absolute differences between adjacent frames. Spatial structural features extracted using edge detection operators The two features are fused to obtain the fused feature: in, This represents the fusion features of the current frame; This represents the temporal difference characteristics calculated from the absolute differences between adjacent frames; This represents the spatial structural features extracted by the edge detection operator; and To integrate the weighting coefficients, and satisfy the following conditions: .

[0044] Spiking neural networks (SNNs) are used to integrate spatiotemporal information of fused features to quantify the salience of the current frame as a key frame.

[0045] Fusion features Each pixel intensity value is mapped to a pulse firing probability within a set time window. Within a time step, if a pixel has a high intensity, then in The higher the probability of generating pulses within a given time step, the better. After normalization, a pulse matrix sequence is generated for network input.

[0046] Construct a 4-layer spiking neural network to receive the above spiking matrix. The network then... The integration-firing process at each time step outputs the pulse sequence of the last layer. Extract the last time step. membrane potential state The first generation is generated by mapping through a fully connected layer and applying the Sigmoid activation function. Frame saliency score : in, Indicates the first The saliency score of the frame as a keyframe; This indicates that the SNN network is in the last layer and at the last time step. The membrane potential state vector; and These are the weight vector and bias term of the fully connected layer, respectively.

[0047] S132. Adaptive Decision Based on Dynamic Statistical Thresholds and Temporal Constraints: To adapt to the differences in different video types (such as action-intensive and smooth video types) and avoid missed or false detections caused by fixed thresholds, this invention adopts an adaptive threshold strategy based on sliding window statistics. According to recent... Calculate an adaptive threshold using historical saliency scores of frames. : in, This represents the keyframe decision threshold for the current frame; and These are the most recent Mean and standard deviation of frame significance scores; It is the adjustment coefficient.

[0048] Combining adaptive thresholds and minimum time interval constraints, the final keyframe selection logic is formulated to ensure a reasonable distribution of keyframes on the timeline (avoiding excessive density). A keyframe decision function is defined. : in, This is the keyframe flag. The timestamp index of the last keyframe; The minimum time interval limit is set. A frame is marked as a foreground keyframe only if the salience score of the current frame exceeds the adaptive threshold and the distance to the previous keyframe meets the minimum interval requirement.

[0049] S14. Obtain the set of foreground keyframes for each camera in the current time period; aggregate the foreground objects of all cameras; use the metric learning algorithm to calculate the similarity distance between any two foreground feature vectors, cluster the foreground objects and remove duplicate objects; based on the deduplication results, construct a non-duplicate foreground main semantic library.

[0050] S141. Cross-view object feature aggregation and global re-identification: Suppose the monitoring system includes... Several cameras with overlapping fields of view. First, acquire data from each camera. The set of foreground keyframes within the current time period. Define the camera. The set of foreground objects is : in, Indicates the first The number of cameras captured the first The semantic feature vector of a foreground object This represents the total number of objects captured by the camera.

[0051] To eliminate duplicate recordings of the same target from different cameras, the system performs global re-identification and discrimination of the foreground object sets from all cameras. This aggregates the foreground objects from all viewpoints into a complete set. The similarity distance between any two foreground feature vectors is calculated using a metric learning algorithm, and objects belonging to the same identity are clustered together. For each identity cluster, only the highest-quality feature vector is retained as the unique representative of that identity, and duplicate objects from other perspectives are marked as redundant and discarded.

[0052] S142. Construction of the foreground main semantic library: Based on the above deduplication results, construct a non-repeating foreground main semantic library. This process can be formally described as: subtracting the set of redundant objects resulting from the repetition of foreground objects from multiple viewpoints from the complete set of foreground objects captured by all cameras. in, This is the final deduplicated foreground main semantic library. Indicates all The complete set of original foreground keyframe objects for each camera; This represents a set of redundant objects that have been identified by the re-identification algorithm as belonging to multi-view duplicate records; This represents the difference operation between sets.

[0053] S2. Extract and encode foreground motion and spatial features based on the foreground main semantic library; extract foreground and background antagonistic semantic feature maps from foreground and background keyframes in the foreground main semantic library based on a antagonistic mechanism; perform frequency domain encoding and decoding on the foreground and background antagonistic semantic feature maps; integrate the encoded foreground motion and spatial features, as well as the foreground and background antagonistic semantic feature maps, and output them as a cross-modal feature set; such as... Figure 3 As shown, it specifically includes: S21. Input the foreground keyframes from the foreground main semantic library into the OpenPose network to extract the skeleton model of 18 key points as foreground motion features, and simultaneously extract the occupancy set of foreground pixels as spatial features. Motion feature extraction: Utilizing the OpenPose pose estimation network for foreground frame sequences The analysis was conducted. The network locates specific joints in the human body using cascaded prediction maps and transforms the probability distribution into deterministic coordinates through non-maximum suppression. Finally, a skeleton model consisting of 18 key points was constructed. As a characteristic of motion: in, Indicates the first The skeleton key point coordinate vector of the frame, this process realizes dimensionality reduction from pixel-level data to topological structure data.

[0054] Spatial feature extraction and encoding: Simultaneously extract the geometric distribution information of foreground objects in the image and obtain the occupancy set of foreground pixels. As spatial features. For the generated motion features. and spatial features Arithmetic coding is used for compression. This coding process maps the feature sequence to a cumulative probability distribution. Lossless entropy encoding is achieved for single floating-point numbers within a given interval.

[0055] S22. After mapping the RGB images of the foreground keyframes and background keyframes in the foreground main semantic library to the antagonistic color space, apply the Laplacian operator to each channel to extract the double antagonistic feature map, and generate the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map. Color space antagonistic transformation: First, the RGB image is transformed using a linear transformation matrix. Mapped to an antagonistic color space, it is decomposed into three independent channels: red-green (RG), blue-yellow (BY), and lightness (I). The transformation formula is as follows: This transformation simulates the antagonistic processing mechanism of retinal ganglion cells in response to wavelength differences.

[0056] Dual-antagonistic edge spatial filtering: To extract salient edges in each channel, a Laplacian operator is applied to each of the three channels. Perform second-order derivative filtering. Define... for: Obtain the dual antagonistic feature map of each channel. : in For pixel coordinates, Indicates the output of the first Each channel (e.g., red-green, blue-yellow, or brightness) is located in the coordinate system. eigenvalues ​​at that location This represents the original value of the input image within the neighborhood of the current pixel. This indicates that the Laplacian operator (convolution kernel) is at the offset. The weighting coefficient at the location, This indicates that a weighted summation is performed within a 3x3 window. This step effectively extracts spatially abrupt changes in color and brightness in the image.

[0057] S23. Perform discrete cosine transform and quantization, zigzag scanning and entropy coding on the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map; perform inverse entropy coding at the decoding end, and output the recovered frequency domain coefficients and double antagonistic semantic feature map.

[0058] Discrete Cosine Transform (DCT) and Quantization: On Dual-Antagonistic Feature Maps The process involves dividing the data into 8×8 blocks and performing a discrete cosine transform on each block to map the spatial domain information to the frequency domain. Then, a quantization matrix is ​​used... The transform coefficients are quantized to remove high-frequency redundancy that is not sensitive to the human eye, generating quantized coefficients. The quantized sparse matrix is ​​then encapsulated into the bitstream after zigzag scanning and entropy encoding.

[0059] Decoding and Feature Recovery: At the decoding end, inverse entropy encoding and dequantization operations are performed on the received bitstream to obtain the recovered frequency domain coefficients. The inverse discrete cosine transform (IDCT) is then performed to recover the bi-antagonistic semantic feature maps for each channel. The recovered bi-antagonistic feature maps DOk' provide crucial color boundary and structural guidance information for subsequent image reconstruction.

[0060] S3. Based on the cross-modal feature set, reconstruct the foreground and background images using a spiking neural network; generate a complete foreground frame sequence using a pose transfer network; perform cross-modal feature fusion to achieve video reconstruction, and output the reconstructed complete video frame sequence. For example... Figure 4 As shown, it specifically includes: S31. The multi-channel edge information of the recovered dual-antagonistic semantic feature map is pulse set encoding, the surface features are filled iteratively through the discrete diffusion dynamics equation, and the three channels are fused by perceptual weights and then mapped back to the RGB space through inverse transformation to generate the reconstructed color foreground keyframe sequence and the reconstructed background keyframe. Pulse set encoding of multi-channel edge information: three dual-opposite channels for each pixel location. Independent spiking neuron sets are configured, each set containing several LIF neurons. A frequency coding strategy is employed to make the firing rate of neurons within each set proportional to the corresponding input edge intensity. Through this process, static edge amplitudes are transformed into a pulse stream with time-domain characteristics, serving as the input excitation for subsequent diffusion dynamics.

[0061] Surface feature filling based on discrete diffusion dynamics: In an SNN architecture, each neuron interacts with its neighboring neurons through horizontal connections. The physical dynamics of this process follow a discretized diffusion equation to describe the spatial diffusion effect of neuronal activity. At time step... ,Location Reconstructed surface strength The update formula is: in, Indicates at time step Time position The intensity of the reconstructed surface features at the location; The intensity of the previous time step; The time constant for controlling the diffusion rate; For discrete Laplace operators; The input is the double antagonistic edge constraint excitation signal.

[0062] Through multiple iterations over time steps, the network drives the initial edge energy to gradually "difflate" into the internal blank areas. As the number of iterations increases, the originally sparse contour lines gradually connect and fill into smooth and continuous grayscale or color patches until a steady state is reached, completing surface reconstruction.

[0063] Perceptual enhancement fusion and color image spatial mapping: completing three independent channels After surface filling, the reconstructed surface feature map is obtained. To simulate individual differences in biological vision and enhance perceptual effects, adjustable perceptual weights are introduced: in , , The weight parameters satisfy the normalization condition.

[0064] The fused antagonistic channel vectors Map back to the RGB color space. Utilize a preset inverse color transformation matrix. Perform the following linear transformation: Through the above process, the system ultimately generates a sequence of keyframes for reconstructing the color foreground. The background keyframes are divided into blocks, and the blocks are reassembled to obtain the reconstructed background keyframes. .

[0065] S32. Perform multi-source feature joint encoding and initialization on the reconstructed color foreground keyframe sequence to generate initial image encoding features; input the initial image encoding features and motion features into the pose attention transfer architecture to generate a complete foreground frame sequence; Human pose transfer networks are built based on the GAN (Generative Adversarial Network) concept. Specifically, given an original image and a new pose (represented by keypoints), the network's goal is to transform the pose of the original image into the target pose. The pose transfer network utilized in this invention employs an encoder-transformer-decoder architecture.

[0066] Joint encoding and initialization of multi-source features: constructing a system based on... An image encoder consisting of convolutional layers with a stride of 2, each followed by instance normalization and ReLU activation, extracts initial image coding features from the reconstructed color foreground keyframe image. .

[0067] To enable the network to perceive the dynamic transformation from the source pose to the target pose, the pose heatmap of the reconstructed color foreground keyframes is first constructed. Compared with the non-keyframe pose heatmap to be generated The input channels are stitched together to generate a 36-channel joint pose input. Subsequently, an attitude encoder with the same structure as the image encoder was used to... Processing is performed to generate an initial joint pose feature code. This joint encoding mechanism effectively captures vector information about changes in a person's posture.

[0068] Feature transformation based on Pose Attention Transfer Architecture (PATN): The converter module consists of This module consists of a cascaded Attention Transfer Block (PATB). Level output image features and posture features As input, the updated features are output after internal processing. and Within each PATB unit, the spatial attention mask is first calculated using the current pose features. The mask indicates which regions in the current feature space need texture reconstruction based on changes in the target pose. The specific generation logic is as follows: in, For the generated spatial attention mask feature map; The pose features output by the previous module; and These represent the first and second layer convolution operations, respectively. This indicates an instance normalization operation; It is a non-linear activation function; This is the Sigmoid activation function.

[0069] Based on the generated mask The module performs residual updates on image features. This process leverages the residual learning concept from known architectures to ensure that the original texture information of keyframes is preserved, while only overlaying the transformed features in masked active regions. in, The updated image features at the current level; Features of the previous level image; It is a 3×3 convolutional layer used to transform image features; This represents element-wise matrix multiplication. Attention mask. This acts as a gating mechanism, determining... Which information should be retained and overlaid into the original features? This allows the network to gradually and locally adjust the transferred image features. The updated image features are then... The transformed pose features are concatenated along the channel dimension to form new pose features. : in and It is another set of two 3×3 convolutional layers. This represents a feature concatenation operation along the channel dimension. It involves encoding the image. By combining the pose features of the (t-1)th PATB, image features are fed back based on pose features, enabling subsequent steps to more accurately locate the area to be adjusted.

[0070] Foreground sequence generation: Through T cascaded PATB operations described above, the initial image and pose features are generated. Transformed gradually and smoothly into the final features The decoder module consists of N transposed convolutional layers and is responsible for encoding the final image features. Upsampled to the target image size, the image is then passed through the output layer to generate the final image. Reconstruct the foreground keyframe image. and decoded motion features (i.e., the aforementioned pose features) are input into the aforementioned encoder-converter-decoder architecture to obtain a series of final images. A complete foreground frame sequence .

[0071] S33. Fuse the complete foreground frame sequence, the static foreground local video slices, and the reconstructed key background frames to output the reconstructed complete video frame sequence.

[0072] Cross-modal feature fusion: combining complete foreground frame sequences Reconstructing key background frames Fusion: in This represents the final reconstructed complete video frame sequence; Indicates the foreground frame sequence Based on the spatial characteristics after decoding Perform coordinate and size adjustments; This is the generated sequence of moving foreground frames; These are the corresponding decoded spatial and scale features; These are the corresponding keyframes for reconstructing the background; This represents the feature overlay and pixel fusion operation of multi-layered images. The above formula represents the reconstructed video frame sequence obtained by fusing the reconstructed foreground frame sequence and the grouped reconstructed key background frames according to spatial features. .

[0073] Analysis of experimental results: This embodiment conducts performance tests on videos captured by five cameras. The video resolution is 240p, the frame rate is 30fps, and each video contains 3000 frames. Four comparative experiments were performed, including traditional benchmark encoders JM, HM, and VTM, and a representative deep learning end-to-end encoder, DVC. Since this embodiment involves joint video coding, to ensure a fair comparison of the performance of each scheme, the traditional encoder and DVC used a weighted average method based on the number of video frames from multiple cameras when calculating the bitrate. Video reconstruction quality was measured using two metrics that effectively measure human visual perception quality: Learned Perceptual Image Patch Similarity (LPIPS) and Deep Image Structure and Texture Similarity. Lower metric values ​​indicate less distortion and higher perceptual quality. The experimental comparison results are as follows: Figure 6 As shown in the figure. Implementation results demonstrate that the pulse-driven multi-camera joint coding framework proposed in this invention exhibits superior performance in multi-camera video systems. The degree of average distortion optimization of the reconstructed video by the proposed scheme compared to VTM and DVC is specifically shown in Table 1.

[0074] Table 1: The degree of optimization of average distortion of the present invention compared to VTM and DVC

[0075] Example 2: A joint encoding and compression system for videos from multiple adjacent cameras, comprising: The data acquisition and semantic deduplication module is used to acquire raw video frame sequences from multiple cameras, and to perform semantic deduplication on the raw video frame sequences using a spiking neural network to generate a foreground main semantic library. The feature extraction and cross-modal coding module is used for: extracting and encoding foreground motion features and spatial features based on foreground frame sequences; extracting foreground and background double-antagonistic semantic feature maps from foreground keyframes and background keyframes in the foreground main semantic library based on a double-antagonistic mechanism; performing frequency domain encoding and decoding on the foreground and background double-antagonistic semantic feature maps; and integrating the encoded foreground motion features, encoded spatial features, foreground double-antagonistic semantic feature maps, and background double-antagonistic semantic feature maps into a cross-modal feature set. The video reconstruction module is used to: reconstruct foreground and background images based on a spiking neural network according to a cross-modal feature set; generate a complete foreground frame sequence based on a pose transfer network; perform cross-modal feature fusion to realize video reconstruction, and output the reconstructed complete video frame sequence.

Claims

1. A joint coding compression method for videos from multiple adjacent cameras, characterized in that, Includes the following steps: S1. Acquire raw video frame sequences from multiple cameras; use a spiking neural network to perform semantic deduplication on the raw video frame sequences and generate a foreground main semantic library; S2. Extract and encode foreground motion features and spatial features based on foreground frame sequences; Based on the dual antagonism mechanism, foreground dual antagonism semantic feature maps and background dual antagonism semantic feature maps are extracted from foreground keyframes and background keyframes in the foreground main semantic library; Frequency domain encoding and decoding are performed on the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map; the encoded foreground motion features, encoded spatial features, foreground double antagonistic semantic feature map and background double antagonistic semantic feature map are integrated and output as a cross-modal feature set; S3. Based on the cross-modal feature set, reconstruct the foreground and background images using a spiking neural network; A complete foreground frame sequence is generated based on a pose transfer network; cross-modal feature fusion is performed to reconstruct the video, and the reconstructed complete video frame sequence is output.

2. The joint coding compression method for videos from multiple adjacent cameras according to claim 1, characterized in that, S1 includes: S11. Segment the original video frame sequence based on a spiking neural network to generate a video scene segment set; S12. Slice the video frames and filter the foreground. Based on the filtering results, separate the foreground and extract the background to generate a complete sequence of background keyframes and foreground frames. S13. Extract multimodal features from the foreground frame sequence and perform feature fusion. Construct a pulse matrix sequence based on the generated fused features. Input the pulse matrix sequence into the constructed 4-layer pulse neural network to generate saliency scores for the video frames. Extract foreground keyframes based on saliency scores; S14. Obtain the set of foreground keyframes for each camera in the current time period; aggregate the foreground objects of all cameras; use the metric learning algorithm to calculate the similarity distance between any two foreground feature vectors, and cluster the foreground objects and remove duplicate objects; Based on the deduplication results, a non-repeating foreground main semantic library is constructed.

3. The joint coding compression method for videos from multiple adjacent cameras according to claim 2, characterized in that, S11 includes: S111. Construct a 6-dimensional illumination feature vector for each video frame based on the original video frame sequence, including statistical features, distribution features, and structural features; calculate the firing probability based on the 6-dimensional illumination feature vector; copy the firing probability along the time axis as the benchmark firing probability for each micro-time step within the time window; generate a random variable that follows a uniform distribution from 0 to 1, compare the random variable with the firing probability, and generate an encoded pulse sequence. S112. Input the generated coded pulse sequence into a three-layer spiking neural network to integrate spatiotemporal features and construct the feature state vector of each frame. S113. Calculate the semantic similarity between adjacent video frames using feature state vectors; Based on semantic similarity, a set of scene change points is defined; the original video frame sequence is segmented using scene change points as boundaries to obtain preliminary video sequence segments; a merging mechanism based on segment statistical characteristics is used to merge the segments; after iterative merging, a set of video scene segments is obtained.

4. The joint coding compression method for videos from multiple adjacent cameras according to claim 2, characterized in that, S12 includes: S121. Input each video frame into the SparseFormer architecture to generate a set of moving foreground marker RoIs; S122. Divide the video frame image into slices to generate a set of mesh slices at different scales; use the set of moving foreground marked RoIs as a priori guide; for each mesh slice, traverse the set of moving foreground marked RoIs and calculate the spatial intersection area between the slice region and each marked RoI; determine whether the mesh slice is a foreground mesh slice containing the target based on the spatial intersection area and retain it. S123. After preprocessing the foreground mesh slices, input them into the pulse target detection backbone network to extract the deep spatial and texture features of the slices; the pulse detection head receives the membrane potential state of the neurons in the output layer of the backbone network at the last time step as a continuous value feature and inputs it into the detector to generate the local coordinate system detection box and the corresponding foreground category confidence score inside each slice, thereby generating the foreground frame sequence. S124. Construct an initial background image using statistical filtering methods by utilizing temporal redundancy information within video segments; generate a candidate background pixel set based on the foreground slice selection results. S125. Based on the candidate background pixel set, a deep learning model is used to generate and complete semantics for regions where effective background information cannot be obtained, and a missing region mask matrix is ​​defined. The initial background image and the missing region mask matrix are input into a pre-trained deep generative inpainting network to output complete background keyframes.

5. The joint coding compression method for videos from multiple adjacent cameras according to claim 2, characterized in that, S13 includes: S131. Multimodal features include temporal difference features and spatial structure features; each pixel intensity value in the fused features is mapped to a pulse firing probability to generate a pulse matrix sequence; the pulse matrix sequence is input into a spiking neural network to generate a saliency score for each video frame; S132. Based on the adaptive threshold strategy and saliency score, calculate the adaptive threshold; combine the adaptive threshold with the minimum time interval constraint, define the keyframe decision function, and mark the foreground frame sequence based on the keyframe decision function to generate foreground keyframes.

6. The joint coding compression method for videos from multiple adjacent cameras according to claim 1, characterized in that, S2 include: S21. Input the foreground keyframes from the foreground main semantic library into the OpenPose network to extract the skeleton model of 18 key points as foreground motion features, and simultaneously extract the occupancy set of foreground pixels as spatial features. S22. After mapping the RGB images of the foreground keyframes and background keyframes in the foreground main semantic library to the antagonistic color space, apply the Laplacian operator to each channel to extract the double antagonistic feature map, and generate the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map. S23. Perform discrete cosine transform and quantization, zigzag scanning and entropy coding on the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map; Inverse entropy encoding is performed at the decoding end, and the recovered frequency domain coefficients and double antagonistic semantic feature map are output.

7. A joint coding compression method for videos from multiple adjacent cameras according to claim 6, characterized in that, S3 include: S31. The multi-channel edge information of the recovered dual-antagonistic semantic feature map is pulse set encoding, the surface features are filled iteratively through the discrete diffusion dynamics equation, and the three channels are fused by perceptual weights and then mapped back to the RGB space through inverse transformation to generate the reconstructed color foreground keyframe sequence and the reconstructed background keyframe. S32. Perform multi-source feature joint encoding and initialization on the reconstructed color foreground keyframe sequence to generate initial image encoding features; input the initial image encoding features and motion features into the pose attention transfer architecture to generate a complete foreground frame sequence; S33. Fuse the complete foreground frame sequence and the reconstructed key background frames to output the reconstructed complete video frame sequence.

8. A joint coding compression method for videos from multiple adjacent cameras according to claim 6, characterized in that, S22 includes: The RGB images of foreground and background keyframes from the foreground main semantic library are mapped to an antagonistic color space using a linear transformation matrix, decomposing them into three independent channels: red-green, blue-yellow, and luminance. A Laplacian operator is then applied to each of the three channels. Perform second-order derivative filtering to obtain the double-antagonistic feature map of each channel.

9. A joint coding compression method for videos from multiple adjacent cameras according to claim 8, characterized in that, S31 includes: For each pixel location, three independent sets of spiking neurons are configured for each channel, with each set containing several LIF neurons. A frequency coding strategy is adopted so that the firing rate of neurons in each set is proportional to the intensity of the corresponding input edge. In the spiking neural network architecture, each neuron interacts with its neighboring neurons through horizontal connections. Through iterative updates over multiple time steps, surface feature filling is completed, resulting in reconstructed surface feature maps for each channel. Based on adjustable perceptual weights and surface feature maps, a fused antagonistic channel vector is generated. The fused antagonistic channel vector is then mapped back to the RGB color space.

10. A joint encoding and compression system for videos from multiple adjacent cameras, characterized in that, include: The data acquisition and semantic deduplication module is used to acquire raw video frame sequences from multiple cameras, and to perform semantic deduplication on the raw video frame sequences using a spiking neural network to generate a foreground main semantic library. The feature extraction and cross-modal coding module is used to: extract and encode foreground motion features and spatial features based on foreground frame sequences; Based on the dual antagonism mechanism, foreground dual antagonism semantic feature maps and background dual antagonism semantic feature maps are extracted from foreground keyframes and background keyframes in the foreground main semantic library; Frequency domain encoding and decoding are performed on the foreground double antagonistic semantic feature map and the background double antagonistic semantic feature map; the encoded foreground motion features, encoded spatial features, foreground double antagonistic semantic feature map and background double antagonistic semantic feature map are integrated and output as a cross-modal feature set; The video reconstruction module is used to reconstruct foreground and background images based on a spiking neural network according to a cross-modal feature set. A complete foreground frame sequence is generated based on a pose transfer network; cross-modal feature fusion is performed to reconstruct the video, and the reconstructed complete video frame sequence is output. The aforementioned joint coding compression system for videos from multiple neighboring cameras is used to implement the joint coding compression method and steps for videos from multiple neighboring cameras as described in claim 1.