Semantic analysis model training method, video data transmission method, electronic equipment and computer readable storage medium
By adopting a semantic analysis model of lightweight Transformer and convolutional hybrid architecture, combining the spatiotemporal attention mechanism and dynamic feature selection module, the high-fidelity contradictions under low bandwidth in ship-strait video transmission and semantic extraction accuracy problems in complex environments are solved, and efficient and real-time video transmission is achieved.
Patent Information
- Application Number
- CN202510714521.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The prior art is difficult to meet the high fidelity needs at low bandwidth in inter-ship shore video transmission, and the model has low semantic extraction accuracy in complex marine environments, and the real-time guarantee mechanism is insufficient.
The semantic analysis model of lightweight Transformer and convolutional hybrid architecture is adopted, combining the spatiotemporal attention mechanism, dynamic feature selection module and two-stage training strategy to improve the model's adaptability to complex marine environments and real-time transmission capabilities under low bandwidth.
The critical content accuracy of video transmission under low bandwidth conditions is achieved to reach ≥95%, which improves the semantic extraction accuracy of the model in dynamic marine environments, and ensures an end-to-end delay of ≤200ms, meeting the real-time response requirements of emergency events.
Smart Images

Figure CN120236248A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and video transmission, and particularly relates to a semantic analysis model training method, a video data transmission method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Video transmission between ship and shore is a core technical means for maritime operation safety monitoring, remote collaboration, and emergency command. By real-time transmitting video data such as the surrounding environment of the ship, equipment status, and personnel operations, the shore-based command center can quickly make decisions and respond to emergencies (such as personnel falling into the water, equipment failures). However, the maritime communication network is limited by satellite links or cellular network coverage, with problems such as extremely low bandwidth (usually 10 kbps - 1 Mbps) and large delay fluctuations (500 ms - 5 s), making it difficult for traditional video transmission technologies to meet the requirements of real-time and reliability.
[0003] In the prior art, there are already related technologies for semantic analysis of videos using deep learning models; however, there are many limitations. For example: Dominance of single-task models: Existing semantic analysis models mostly focus on single tasks (such as object detection or action recognition), lacking multi-task joint modeling of objects, actions, and scenes. This results in fragmented output information of the model, and multiple inferences are required to obtain complete semantics, increasing the computational delay (single-frame processing time ≥ 100 ms). Poor scene adaptability: General video datasets (such as COCO, Kinetics) lack coverage of the marine environment (such as waves, fog, low light), resulting in a significant decrease in the semantic extraction accuracy of the model in the ship-shore scenario (experiments show that the mAP of object detection of conventional models decreases by ≥ 30% under fog interference). Lack of temporal continuity: Existing models do not explicitly constrain the coherence of semantics between adjacent frames, resulting in jumps in the transmitted semantic descriptions (such as sudden changes in object positions), affecting the authenticity of shore-based video restoration.
[0004] Meanwhile, in the existing video data transmission technologies, there are also some problems, such as: High bandwidth dependence: Although traditional video compression technologies (such as H.265) can reduce the amount of data, they still need to sacrifice resolution or frame rate (such as dropping to 240p@5fps) under extremely low bandwidth (< 50 kbps), and key details (such as small floating objects, personnel actions) are severely lost. Static encoding strategy: Existing transmission schemes adopt a fixed compression ratio and cannot dynamically adjust the semantic granularity according to bandwidth fluctuations. For example, when the bandwidth drops suddenly, complete scene information is still transmitted, resulting in a sharp increase in transmission delay (≥ 1 s). Insufficient end-to-end optimization: Existing "video → text → video" schemes (such as reverse applications based on text-to-image technologies) do not design dedicated encoding and decoding protocols for the ship-shore scenario, resulting in poor compatibility between semantic information and restoration models, and deformed or logically incorrect restored videos (such as distorted ship movement trajectories).
[0005] Combined with the deficiencies of the above technologies, it fails to solve the following core problems: the contradiction between low bandwidth and high fidelity: how to ensure that the accuracy of the key content (targets, actions, scenes) of the shore-based restored video is ≥ 95% on the premise that the transmitted data volume is compressed to 1% - 5% of the original video. The semantic robustness in complex environments: how to improve the semantic extraction accuracy of the model in dynamic marine environments (waves, fog, light changes) to avoid false detections and missed detections. The lack of a real-time guarantee mechanism: how to achieve an end-to-end delay ≤ 200 ms (including ship-end processing, transmission, and shore-based restoration) to meet the real-time response requirements for emergency events (such as a person falling into the water).
[0006] Therefore, there is an urgent need for a full-link optimization solution from model training to transmission protocols to break through the technical bottlenecks of low bandwidth, high real-time performance, and complex scene adaptation. Summary of the Invention
[0007] The object of the present invention is to provide a semantic analysis model training method, a video data transmission method, an electronic device, and a computer-readable storage medium to solve the adaptability of the network model to complex marine environments and the real-time video transmission in low-bandwidth scenarios.
[0008] The present invention provides a semantic analysis model training method for video data transmission between ships and shores, which includes the following steps: Step 1: Construct a semantic analysis model, and the model includes: An input module that receives RGB images and optical flow information and splices them into 5 channels for input; A backbone network that adopts a lightweight Transformer and a convolutional hybrid architecture MobileVit-S and embeds a spatio-temporal attention mechanism; A multi-task fusion branch, including an object detection branch, an action recognition branch, and a scene encoding branch, which respectively output object categories and positions, action labels, and scene labels; A dynamic feature selection module that dynamically closes non-critical task branches according to the current bandwidth status; Step 2: Train the semantic analysis model: adopt a two-stage training strategy, including: Stage 1: Pre-train on a general video dataset, using the AdamW optimizer and Focal Loss; Stage 2: Fine-tune on ship-shore scene data, adopt a curriculum learning strategy to gradually increase the weights of complex samples, and add a temporal smoothing loss; Among them, the MobileViT-S is composed of alternating stacks of convolutional layers and Transformer modules, and a total of 6 stages, namely Stage1 to Stage6, are included; among them, the model types of Stage1, Stage3, and Stage5 are Conv2D + BN + Swish, the model types of Stage2 and Stage4 are MobileViT Block, and the model type of Stage6 is Global Context Transformer; The spatio-temporal attention mechanism introduces 3D spatio-temporal position encoding in the Transformer layer to capture the temporal dependence between video frames. Among them, the 3D spatio-temporal position encoding consists of two parts: 2D spatial encoding and 1D temporal encoding; specifically, local window spatial attention is embedded in the Stage2 and Stage4, and global temporal attention is embedded in the Stage6.
[0009] A semantic analysis model training method as described above is further preferably that the specific structure of the MobileViT Block is: CNN → Unfold → Transformer → Fold → CNN, which fuses local-global features, divides the image into local blocks through the Unfold operation, inputs them into the Transformer to capture global relationships, and then restores the spatial structure through Fold. At the same time, depthwise separable convolution is used to replace part of the fully connected layer to reduce the number of parameters.
[0010] A semantic analysis model training method as described above is further preferably that the implementation of the dynamic feature selection module includes the following steps: Monitor the bandwidth status and classify it: Monitor the ship-shore communication bandwidth in real time, and divide the bandwidth into three levels: high, medium, and low according to predefined thresholds; Dynamically configure task branches: When the bandwidth is high, all task branches are enabled; when the bandwidth is medium, the scene encoding branch is closed, and the object detection and action recognition branches are retained; when the bandwidth is low, only the object detection branch is enabled, and the detection granularity is reduced; Introduce an adaptive adjustment mechanism: Predict the bandwidth fluctuation trend based on a lightweight decision tree model, and adjust the task branch configuration in advance; if key objects are lost for 3 consecutive frames, force the action recognition branch to start and trigger high-optimization-level transmission. A semantic analysis model training method as described above is further preferably that before training the semantic analysis model, it also includes preprocessing operations on video data to extract RGB images and optical flow information from the video data.
[0011] A semantic analysis model training method as described above is further preferably that the first stage specifically includes: Configure a general dataset, using the publicly available video dataset Kinetics-400, covering daily actions and complex scenarios; Perform data augmentation by means of random cropping, color jittering, and motion blur simulation; Set the initial learning rate of the AdamW optimizer to , with a weight decay of 0.05, a batch size of 32, and an input resolution of 256×256; Set the training objectives, with the mAP@0.5 of object detection ≥ 75%, the Top-1 accuracy of action recognition ≥ 82%, and the F1-score of scene classification ≥ 0.88.
[0012] A method for training a semantic analysis model as described above, further preferably, the second stage specifically includes: Collect real ship-shore monitoring videos and annotate them. The annotation information includes object categories, action labels, and scene labels; Perform data augmentation by simulating fog, sea wave jitter, and salt spray noise at sea; Set a curriculum learning strategy, including using only simple scene samples in the 1st - 50th epoch; gradually adding complex scene samples from the 51st - 100th epoch, with the proportion increasing from 20% to 80%; full-scale complex scene training in the 101st - 150th epoch, and adding adversarial samples; Design a multi-task joint loss function, introduce a temporal smoothing constraint loss, and the joint loss function is: , where : The Focal Loss of object detection; : The cross-entropy loss of action recognition; : The label smoothing loss of scene classification; : The temporal smoothing constraint loss; Dynamically adjust the weights of the multi-task branches: The initial weights are α = 1.0, β = 0.8, γ = 0.5, λ = 0.2, where α represents the weight corresponding to object detection, β represents the weight corresponding to action recognition, γ represents the weight corresponding to scene classification, and λ represents the weight of the temporal smoothing constraint loss; linearly increase λ as the training progresses, and finally λ = 0.6.
[0013] The present invention also discloses a video data transmission method, which is applied to data communication between a ship and the shore, and includes: deploying the semantic analysis model trained by the semantic analysis model training method as described above on the ship-side edge device; generating video frames based on the GAN+Diffusion model at the shore-based end and performing video restoration operations.
[0014] A video data transmission method as described above, further preferably, the transmission method specifically is: including the following steps: Extracting video semantic information: Using a semantic analysis model to perform semantic analysis on video data, extract key information, and generate a structured semantic description, where the key information includes: objects, actions, and scenes; Semantic information compression: Encoding and compressing the semantic description to reduce the data transmission volume; Data transmission: Sending the compressed semantic information to the shore-based terminal according to the transmission protocol; Semantic information restoration: The shore-based terminal decodes the semantic information and restores it to video frames using a generation module; Video reconstruction: Combining the restored video frames in chronological order to generate complete video data; Video post-processing: Using optical flow-guided interpolation to insert intermediate frames into low-frequency segments to increase the frame rate; Applying temporal filtering to smooth the target position; Evaluating the quality of the generated video using objective metrics.
[0015] The present invention also discloses an electronic device, including one or more processors; a storage device for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the above-mentioned semantic analysis model training method or the above-mentioned video data transmission method.
[0016] The present invention also discloses a computer-readable storage medium, on which a computer program is stored, which, when executed by a processor of an electronic device, causes the electronic device to execute the above-mentioned semantic analysis model training method or the above-mentioned video data transmission method.
[0017] The beneficial effects of the present invention are as follows: The semantic analysis model training method provided by the present invention is used for video data transmission between ship and shore. The method includes the following steps: Step 1: Construct a semantic analysis model, which includes: an input module that receives RGB images and optical flow information and splices them into 5 channels for input; a backbone network that adopts a lightweight Transformer and a convolutional hybrid architecture MobileVit-S and embeds a spatio-temporal attention mechanism; a multi-task fusion branch that includes an object detection branch, an action recognition branch, and a scene encoding branch, and outputs object categories and positions, action labels, and scene labels respectively; a dynamic feature selection module that dynamically closes non-critical task branches according to the current bandwidth status; Step 2: Train the multi-task fusion network: adopt a two-stage training strategy, including: Stage 1: Pre-train on a general video dataset, using the AdamW optimizer and Focal Loss; Stage 2: Fine-tune on ship-shore scene data, adopt a curriculum learning strategy to gradually increase the weights of complex samples, and add a temporal smoothing loss. Moreover, a video data transmission method is also provided, which deploys the trained semantic analysis model to the ship-side edge device, and generates video frames based on the GAN+Diffusion model at the shore-based end and performs video restoration operations. By constructing a multi-task fusion branch, the present invention synchronously extracts object, action, and scene semantic information in the video, and improves the adaptability of the model to complex marine environments based on a two-stage training strategy; and solves the problem of real-time video transmission in low-bandwidth scenarios, and can be applied to fields such as ship monitoring and remote collaboration, significantly improving transmission efficiency and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of the steps of the training method of the semantic analysis model of the present invention; Figure 2 It is a flowchart of the steps of the dynamic feature selection module of the present invention; Figure 3 It is a flowchart of the steps of the video data transmission method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, the training method of the semantic analysis model and the video data transmission method proposed according to the present invention, including their specific implementation manners, structures, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0022] The following specifically describes the specific solutions of the training method of the semantic analysis model and the video data transmission method provided by the present invention in conjunction with the accompanying drawings.
[0023] Please refer to Figure 1 , which shows the flowchart of the steps of the training method of the semantic analysis model provided by an embodiment of the present invention. As Figure 1 shown, the training method of the semantic analysis model provided by the embodiment of the present application, which is used for video data transmission between ship and shore, includes the following steps 1 to 2: Step 1: Construct a semantic analysis model, and the model includes: An input module that receives RGB images and optical flow information and splices them into 5 channels for input; A backbone network that adopts a lightweight Transformer and convolutional hybrid architecture MobileViT-S and embeds a spatio-temporal attention mechanism; A multi-task fusion branch that includes an object detection branch, an action recognition branch, and a scene encoding branch, and respectively outputs the object category and location, action label, and scene state; A dynamic feature selection module that dynamically closes non-critical task branches according to the current bandwidth state; It should be noted that: to solve the problems of multiple inferences and redundant calculations caused by separately calling object detection, action recognition, and scene classification models in the traditional solution, the present invention proposes a semantic analysis model, which solves the core problems of low semantic extraction efficiency, poor environmental adaptability, and insufficient transmission flexibility in the ship-shore scenario through feature sharing, joint optimization, and dynamic control. It is the technical cornerstone for the present invention to achieve the full-link optimization of "high-precision semantic extraction - low-bandwidth transmission - real-time restoration". The network model includes an input module, a backbone network, a multi-task fusion branch, and a dynamic feature selection module.
[0024] To simultaneously capture spatial static features and temporal dynamic features, thereby enhancing the model's understanding ability of video content, in the input module, the collected RGB images and optical flow information are concatenated to obtain the corresponding 5-channel input. Among them, the RGB images (3 channels) provide static spatial information such as color, texture, and object shape, such as the color of the ship and the shape of the waves; the optical flow information (2 channels) characterizes the temporal motion information by calculating the pixel displacement between adjacent frames (X / Y direction motion vectors), such as the direction of the ship's travel and the fluctuation frequency of the waves. By fusing the two, "what" and "how to move" can be modeled simultaneously. For specific channel concatenation, functions such as torch.cat in the deep learning framework PyTorch or tf.concat in TensorFlow can be used to concatenate along the channel dimension.
[0025] The backbone network of the present invention adopts a lightweight Transformer and convolutional hybrid architecture (MobileViT-S), which reduces the computational complexity while ensuring the model accuracy and adapts to the deployment requirements of shipboard edge devices. The following is the detailed structural design: 1. Input preprocessing layer Input format: 5-channel tensor (3-channel RGB + 2-channel optical flow), resolution 384×384.
[0026] Preprocessing operations: Channel normalization: For the RGB channels, the pixel values are normalized to [0,1]; for the optical flow channels, the displacement values are normalized to [-1,1].
[0027] Spatial alignment: Ensure that the optical flow strictly matches the size of the RGB image through bilinear interpolation.
[0028] 2. Feature extraction backbone The backbone network is alternately stacked with convolutional (CNN) and Transformer modules, and a total of 6 stages (Stage 1~6) are included. The specific structure is shown in Table 1: Table 1
[0029] Key module description: 1. Structure of the MobileVit Block: CNN → Unfold → Transformer → Fold → CNN. Its advantages are: (1) Local-global feature fusion. The image is divided into local blocks (such as 8×8) through the Unfold operation, and the input Transformer captures the global relationship, and then the spatial structure is restored through Fold; (2) Lightweight design: Compared with the standard ViT, the number of parameters is reduced by 70% (by replacing some fully connected layers with depthwise separable convolutions).
[0030] 2. Global Context Transformer: Its function is to perform global context modeling on the 12×12×256 feature map output by Stage5.
[0031] Here, it needs to be specifically explained that the alternating stacking of convolution and Transformer means that modules mainly based on CNN and modules mainly based on Transformer are alternately used between different stages (Stages) of the backbone network, rather than strictly restricting each stage to be a pure CNN or a pure Transformer. Its alternating logic is as follows: Pure CNN stages (Stage1 / 3 / 5): Responsible for fast downsampling, feature compression, and low-level / high-level feature extraction; Hybrid module stages (Stage2 / 4): Through the cooperation of CNN and Transformer in MobileViTBlock, local-global feature fusion is achieved; Pure Transformer stage (Stage6): Focus on long-term temporal dependency modeling. This design forms an alternating pattern of "CNN → Hybrid → CNN → Hybrid → CNN → Transformer" at the stage level. This design not only improves computational efficiency but also realizes complementary functions. Specifically: Pure CNN stages (Stage1, 3, 5) quickly process high-resolution inputs, avoiding the computational overhead of global attention; Hybrid module stages (Stage2, 4) introduce Transformer at appropriate resolutions (such as 96×96, 24×24) to balance the computational cost and semantic modeling requirements.
[0032] At the same time, the backbone network also embeds a spatio-temporal attention mechanism, which is realized by separating 3D position encoding and local-global attention, and is divided into two parts: spatial attention (within the local window) and temporal attention (cross-frame global). Specifically: For the position encoding of spatial attention, 2D sine position encoding is added to each pixel within the local window to capture the relative spatial position. The calculation range is to calculate self-attention only within the local window, reducing the computational amount, and the complexity drops from to .
[0033] For the temporal position encoding of temporal attention, 1D temporal encoding is added to the same spatial position of consecutive video frames to represent the temporal order. Cross-frame attention is to perform global self-attention calculation on the multi-frame feature sequence in the Global Context Transformer of Stage6 to capture long-term temporal dependencies (such as the movement trajectory of ships).
[0034] The specific design is as follows: local window spatial attention is embedded in the Stage2 and Stage4, and global temporal attention (cross-frame) is embedded in the Stage6. Through the above design, that is, applying the attention design only in the Stage2, Stage4, and Stage6, it is the comprehensive optimization result based on computational efficiency, feature abstraction requirements, and real-time constraints: 1. Early stage (Stage1 / 3): Convolution efficiently extracts low-level features to avoid computational explosion at high resolutions; 2. Middle stage (Stage2 / 4): Local window attention models medium-grained semantics to balance accuracy and speed; 3. Late stage (Stage6): Global attention captures long-term temporal dependencies to improve the understanding ability of actions and scenes.
[0035] Therefore, the lightweight Transformer and convolution hybrid architecture adopted by the backbone network proposed in the present invention, with the spatio-temporal attention mechanism embedded, is the core technical support for achieving the tripartite goal of "high-precision - low-latency - low-power consumption" in the ship-shore scenario. At the same time, the design goal of MobileViT-S is to minimize the number of parameters and computational amount while maintaining relatively high accuracy, to adapt to the real-time inference requirements of edge devices (such as the Jetson AGX Xavier on the ship side).
[0036] The multi-task fusion branch is that the backbone network outputs multi-scale feature maps (Stage2~Stage6), which are distributed to three branches through task adapters: the object detection branch, the action recognition branch, and the scene encoding branch, where: For the object detection branch, the input is the output of Stage4 (24×24×128) + the output of Stage5 (12×12×256); its structure is as follows: 1. Feature fusion: The Stage5 features are upsampled to 24×24 through deconvolution and concatenated with the Stage4 features (24×24×384); 2. Detection head: The YOLO-style anchor box mechanism is adopted, and the outputs are: object categories (80 categories + 10 ship-shore extended categories); bounding box categories (x1, y1, x2, y2, normalized to [0,1]); confidence scores (0~1); 3. Post-processing: Non-maximum suppression (NMS, IOU threshold 0.5) is adopted.
[0037] For the action recognition branch: The input is the output of Stage3 (48×48×64) + the temporal optical flow features (5 consecutive frames); its structure is as follows: 1. Spatio-temporal feature extraction, including spatial features (global average pooling is performed on the output of Stage3 to obtain a 64-dimensional vector) and temporal features (the optical flow fields of 5 consecutive frames are input into the GRU network, and a 128-dimensional temporal vector is output); 2. Fusion classification, including concatenating spatial and temporal features (64 + 128 = 192 dimensions) and outputting the probabilities of 20 actions through a fully connected layer (192→64→20).
[0038] Scene Encoding Branch: The input is the output of Stage6 (1×1×512 global features); its structure is as follows: 1. Multi-label classification, including Fully Connected Layer 1 (512→256, activation function Swish) and Fully Connected Layer 2 (256→32, output scene labels (such as "wave level", "visibility")); 2. The output format is a 32-dimensional vector, with each bit corresponding to a binary classification of a scene attribute (Sigmoid activation).
[0039] The Dynamic Feature Selection Module implements a branch-level control and adaptive adjustment mechanism to achieve real-time response under bandwidth fluctuations. The specific structure is as follows: The input is the intermediate feature maps of each branch (such as the 24×24×384 features of the object detection branch); Control logic: 1. Bandwidth monitoring, real-time acquisition of network bandwidth (such as 10Kbps), and dividing into gears according to thresholds; 2. Branch switch, the specific operation is: start all branches at high bandwidth; close the scene encoding branch at medium bandwidth and cache its input features (24×24×384); at low bandwidth, only enable the object detection branch and reduce the output granularity (confidence threshold 0.5→0.8); 3. Introduce an adaptive adjustment mechanism: predict the bandwidth fluctuation trend based on a lightweight decision tree model and adjust the task branch configuration in advance; if key targets are lost for 3 consecutive frames, forcibly start the action recognition branch and trigger high-optimization-level transmission.
[0040] Step 2: Train the semantic analysis model: Adopt a two-stage training strategy, including: Stage 1: Pre-train on a general video dataset, using the AdamW optimizer and Focal Loss; Stage 2: Fine-tune on the ship-shore scene data, adopt a curriculum learning strategy to gradually increase the weights of complex samples, and add temporal smoothing loss.
[0041] It should be noted that: The model training strategy includes three parts: dataset construction, two-stage training method, and loss function design. Specifically: Dataset construction: A dedicated dataset for ship-shore scenes, which collects video data of scenes including ships, marine environments, and personnel operations, and annotates semantic labels (objects, actions, scenes); Data augmentation, simulating marine environment interferences (such as fog, wave jitter) to improve the robustness of the model.
[0042] Two-stage training method: Stage 1 (Pre-training): Train a multi-task model on a general video dataset to initially learn the semantic extraction ability.
[0043] Stage 2 (Fine-tuning): Adopt a curriculum learning strategy on the ship-shore dataset to gradually increase the weights of complex scene samples and optimize the adaptability of the model to the marine environment.
[0044] The specific design of the two-stage training strategy is as follows: Stage 1: Pre-training on General Datasets Dataset Configuration: Use publicly available video datasets (Kinetics-400, AVA v2.2) that cover daily actions and complex scenarios.
[0045] Data Augmentation: Random cropping, color jittering, and motion blur simulation (intensity 5% - 15%).
[0046] Loss Function Design: Object Detection Branch: Use Focal Loss (α = 0.25, γ = 2) to address class imbalance issues (such as missed detections of small objects); Action Recognition Branch: Use cross-entropy loss (with class weights, higher weights are assigned to rare actions); Scene Classification Branch: Use label smoothing loss (smooth = 0.1) to improve generalization.
[0047] Training Parameters: Optimizer: AdamW (initial learning rate , weight decay 0.05).
[0048] Batch Size: 32 (single-GPU training), input resolution 256×256.
[0049] Training Objectives: Make the model achieve the following performance in general scenarios: Object detection mAP@0.5 ≥ 75%; Action recognition Top-1 accuracy ≥ 82%; Scene classification F1-score ≥ 0.88.
[0050] Stage 2: Fine-tuning on Ship-Shore Scenarios Dataset Construction: Collect real ship-shore surveillance videos (≥10,000 segments) and annotate the following extended information: Object categories: Add ship-shore specific categories such as "lifeboat", "buoy", "oil spill area", etc.; Action labels: Define emergency event actions such as "rope breakage", "cargo tilt", etc.; Scene labels: Annotate wave height (0.5m - 5m) and visibility (100m - 1000m).
[0051] Data Augmentation: Simulate marine environment interference: Add fog (random concentration), wave jitter (random affine transformation), and salt spray noise (SPNoise probability 5%).
[0052] Curriculum Learning Strategy: Epoch 1 - 50: Only use simple scenario samples (visibility ≥ 500m, sea wave ≤ 1m); Epoch 51 - 100: Gradually add complex scenario samples (visibility < 500m, sea wave > 1m), with the proportion increasing from 20% to 80%; Epoch 101 - 150: Full - scale complex scenario training, and add adversarial samples (generate extreme weather videos through GAN).
[0053] Adjustment of loss function: Loss function design: Multi - task joint loss, the function is as follows: , where : FocalLoss for object detection; : Cross - entropy loss for action recognition; : Label smoothing loss for scene classification; : Temporal smoothing constraint loss (penalize semantic jumps between adjacent frames).
[0054] Dynamic adjustment of multi - task weights: Initial weights: α = 1.0, β = 0.8, γ = 0.5, λ = 0.2; Linearly increase λ (weight of temporal smoothing loss) with the training progress, and finally λ = 0.6.
[0055] The present invention realizes progressive optimization from general ability to scene specialization. The two - stage strategy has clear division of labor: Pre - training: Build the initial robustness of the model to noise, avoiding "overfitting" to specific transmission conditions; Fine - tuning: Achieve precise optimization for the extremely low - bandwidth and high - dynamic characteristics of ship - shore communication.
[0056] This design ensures that the model not only has the semantic extraction ability for general scenarios but also deeply adapts to the transmission requirements of ship - shore communication, which is the core basis for the "training - transmission" full - link protection in the present invention.
[0057] Meanwhile, in order to adapt to the real - time constraints of ship - shore communication, pre - processing operations of video data are required to extract RGB images and optical flow information from the original video data.
[0058] 1. Video frame division and decoding Through hardware - accelerated decoding, such as NVDEC, decode the video stream data into an RGB frame sequence, maintaining the original 1080p respectively. Then, extract key frames according to the preset sampling rate (such as 5fps) to reduce the subsequent processing load.
[0059] 2. Optical flow calculation and optimization Based on the obtained continuous RGB frame sequence, the TV-L1 optical flow algorithm (implemented by OpenCV) is adopted to balance accuracy and computational efficiency; the RGB frames are downsampled to the model input size (such as 384×384) to reduce the amount of optical flow calculation; the optical flow calculation is accelerated through CUDA (such as using the OpticalFlow API) to achieve real-time processing.
[0060] 3. Standardization of data format Normalize the RGB images and the optical flow respectively: normalize the pixel values from [0, 255] to [0, 1]; normalize the optical flow displacement values to [-1, 1] to adapt to the model input range.
[0061] Please refer to Figure 2 which shows the flowchart of the steps of the dynamic feature selection module provided by an embodiment of the present invention. As Figure 2 shown, the dynamic feature selection module provided by the embodiment of the present application includes three steps, which are respectively: 1. Bandwidth monitoring and classification: Monitor the ship-shore communication bandwidth in real time, and divide the bandwidth into three levels of high, medium, and low according to predefined thresholds: High bandwidth: ≥100 kbps; Medium bandwidth: 10 kbps~100 kbps; Low bandwidth: <10 kbps.
[0062] 2. Dynamic configuration of task branches: High bandwidth: Enable all task branches (object detection, action recognition, scene encoding); Medium bandwidth: Close the scene encoding branch and retain the object detection and action recognition branches; Low bandwidth: Only enable the object detection branch and reduce the detection granularity (increase the confidence threshold from 0.5 to 0.8).
[0063] 3. Adaptive adjustment mechanism: Predict the bandwidth fluctuation trend based on a lightweight decision tree model and adjust the task branch configuration in advance; If key targets are lost (such as a person falling into the water) are detected for 3 consecutive frames, force the action recognition branch to be enabled and trigger high-priority transmission.
[0064] Through the implementation of the above dynamic feature selection module, non-critical branches can be dynamically closed to balance the semantic information volume and transmission efficiency under bandwidth fluctuations; At the same time, force the key branches to be enabled and perform redundant transmission to ensure the reliable handling of emergency events such as personnel safety.
[0065] Please refer to Figure 3 which shows the flowchart of the steps of the video data transmission method provided by an embodiment of the present invention. As Figure 3 shown, the video data transmission method provided by the embodiment of the present application includes five steps, which are respectively: Extract video semantic information: Use a semantic analysis model to perform semantic analysis on video data, extract key information, and Generate a structured semantic description, where the key information includes: objects, actions, and scenes; Semantic information compression: Encode and compress the semantic description to reduce the amount of data transmitted; Data transmission: Send the compressed semantic information to the shore-based terminal according to the transmission protocol; Semantic information restoration: The shore-based terminal decodes the semantic information and uses the generation module to restore it to video frames; Video reconstruction: Combine the restored video frames in chronological order to generate complete video data.
[0066] It should be noted that for the transmission of the above video data, hardware and software deployment need to be carried out in advance in the ship-shore application scenario. Deploy the above-mentioned trained semantic analysis model on the ship-side edge device; perform video restoration operations on the shore-based side, and generate video frames based on the GAN+Diffusion model.
[0067] 1. Ship-side deployment and semantic extraction: Step 1.1 Ship-side edge device configuration Hardware platform: The ship-side uses an NVIDIA Jetson AGX Xavier edge computing device with 32GB of storage and 8GB of memory.
[0068] Model deployment: Deploy the above-mentioned trained semantic analysis model (MobileViT-S architecture, 25MB in size) to the device through the TensorRT acceleration engine. The input resolution is set to 384×384, supporting real-time processing (≤35ms / frame).
[0069] Step 1.2 Video semantic extraction Input video stream: The on-board camera captures a 1080P@30fps video stream.
[0070] To adapt to the real-time constraints of ship-shore communication, preprocessing operations on the video data are required to extract RGB images and optical flow information from the original video data. It includes video frame division and decoding, optical flow calculation and optimization, and data format standardization. The above three steps have been described before and will not be elaborated here.
[0071] Semantic analysis process: Object detection: Identify targets such as ships, personnel, and lifeboats, and output the bounding box coordinates (normalized to [0,1]) and confidence levels.
[0072] Action recognition: Based on the optical flow information of 5 consecutive frames, determine the action category (such as "ship accelerating", "person waving").
[0073] Scene encoding: Classify the sea wave level (1-5 levels) and visibility (low / medium / high).
[0074] Structured output: Generate semantic descriptions in JSON format { "timestamp": 1678901234, "objects": {"class": "ship", "bbox": [0.2, 0.5, 0.8, 0.7], "confidence": 0.92}, {"class": "person", "bbox": [0.4, 0.3, 0.5, 0.6], "confidence": 0.88} , "actions": ["ship_accelerating"], "scene": {"wave_level": 3, "visibility": "low"} } 2. Semantic information compression and transmission: Step 2.1 Semantic compression Encoding protocol: Convert JSON data into Protobuf binary format, with field definitions as follows: message SemanticData { uint64 timestamp = 1; repeated Object objects = 2; repeated string actions = 3; Scene scene = 4; } Compression algorithm: First-level semantics (object information): Huffman coding (the pre-trained dictionary is based on high-frequency words in ship-shore scenarios); Second-level semantics (scene information): Compressed by the LZ77 algorithm, with a compression rate ≥ 60%.
[0075] Step 2.2 Data transmission Transmission protocol: Based on the MQTT protocol, supporting low-bandwidth adaptation.
[0076] Chunked transmission: Split the Protobuf data into 1KB data chunks, and append CRC check codes to each chunk.
[0077] Priority Scheduling: Key semantics (such as "person falling into water") are marked as high priority and sent first.
[0078] Bandwidth Adaptation: Bandwidth > 100 kbps: Transmit the complete semantics (target + action + scenario); Bandwidth 10 - 100 kbps: Only transmit the target and action; Bandwidth < 10 kbps: Only transmit high-confidence targets (confidence ≥ 0.8).
[0079] 3. Shore-based Video Restoration and Reconstruction: Step 3.1 Semantic Decoding Data Verification: Restore the complete Protobuf data through CRC verification and decode it into a structured semantic description (including target, action, and scenario labels).
[0080] Error Handling: If the packet loss rate > 5%, trigger a retransmission request.
[0081] Step 3.2 Video Frame Restoration The core of this step is to restore the received structured semantic information (such as target, action, and scenario labels) into high-quality video frames through a hybrid architecture of Generative Adversarial Network (GAN) and Diffusion Model (Diffusion Model). Specifically, the action information is explicitly incorporated into the generation process in this step to ensure cross-frame motion coherence and physical rationality. The following are the specific implementation details: Step 3.2.1 Generative Model Architecture Design Overall Architecture: Adopt a hybrid design of StyleGAN2 + Latent Diffusion Model (LDM).
[0082] GAN Branch (Target Generation): Input: Target category, location, and action labels; Action Condition Injection: class ActionEncoder(nn.Module): def __init__(self): super().__init__() self.embed = nn.Embedding(num_actions, 256) self.fc = nn.Linear(256, 512) def forward(self, action_label): action_embed = self.embed(action_label) return self.fc(action_embed) Conditional Encoding: Map the target semantics (category, location) to a latent vector (Latent code). The code is as follows: class ObjectEncoder(nn.Module): def forward(self, class_label, bbox): class_embed = nn.Embedding(num_classes, 512)(class_label) pos_embed = fourier_embedding(bbox) latent = torch.cat([class_embed, pos_embed], dim=1) return latent Input of StyleGAN2 Generator: The base latent vector z is sampled from a normal distribution; the semantic latent vector c modulates the style parameters of the generator through a fully connected layer; fuse the base latent vector z, semantic latent vector c (category + location), and action latent vector . Among them, the base latent vector z is a random noise vector used to control the diversity of the generated content (such as object pose, texture details); the semantic latent vector c is a conditional vector encoded from the target semantics (category, location) used to control the semantic attributes of the generated content (such as object category, location, size); the action latent vector is a conditional vector encoded from the action label used to explicitly control the action attributes of the generated content (such as a ship accelerating, a person waving).
[0083] Diffusion Branch (Background Generation): Based on LDM, progressively denoise to generate a dynamic background (such as waves, sky), supporting fine-grained texture control.
[0084] Input: Decoded semantic information (object category, location, action label, scene description); Output: A single-frame image (resolution 1080p), frame rate 30fps (boosted to 90fps through frame interpolation).
[0085] Generation Process: 1. Generate the initial image of the object (256×256) based on the latent vector; 2. Gradually increase the resolution to 1024×1024 through the upsampling module; 3. Apply Adaptive Instance Normalization (AdaIN) to fuse scene lighting conditions (such as fog density).
[0086] Spatial constraint: According to the bounding box coordinates in the semantics, precisely crop and fit the generated target to the specified area of the background.
[0087] The specific implementation details are as follows: Scene encoding: Input scene labels (such as "wave level 3", "low visibility") into the CLIP text encoder to generate scene description vectors.
[0088] Diffusion model input: Initial noise image (resolution 1024×1024); the scene description vector guides the denoising process through the cross-attention mechanism. Denoising steps: Use 1000 steps of DDPM (Denoising Diffusion Probabilistic Models) to gradually refine the background details; Key optimizations: Dynamic noise scheduling, adaptively adjust the noise intensity according to scene complexity (such as increasing noise diversity in high wave level areas, adding wake details when the ship accelerates); Local refinement, perform an additional 50 steps of local denoising on the periphery of the target area (such as the wake of the ship) to enhance realism.
[0089] Step 3.2.2 Optical Flow-Guided Temporal Generation Predict the optical flow field of future frames based on action labels to guide target displacement.
[0090] The optical flow constraint loss is:
[0091] Step 3.2.3 Physical Constraint Modeling Apply physical constraints to the generated target displacement, such as maximum speed limit.
[0092] def apply_speed_limit(prev_bbox, curr_bbox, max_speed): dx = curr_bbox.x - prev_bbox.x dy = curr_bbox.y - prev_bbox.y speed = np.sqrt(dx××2 + dy××2) if speed>max_speed: dx ×= max_speed / speed dy ×= max_speed / speed return dx, dy Step 3.2.4 Target-background Fusion and Temporal Consistency Optimization Step 3.2.4.1 Optical Flow-guided Interpolation Input: Two consecutive generated images and their optical flow fields .
[0093] The interpolation method is as follows: def interpolate(I0, I1, F, alpha=0.5): F0_1 = alpha × F F1_0 = -(1 - alpha) × F I0_warp = warp(I0, F0_1) I1_warp = warp(I1, F1_0) return (I0_warp + I1_warp) / 2 Output: Insert an intermediate frame (e.g., ), increasing the frame rate from 30fps to 90fps.
[0094] Step 3.2.4.2 Temporal Filtering Kalman filtering: Smoothly filter the target position coordinates to eliminate generated jitter; Motion consistency loss: Constrain the motion acceleration of the target in adjacent frames during the generation process.
[0095] Step 3.3 Video Reconstruction The goal of this step is to combine the discrete video frames restored at the shore-based end into a continuous video stream in chronological order, ensuring temporal integrity. It includes the following three steps: Step 3.3.1 Frame Buffering and Sorting Receive the video frames restored at the shore-based end (single-frame resolution is 1080p), with each frame attached with a timestamp (parsed from semantic information); Sort the frames according to the timestamps. If out-of-order frames are detected (e.g., due to network jitter), trigger the rearrangement mechanism.
[0096] Step 3.3.2 Frame Rate Synchronization According to the preset output frame rate (e.g., 30fps), calculate the theoretical duration of each frame (33.3ms / frame); If the time interval between adjacent frames exceeds the threshold (e.g., 50ms), insert duplicate frames or interpolated frames to fill the gap.
[0097] Step 3.3.3 Generate Video Stream Encode the ordered frame sequence into a standard video format (such as MP4, HLS) using the FFmpeg or GStreamer library.
[0098] Step 3.4 Video Post-Processing The goal of this step is to improve the fluency and quality of the generated video stream and verify its compliance with application standards.
[0099] Temporal Consistency Optimization: Use optical flow-guided interpolation to insert intermediate frames into low-frame-rate segments and increase the frame rate to 90fps; Apply temporal filtering (such as Kalman filtering) to smooth the target position coordinates and eliminate generated jitter; Quality Assessment: Use objective metrics (peak signal-to-noise ratio PSNR, structural similarity SSIM, and perceptual similarity LPIPS) to evaluate the quality of the generated video.
[0100] Adversarial Post-Processing: Apply the super-resolution model ESRGAN to enhance details in low-quality regions (such as blurred targets); use a pre-trained PatchGAN discriminator for adversarial fine-tuning to improve the authenticity of the generated images.
[0101] Step 3.5 Real-Time Assurance Technology Step 3.5.1 Distributed Generation Pipeline GPU Cluster Division of Labor: Node1 is responsible for target generation (GAN); Node2 is responsible for background generation (Diffusion); Node3 is responsible for fusion and interpolation.
[0102] Step 3.5.2 Caching and Pre-Generation Predictive Rendering: Based on the semantics of future frames predicted by LSTM, pre-generate 3 frames for caching; Dynamic Load Balancing: Dynamically allocate generation tasks according to GPU load (such as pre-generating non-critical backgrounds during low load).
[0103] The ship-shore video data transmission method provided by the present invention achieves the following technical effects: 1. Hybrid Generation Architecture: Through the hybrid generation architecture GAN+Diffusion, it can not only focus on high-resolution target generation but also refine the background, taking into account both efficiency and quality; ensure temporal smoothness through optical flow interpolation and Kelman filtering. 2. Semantic-Generation Deep Coupling: Semantic labels directly modulate the generation style parameters to achieve precise and controllable video synthesis; the scene description vector guides the denoising of the diffusion model to enhance the environmental realism. 3. Real-Time Optimization: Distributed pipeline + predictive rendering, breaking through the single-device computing power bottleneck; dynamic load of the GPU cluster to ensure a high frame rate output of 90fps.
[0104] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for training a semantic analysis model, which is used for video data transmission between ship and shore, characterized in that, The method includes the following steps: Step 1: Construct a semantic analysis model, which includes: An input module that receives RGB images and optical flow information and splices them into 5 channels for input; A backbone network that adopts a lightweight Transformer and convolutional hybrid architecture MobileViT-S and embeds a spatio-temporal attention mechanism; A multi-task fusion branch that includes an object detection branch, an action recognition branch, and a scene encoding branch, and outputs the object category and location, action label, and scene label respectively; A dynamic feature selection module that dynamically closes non-critical task branches according to the current bandwidth status; Step 2: Train the semantic analysis model: Adopt a two-stage training strategy, including: Stage 1: Pre-train on a general video dataset, using the AdamW optimizer and Focal Loss; Stage 2: Fine-tune on ship-shore scene data, adopt a curriculum learning strategy to gradually increase the weights of complex samples, and add a temporal smoothing loss; Among them, the MobileViT-S is alternately stacked by convolutional layers and Transformer modules, and a total of 6 stages Stage1~Stage6 are included; among them, the model types of Stage1, Stage3, and Stage5 are Conv2D+BN+Swish, the model types of Stage2 and Stage4 are MobileViT Block, and the model type of Stage6 is Global Context Transformer; The spatio-temporal attention mechanism introduces 3D spatio-temporal position encoding in the Transformer layer to capture the temporal dependence between video frames. Among them, the 3D spatio-temporal position encoding consists of two parts: 2D spatial encoding and 1D time encoding; specifically designed to embed local window spatial attention in the Stage2 and Stage4 stages and global time attention in the Stage6 stage.
2. The semantic analysis model training method according to claim 1, wherein The specific structure of the MobileViT Block is: CNN → Unfold → Transformer → Fold → CNN, which fuses local-global features, slices the image into local blocks through the Unfold operation, inputs them to the Transformer to capture global relationships, and then restores the spatial structure through Fold. At the same time, depthwise separable convolutions are used to replace some fully connected layers to reduce the number of parameters.
3. The semantic analysis model training method according to claim 1, characterized in that The implementation of the dynamic feature selection module includes the following steps: Monitor the bandwidth status and classify it: Real-time monitor the communication bandwidth between the ship and the shore, and divide the bandwidth into High, medium, and low levels; Dynamically configure task branches: When the bandwidth is high, enable all task branches; when the bandwidth is medium, close the scene encoding branch, Retain the object detection and action recognition branches; when the bandwidth is low, only enable the object detection branch and reduce the detection granularity; Introduce an adaptive adjustment mechanism: Predict the bandwidth fluctuation trend based on a lightweight decision tree model, and adjust the task branch Configuration in advance; if key objects are detected to be lost for 3 consecutive frames, force the action recognition branch to be started and trigger high-priority transmission.
4. The semantic analysis model training method according to claim 1, wherein Before training the semantic analysis model, it also includes preprocessing operations on video data to extract RGB images and optical flow information from the video data.
5. The semantic analysis model training method according to claim 1, wherein The specific content of the first stage includes: Configure a general dataset, using the publicly available video dataset Kinetics-400, which covers daily actions and complex scenarios; Adopt data augmentation through random cropping, color jittering, and motion blur simulation; Set the initial learning rate of the AdamW optimizer to , the weight decay to 0.05, the batch size to 32, and the input resolution to 256×256; set the training objectives, with the mAP@0.5 of object detection ≥ 75%, the Top-1 accuracy of action recognition ≥ 82%, and the F1-score of scene classification ≥ 0.
88.
6. The semantic analysis model training method according to claim 1, wherein The specific content of the second stage includes: Collect real ship-shore monitoring videos and annotate them. The annotation information includes target categories, action labels, and scene labels; Simulate fog, sea wave jitter, and salt spray noise at sea for data augmentation; Set a curriculum learning strategy, including using only simple scene samples in the 1st - 50th epoch; gradually adding complex scene samples from the 51st - 100th epoch, with the proportion increasing from 20% to 80%; full-scale complex scene training in the 101st - 150th epoch and adding adversarial samples; Design a multi-task joint loss function, introduce a temporal smoothing constraint loss, and the joint loss function is: , where : Focal Loss for object detection; : Cross-entropy loss for action recognition; : Label smoothing loss for scene classification; : Temporal smoothing constraint loss; Dynamically adjust the weights of multi-task branches: the initial weights are α = 1.0, β = 0.8, γ = 0.5, λ = 0.2, where α represents the weight corresponding to object detection, β represents the weight corresponding to action recognition, γ represents the weight corresponding to scene classification, and λ represents the weight of the temporal smoothing constraint loss; linearly increase λ as the training progresses, and finally λ = 0.
6.
7. A video data transmission method, applied to data communication between a ship and the shore, characterized in that It includes: Deploy the semantic analysis model trained by the semantic analysis model training method described in any one of claims 1 - 6 on the shipboard edge device; Generate video frames based on the GAN + Diffusion model at the shore-based end and perform video restoration operations.
8. The video data transmission method according to claim 7, wherein It includes the following steps: Extract video semantic information: Use the semantic analysis model to perform semantic analysis on video data, extract key information, and generate a structured semantic description, where the key information includes: objects, actions, and scenes; Semantic information compression: Encode and compress the semantic description to reduce the data transmission volume; Data transmission: Send the compressed semantic information to the shore-based end according to the transmission protocol; Semantic information restoration: The shore-based end decodes the semantic information and restores it to video frames using the generation module; Video reconstruction: Combine the restored video frames in chronological order to generate complete video data; Video post-processing: Use optical flow-guided interpolation to insert intermediate frames into low-frequency segments to increase the frame rate; Apply temporal filtering to Smooth the target position; Use objective metrics to evaluate the quality of the generated video.
9. An electronic device, characterized in that, It includes one or more processors; a storage device for storing one or more computer programs, which when executed by the one or more processors, enable the electronic device to implement the semantic analysis model training method described in any one of claims 1 - 6 or the video data transmission method described in any one of claims 7 - 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which when executed by the processor of the electronic device, enables the electronic device to execute the semantic analysis model training method described in any one of claims 1 - 6 or the video data transmission method described in any one of claims 7 - 8.
Citation Information
Patent Citations
End-to-end time sequence action detection method, electronic equipment and storage medium
CN117079188A
System and method for efficiently amalgamated CNN-transformer architecture for mobile vision applications
US20240193404A1
Cited By
Multi-scene disaster risk early warning method and device fusing multi-source vision and deep learning, and storage medium
CN121259678A
Multi-scene disaster risk early warning method and device fusing multi-source vision and deep learning and storage medium
CN121259678B
Dynamic visual prediction method and device for key frame in SMILE (Small Model Inter-Link Exchange), medium and program product
CN122347619A
Dynamic visual prediction method, device, medium and program product for key frames in SMILE surgery
CN122347619B