Satellite communication gateway data processing method based on multi-modal large model
Patent Information
- Application Number
- CN202611301474.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
为克服现有卫星通信网关数据处理方法所存在的窄带信道带宽严重受限、数据回传冗余度高以及边缘语义理解能力缺失等技术问题,本发明提供一种适配窄带信道、精简回传数据、增强边缘语义理解的基于多模态大语言模型的卫星通信网关数据处理方法
本发明通过在卫星通信网关侧同步处理音视频数据并执行跨模态联合语义理解,实现了对监测事件的端侧闭环判别,摆脱了对中心平台分析能力的依赖;在此基础上,模型生成的语义重要性图或分数不仅支撑图像与音频的差异化编码,输出比特流形式的语义紧凑表,显著降低同等事件信息量所需的回传数据量,更与实时获取的北斗或NTN信道状态共同参与压缩目标决策,使每一次上行传输的数据长度均能精准适配窄带信道的实际承载能力,从而达到了适配窄带信道、精简回传数据、增强边缘语义理解的效果。
Smart Images

Figure CN122824286A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite communication and edge intelligence cross-technology, and specifically relates to a satellite communication gateway data processing method based on a multimodal large model. Background Technology
[0002] In scenarios such as emergency communication, forest fire prevention, power line inspection and IoT monitoring in remote areas, terrestrial mobile communication networks often have coverage blind spots. Existing technologies usually use Beidou short message communication or non-terrestrial networks (NTN) as wide-area backhaul channels to achieve data reporting in areas without network coverage. However, this technical approach has the following drawbacks: The bandwidth of narrowband channels is severely limited: the single transmission capacity of Beidou short message is only 229 bytes for a Level 2 card, and the typical rate of NTN narrowband link is less than 10 kbps, which cannot carry the original video stream or the original audio stream, resulting in a lack of on-site situational awareness. High data backhaul redundancy: Traditional gateways only perform data aggregation and forwarding functions, uploading all the collected raw audio and video data or simple threshold alarm information, which includes a large amount of low-value content such as static images and environmental noise, which not only exacerbates satellite link congestion but also increases the parsing burden on the central platform. Lack of edge semantic understanding capabilities: Existing gateways do not integrate models with multimodal content understanding capabilities, and cannot perform joint semantic analysis of video frames and audio streams locally. They also cannot independently determine the authenticity of events and extract key elements based on task monitoring content. These key elements include timestamps, geographic coordinates, event types, and event text descriptions, resulting in coarse-grained reported information and low effective information density, making it difficult to support accurate GIS map plotting and hierarchical response. Summary of the Invention
[0003] (a) Technical problems to be solved To overcome the technical problems of existing satellite communication gateway data processing methods, such as severely limited narrowband channel bandwidth, high data backhaul redundancy, and lack of edge semantic understanding, this invention provides a satellite communication gateway data processing method based on a multimodal large language model that adapts to narrowband channels, simplifies backhaul data, and enhances edge semantic understanding.
[0004] (II) Technical Solution This invention is achieved through the following technical solution: This invention proposes a data processing method for satellite communication gateways based on a multimodal large model. The data processing method is implemented based on a satellite communication gateway, which is configured with a multimodal large language model. The processing method includes: S1: The satellite communication gateway receives video and audio data; S2: The satellite communication gateway inputs the video data and audio data into the multimodal large language model, which performs video content recognition and inputs the task monitoring content into the multimodal large language model to determine whether a monitoring event has occurred; S3: If a monitoring event is determined to have occurred, the satellite communication gateway generates structured data containing timestamps, geographic coordinates, event type, and event text description, and stores it in the pending record; S4: Once the satellite communication gateway generates a record to be sent, the sending task is activated; S5: The satellite communication gateway reads the data in the record to be transmitted and obtains the signal status of the currently available communication channels; S6: The satellite communication gateway inputs the data in the record to be transmitted and the signal status into the multimodal large language model, and the multimodal large language model outputs the adapted communication channel and the corresponding compressed data length; S7: The satellite communication gateway determines whether the length of the compressed data matches the selected communication channel. If they do not match, it returns to step S5. If they match, it encapsulates the data into the format corresponding to the selected communication channel and sends it to the central platform. S8: The central platform extracts geographic coordinates and event text descriptions, plots corresponding markers on the GIS map, and triggers alarms.
[0005] Preferably, in step S2, the multimodal large language model performs semantic compression on the input video and audio data, including the following collaborative processing steps: S21: Parse video data into image sequences or video blocks, and parse audio data into spectrograms or speech feature vectors; S22: Perform object detection, scene recognition and text extraction on image sequences or video blocks, perform speech recognition, speaker recognition, emotion recognition and event detection on spectrograms or speech feature vectors, and complete multimodal semantic analysis; S23: Generate a semantic importance graph or semantic importance score based on the multimodal semantic analysis; S24: Input the semantic importance graph or semantic importance score and communication channel information into the attention-based policy network, and output the appropriate semantic data length; S25: Based on the semantic data length, the image sequence is encoded using non-uniform coding based on importance graphs or semantic masking coding based on Transformer, and the speech feature vector is encoded using variable rate coding or semantic token stream coding. S26: Output a semantically compact table in bitstream form.
[0006] Preferably, the multimodal large language model is a CLIP variant, a Flava variant, or an ImageBind variant.
[0007] Preferably, the satellite communication gateway integrates a BeiDou satellite communication unit, an NTN satellite communication unit, a camera, a microphone unit, a LoRa communication unit, a WiFi communication unit, and a digital intercom unit. The satellite communication gateway establishes local wireless access connections with several near-end devices via the LoRa communication unit, WiFi communication unit, or digital intercom unit, and establishes a wide-area narrowband backhaul connection with the central platform via the BeiDou satellite communication unit or NTN satellite communication unit. The near-end devices are equipped with one or more of the following: a camera, a microphone unit, a LoRa communication unit, a WiFi communication unit, and a digital intercom unit, used to collect on-site audio and video signals or status information, and upload raw data streams to the satellite communication gateway. Preferably, the data stored in the record to be sent includes timestamps, geographic coordinates, event type, and event text description.
[0008] Preferably, when the central platform draws corresponding markers using GIS maps, the marker style includes one or more of the following: flashing icons, heat maps, or pop-up prompts.
[0009] Preferably, the task monitoring content is issued and dynamically updated by the central platform, and stored in the local cache after being received by the satellite communication gateway for use by the multimodal large language model for judgment.
[0010] Preferably, the satellite communication gateway inputs the signal values of the NTN satellite communication unit, the signal values of the Beidou satellite communication unit, and the data in the record to be transmitted into the multimodal large language model, and the multimodal large language model outputs the compressed data length adapted to the current communication channel transmission mode and rate.
[0011] Preferably, when there are multiple records to be sent, the satellite communication gateway prioritizes each record according to the event type and executes the sending task in the order of priority.
[0012] (III) Beneficial Effects Compared with the prior art, the present invention has the following advantages: This invention achieves closed-loop discrimination of monitored events at the end side by synchronously processing audio and video data and performing cross-modal joint semantic understanding at the satellite communication gateway side, thus eliminating the dependence on the analysis capabilities of the central platform. On this basis, the semantic importance graph or score generated by the model not only supports differentiated encoding of images and audio, but also outputs a semantically compact table in bitstream form, significantly reducing the amount of backhaul data required for the same amount of event information. Furthermore, it participates in the compression target decision together with the real-time acquired BeiDou or NTN channel status, so that the data length of each uplink transmission can be accurately adapted to the actual carrying capacity of the narrowband channel, thereby achieving the effects of adapting to narrowband channels, simplifying backhaul data, and enhancing edge semantic understanding. Attached Figure Description
[0013] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart of the data processing method described in this invention.
[0014] Figure 2 This is a flowchart illustrating the collaborative processing of semantic compression of input video and audio data by the multimodal large language model described in this invention.
[0015] Figure 3 This is a schematic diagram showing the connection between the satellite communication gateway, central platform, and near-end equipment of the present invention. Detailed Implementation
[0016] In this technical solution: To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] Reference Figure 1 As shown, this invention proposes a satellite communication gateway data processing method based on a multimodal large model. The data processing method is implemented based on a satellite communication gateway configured with a multimodal large language model. The processing method includes: S1: The satellite communication gateway receives video and audio data; S2: The satellite communication gateway inputs the video data and audio data into the multimodal large language model, which performs video content recognition and inputs the task monitoring content into the multimodal large language model to determine whether a monitoring event has occurred; S3: If a monitoring event is determined to have occurred, the satellite communication gateway generates structured data containing timestamps, geographic coordinates, event type, and event text description, and stores it in the pending record; S4: Once the satellite communication gateway generates a record to be sent, the sending task is activated; S5: The satellite communication gateway reads the data in the record to be transmitted and obtains the signal status of the currently available communication channels; S6: The satellite communication gateway inputs the data in the record to be transmitted and the signal status into the multimodal large language model, and the multimodal large language model outputs the adapted communication channel and the corresponding compressed data length; S7: The satellite communication gateway determines whether the length of the compressed data matches the selected communication channel. If they do not match, it returns to step S5. If they match, it encapsulates the data into the format corresponding to the selected communication channel and sends it to the central platform. S8: The central platform extracts geographic coordinates and event text descriptions, plots corresponding markers on the GIS map, and triggers alarms.
[0018] Among them, reference Figure 2 As shown, in step S2, the multimodal large language model performs semantic compression on the input video and audio data, including the following collaborative processing steps: S21: Parse video data into image sequences or video blocks, and parse audio data into spectrograms or speech feature vectors; S22: Perform object detection, scene recognition and text extraction on image sequences or video blocks, perform speech recognition, speaker recognition, emotion recognition and event detection on spectrograms or speech feature vectors, and complete multimodal semantic analysis; S23: Generate a semantic importance graph or semantic importance score based on the multimodal semantic analysis; S24: Input the semantic importance graph or semantic importance score and communication channel information into the attention-based policy network, and output the appropriate semantic data length; S25: Based on the semantic data length, the image sequence is encoded using non-uniform coding based on importance graphs or semantic masking coding based on Transformer, and the speech feature vector is encoded using variable rate coding or semantic token stream coding. S26: Output a semantically compact table in bitstream form.
[0019] The multimodal large language model is a CLIP variant, a Flava variant, or an ImageBind variant; This enables the model output to not only include event determination results, but also generate semantic control signals that can directly drive encoder parameter adjustments, thereby constructing an end-to-end controllable processing link from raw audio and video input to bitstream-level semantic compact table output.
[0020] Among them, reference Figure 3As shown, the satellite communication gateway integrates a BeiDou satellite communication unit, an NTN satellite communication unit, a camera, a microphone unit, a LoRa communication unit, a WiFi communication unit, and a digital intercom unit. The satellite communication gateway establishes local wireless access connections with several near-end devices through the LoRa communication unit, WiFi communication unit, or digital intercom unit, and establishes a wide-area narrowband backhaul connection with the central platform through the BeiDou satellite communication unit or NTN satellite communication unit. The near-end devices are equipped with one or more of the following: a camera, a microphone unit, a LoRa communication unit, a WiFi communication unit, and a digital intercom unit, used to collect on-site audio and video signals or status information, and upload raw data streams to the satellite communication gateway. The data stored in the pending record includes timestamps, geographic coordinates, event type, and event text description.
[0021] When the central platform draws corresponding markers using GIS maps, the marker style includes one or more of the following: flashing icons, heat maps, or pop-up prompts.
[0022] The task monitoring content is issued and dynamically updated by the central platform. After being received by the satellite communication gateway, it is stored in the local cache for use by the multimodal large language model for judgment.
[0023] The satellite communication gateway inputs the signal values of the NTN satellite communication unit, the signal values of the Beidou satellite communication unit, and the data in the record to be transmitted into the multimodal large language model, and the multimodal large language model outputs the compressed data length adapted to the current communication channel transmission mode and rate.
[0024] When there are multiple pending records, the satellite communication gateway prioritizes each pending record according to the event type and executes the transmission task in the order of priority.
[0025] The fundamental principle of the satellite communication gateway data processing method proposed in this invention is to use a multimodal large language model as the intelligent kernel to achieve end-to-end closed-loop processing from raw audio and video signals to high-fidelity semantic events on resource-constrained satellite-to-ground edge nodes. The multimodal large language model does not simply call pre-trained weights, but has undergone structured adaptation for satellite communication scenarios: The input terminal synchronously receives the video stream captured by the camera and the audio stream captured by the microphone unit, which are then converted into image sequences or video blocks, spectrograms or speech feature vectors by the internal parsing module. Subsequently, cross-modal alignment and joint reasoning are performed in a unified semantic space to complete target detection, scene recognition, OCR text extraction, speech recognition, speaker separation, emotion classification and multimodal event fusion judgment. The final output not only includes event type labels, but also generates a quantifiable and spatially localizable semantic importance graph or token-level importance score, providing a physical basis for subsequent compression; Furthermore, the system architecture of the data processing method relies on a highly integrated satellite communication gateway hardware platform. Its core configuration includes a BeiDou satellite communication unit, an NTN satellite communication unit, a LoRa communication unit, a WiFi communication unit, a digital intercom unit, and a camera and microphone unit. Among them, LoRa, WiFi, and digital intercom constitute a near-field access layer, used to aggregate raw audio and video streams uploaded by surrounding sensors or mobile terminals. The BeiDou and NTN units constitute a wide-area backhaul layer, supporting dual-mode capabilities of narrowband short message and broadband transmission. The status parameters of all communication units are collected in real time and used as context input to the multimodal large model, ensuring that the model's decisions are always anchored to real link conditions. Furthermore, the entire data processing method exhibits a strong feedback closed-loop characteristic. That is, when the gateway determines that a monitoring event has occurred, it immediately generates a structured record containing timestamps, geographic coordinates, standardized event type codes, and natural language text descriptions, and stores it in a local queue to be sent. The queue is preset with priority based on event type to ensure that high-risk events are given priority in the sending schedule; before each transmission, the gateway dynamically reads the measured signal status of each communication channel and inputs it together with the content to be transmitted into the multimodal large model, which outputs the optimal channel selection and the corresponding target compression length. If the actual compression result does not match the length, a recompression-reevaluation loop is automatically triggered until the matching condition is met, and then the data is encapsulated into the frame format required by the selected channel protocol for uplink transmission. Compared with existing satellite communication data processing methods, the advantages of this invention are: Under bandwidth constraints, semantic importance-driven non-uniform image coding and semantic token stream audio coding significantly reduce the transmission bit rate required for the same amount of event information. In unstable link scenarios, the joint decision-making mechanism of channel state and semantic features greatly improves the success rate of single transmission and effectively reduces the number of retransmissions. In terms of system response timeliness, the time taken for the end side to complete the entire process of event identification, compression, encapsulation and transmission is significantly shortened, which significantly improves the timeliness and determinism of event response compared with the scheme that relies on the analysis of the central platform.
[0026] This invention achieves information alignment and mutual verification between audio and video modalities by synchronously receiving video and audio data at the satellite communication gateway side and performing joint semantic analysis by the same multimodal large language model to determine the monitored events, thereby improving the accuracy of event recognition and environmental robustness. By generating a quantifiable and spatially localizable semantic importance map or semantic importance score based on the joint semantic analysis, and using it as the basis for encoding control, the compression process of image sequence and speech feature vectors focuses on the key regions and key semantic units required for event discrimination, avoiding invalid encoding of background, silent segments and redundant textures. By inputting the structured data in the record to be transmitted and the signal status of the currently available communication channel into the same multimodal large language model, the model outputs the adapted communication channel and the corresponding compressed data length, and automatically returns to the re-evaluation step when there is a mismatch. This establishes a dynamic closed-loop mapping relationship between semantic understanding results and channel carrying capacity, ensuring the data validity of each uplink transmission. By prioritizing multiple pending records based on event type and executing the sending tasks in sequence, deterministic allocation of limited communication resources among different event urgency levels is achieved.
[0027] Example 1: Dynamic Link Adaptive Transmission for Field Geological Disaster Monitoring This embodiment is deployed in high-risk areas for geological disasters such as landslides and debris flows in mountainous areas. The satellite communication gateway integrates a camera, microphone, Beidou short message unit, NTN broadband unit and LoRa communication unit; the surrounding area is connected to rain gauges, tilt sensors and vibration sensors via LoRa to upload environmental status data in real time. When the gateway receives the video stream of the mountain captured by the camera and the audio stream of rock fracture captured by the microphone, it synchronously inputs them into the lightweight adapted CLIP variant model. The model parses the video into an image sequence and the audio into a spectrogram. In a unified semantic space, it jointly performs target detection (identifying crack propagation), scene recognition (judging it as a loose accumulation on a steep slope), OCR text extraction (reading the monitoring pile number), speech recognition (identifying brittle fracture sounds such as "crack"), emotion recognition (determining the tension of the audio), and multimodal event fusion judgment. Finally, it generates a semantic importance map and outputs the event "local instability of the mountain has occurred". The gateway then generates a structured record containing a precise timestamp, BeiDou time-synchronized geographic coordinates, event type code "LANDSLIDE_LOCAL", and text description "continuous rockfall on the eastern slope, accompanied by a high-frequency crisp sound, suspected to be a shallow landslide". This record is stored in the pending queue and is automatically prioritized for scheduling because the event type is a Level 1 warning. Before transmission, the gateway reads the BeiDou signal-to-noise ratio and NTN link delay in real time, and inputs them together with the event log into the same model; Model output: The current BeiDou signal is stable but the bandwidth is extremely narrow. It is recommended to enable the BeiDou short message channel and control the length of the compressed data to within 229 bytes. Based on this, the gateway should use non-uniform encoding based on importance graphs for the image sequence (only retaining high-precision frames of the slope movement area) and semantic token stream encoding for the spectrogram (only extracting the "break band energy surge" feature) to generate a 229-byte bit stream, which will be encapsulated in the BeiDou standard short message format and sent out. After receiving the data, the central platform analyzes the geographic coordinates, marks the corresponding location on the GIS map with a flashing red icon, and automatically overlays an orange heat map (covering a radius of 500 meters) based on the event text description, simultaneously triggering audible and visual alarms and SMS notifications.
[0028] Example 2: Collaborative Discrimination of Multi-Source Heterogeneous Events for Intelligent Border Patrol This embodiment is applied to unmanned border outposts. The gateway connects to a high-definition infrared PTZ camera and a directional microphone array via WiFi, and connects to the audio and video transmission streams of patrolling soldiers' terminals via a digital intercom unit. It also has Beidou + NTN dual-mode transmission capability. The gateway synchronously receives the video stream from the PTZ camera and the audio stream of distant human voices collected by the microphone array. It inputs the data into the Flava variant model. On the video side, it performs target detection (identifying the outline of the person crossing the border), scene recognition (determining it to be a nighttime shrubland boundary zone), and OCR to extract the boundary marker text. On the audio side, it performs speech recognition (transcription of "someone is approaching"), speaker recognition (matching the pre-stored border guard voiceprint database), emotion recognition (determining it to be a vigilant tone), and event detection (identifying abnormal footstep rhythm). After multimodal fusion, the model generates a semantic importance score and determines it as "suspected border crossing behavior". At the same time, it outputs the keyframe ROI coordinates and the start and end points of key speech segments. The gateway generates a structured record containing a timestamp, WGS-84 geographic coordinates, event type code "CROSSING_SUSPECTED", and text description "In the area of XX.XXXX°N latitude and XX.XXXX°E longitude, infrared video detected two unmarked individuals crossing the border, and the synchronized audio contained unauthorized voices and abnormal footsteps." This record is entered into the pending queue and is preset to the highest priority as a "border crossing" event. Before transmission, the gateway detected a sudden congestion in the NTN link, while the BeiDou signal was good. After comprehensive evaluation by the model, the output was: enable the BeiDou channel, target compression length ≤ 229 bytes, the gateway performs Transformer semantic masking encoding on the key frame ROI region (only retaining the difference between human body hot spots and background regions), and uses variable rate encoding (focusing on the 1–3kHz frequency band) on key speech segments to generate a 229-byte bit stream, which is then encapsulated and transmitted. After receiving the data, the central platform marks the corresponding coordinates on the GIS map with a flashing yellow icon, and generates a light blue heat map based on the "2 people" in the event text and the geographical coordinates, to assist commanders in assessing the scale and path of the infiltration.
[0029] Example 3: Efficient Backhaul of Semantic Compact Tables for Equipment Inspection in Large Hydropower Stations This embodiment is deployed in a remote mountainous hydropower station. The gateway collects video of the turbine unit's operation through a camera, collects audio of bearing vibration through a microphone unit, and accesses data streams from sensors such as temperature and pressure via LoRa. The backhaul relies on BeiDou short message service (narrowband, low power consumption). The gateway inputs the unit's video stream and vibration audio stream into the ImageBind variant model. The video is parsed into video blocks, and the audio is parsed into speech feature vectors. The model performs target detection (identifying the oil leak point), scene recognition (determining it to be the oil filter chamber next to Unit A), and OCR to extract equipment nameplate information. On the audio side, it performs speech recognition (transing the inspector's statement "Bearing temperature is too high"), speaker recognition (confirming it to be a certified inspector), emotion recognition (determining it to be a calm statement), and event detection (identifying typical bearing fault frequency components). After multimodal analysis, the model generates a semantic importance map, focusing on the pixel blocks of the oil leak area and the fault frequency band feature vectors. The gateway generates a structured record containing a timestamp, geographic coordinates, event type code "TURBINE_BEARING_WARN", and text description "Fresh oil stains found on the floor of the oil filter chamber of Unit A; vibration spectrum of the B bearing area shows that the amplitude of the 237Hz main frequency exceeds the threshold; the inspector simultaneously reports abnormal temperature". This record enters the pending queue and is processed according to the priority in "Equipment Warning". Before transmission, the gateway reads the BeiDou signal value, the model outputs a target compressed length of ≤229 bytes, the gateway uses non-uniform encoding based on importance graph for the video block (encoding only the oil leak area and its adjacent equipment nameplates), and uses semantic token stream encoding for the vibration feature vector (encapsulating only key tokens such as "237Hz@+12dB" and "temperature>72℃"), outputs a 229-byte bit stream, and encapsulates it into a BeiDou short message for transmission; After receiving the data, the central platform marks the corresponding equipment location on the GIS map of the power plant with a flashing green icon, and generates a cyan thermal zone based on the location of the oil stain and the coordinates of the vibration source to help maintenance personnel quickly locate the faulty components.
Claims
1. A data processing method for a satellite communication gateway based on a multimodal large model, wherein the data processing method is implemented based on a satellite communication gateway, characterized in that: The satellite communication gateway is configured with a multimodal large language model, and the processing method includes: Satellite communication gateways receive video and audio data; The satellite communication gateway inputs the video and audio data into the multimodal large language model, which performs video content recognition and inputs the task monitoring content into the multimodal large language model to determine whether a monitoring event has occurred. If a monitoring event is detected, the satellite communication gateway generates structured data containing timestamps, geographic coordinates, event type, and event text description, and stores it in the pending record. Once the satellite communication gateway generates a record to be sent, it activates the sending task. The satellite communication gateway reads the data in the record to be transmitted and obtains the signal status of the currently available communication channels; The satellite communication gateway inputs the data in the record to be transmitted and the signal status into the multimodal large language model, and the multimodal large language model outputs the adapted communication channel and the corresponding compressed data length. The satellite communication gateway determines whether the length of the compressed data matches the selected communication channel. If they do not match, it returns to the step of "the satellite communication gateway reads the data in the record to be sent and obtains the signal status of the currently available communication channel". If they match, the data is encapsulated into the format corresponding to the selected communication channel and sent to the central platform. The central platform extracts geographic coordinates and event text descriptions, plots corresponding markers on a GIS map, and triggers alarms.
2. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: The multimodal large language model performs semantic compression on the input video and audio data, including the following collaborative processing steps: Video data is parsed into image sequences or video blocks, and audio data is parsed into spectrograms or speech feature vectors. Perform object detection, scene recognition, and text extraction on image sequences or video blocks; perform speech recognition, speaker recognition, emotion recognition, and event detection on spectrograms or speech feature vectors; and complete multimodal semantic analysis. Based on the multimodal semantic analysis, a semantic importance graph or semantic importance score is generated; Input the semantic importance graph or semantic importance score and communication channel information into the attention-based policy network, and output the appropriate semantic data length; Based on the semantic data length, the image sequence is encoded using non-uniform coding based on importance graphs or semantic masking coding based on Transformer, and the speech feature vector is encoded using variable rate coding or semantic token stream coding. The output is a semantically compact table in bitstream form.
3. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: The multimodal large language model is a CLIP variant, a Flava variant, or an ImageBind variant.
4. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: The satellite communication gateway integrates a BeiDou satellite communication unit, an NTN satellite communication unit, a camera, a microphone unit, a LoRa communication unit, a WiFi communication unit, and a digital intercom unit. The satellite communication gateway establishes local wireless access connections with several near-end devices through the LoRa communication unit, the WiFi communication unit, or the digital intercom unit, and establishes a wide-area narrowband backhaul connection with the central platform through the BeiDou satellite communication unit or the NTN satellite communication unit.
5. The satellite communication gateway data processing method based on a multimodal large model according to claim 4, characterized in that: The near-end device is equipped with one or more of the following: a camera, a microphone unit, a LoRa communication unit, a WiFi communication unit, and a digital intercom unit, for collecting on-site audio and video signals or status information, and uploading raw data streams to the satellite communication gateway.
6. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: The data stored in the pending record includes timestamps, geographic coordinates, event type, and event text description.
7. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: When the central platform draws corresponding markers using GIS maps, the marker styles include one or more of the following: flashing icons, heat maps, or pop-up prompts.
8. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: The task monitoring content is issued and dynamically updated by the central platform. After being received by the satellite communication gateway, it is stored in the local cache for use by the multimodal large language model for judgment.
9. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: The satellite communication gateway inputs the signal values of the NTN satellite communication unit, the signal values of the Beidou satellite communication unit, and the data in the record to be transmitted into the multimodal large language model, and the multimodal large language model outputs the compressed data length adapted to the current communication channel transmission mode and rate.
10. The satellite communication gateway data processing method based on a multimodal large model according to claim 1, characterized in that: When there are multiple pending records, the satellite communication gateway prioritizes each pending record according to the event type and executes the transmission task in the order of priority.