Key frame generation method and device based on semantic change degree and semantic communication system
By adaptively generating keyframes by quantizing semantic change degree, the problems of frequent transmission of redundant information and insufficient real-time performance in existing semantic communication systems are solved, and efficient and reliable semantic transmission in complex environments is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN 712 COMM & BROADCASTING CO LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing semantic communication systems lack explicit modeling of the degree of semantic change, resulting in the frequent transmission of redundant information during the semantically stable phase, while it is difficult to trigger effective transmission in a timely manner when critical semantic changes occur. This increases computational and communication costs and reduces the real-time performance and robustness of the system.
By acquiring multimodal perception data for semantic encoding, the degree of change of the current semantic state relative to the historical semantic state is quantified, key frames are adaptively determined and key frame semantic descriptions are triggered. A multi-level semantic change range and semantic change trend determination mechanism is adopted, and key frames are generated in combination with task perception information.
While ensuring the integrity of key semantic information, it reduces communication overhead and improves the stability and real-time performance of the semantic communication system in complex environments, adapting to the semantic communication needs of different scenarios.
Smart Images

Figure CN121907409A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of semantic communication technology, and in particular relates to a keyframe generation method, apparatus and semantic communication system based on semantic variability. Background Technology
[0002] In existing semantic communication systems, semantic encoding is typically performed on perceptual data such as images and videos, and semantic information is transmitted at fixed time intervals or frame rates to reduce the bandwidth overhead of raw data transmission. The receiving end often combines generative or diffusion models to reconstruct or recover the received semantic information, thereby supporting downstream perception and decision-making tasks.
[0003] However, the aforementioned semantic communication methods generally lack explicit modeling of the degree of semantic change. They typically assume that the semantics of the scene change continuously and uniformly over time, still relying on periodic triggering or simple threshold judgments to determine when to send semantic information. This method struggles to distinguish between semantically stable phases and phases of significant semantic change, easily leading to the frequent transmission of redundant semantic information when semantic changes are minor, while failing to trigger effective transmission in a timely manner when critical semantic changes occur.
[0004] Especially in semantic communication systems based on generative or diffusion models, the semantic reconstruction process has high computational and communication costs. If there is a lack of a screening mechanism for semantic changes, the system load will be further amplified, reducing the overall real-time performance and robustness.
[0005] Therefore, existing technologies lack a technical solution that can adaptively determine keyframes based on the degree of semantic change and trigger semantic generation and transmission. There is an urgent need for a keyframe generation mechanism oriented towards semantic change perception in order to improve the efficiency and reliability of semantic communication systems in dynamic scenarios. Summary of the Invention
[0006] In view of this, this application aims to propose a keyframe generation method, apparatus and semantic communication system based on semantic variability to solve at least one of the above problems.
[0007] To achieve the above objectives, the technical solution of this application is implemented as follows:
[0008] Firstly, this application provides a keyframe generation method based on semantic variability, applied to the sending end of a semantic communication system, including:
[0009] The system acquires multimodal sensing data at the current moment and performs semantic encoding processing on the multimodal sensing data to generate a current semantic representation; wherein the multimodal sensing data includes one or more combinations of visual, audio stream, and sensor time-series signal data.
[0010] The semantic change degree is determined based on the current semantic representation and the pre-acquired historical semantic information. The semantic change degree is used to quantify the degree of change of the current semantic state relative to the historical semantic state. The historical semantic information includes one or more historical semantic representations or their statistical features.
[0011] The semantic change degree is compared with a preset change threshold. Based on the comparison result, the current moment is determined to be a key frame and the generation of a key frame semantic description is triggered. Semantic encoding is performed based on the generated key frame semantic description and sent to the receiving end through the communication channel.
[0012] Secondly, based on the same inventive concept, this application also provides a keyframe generation apparatus based on semantic variability, comprising:
[0013] The perception data acquisition module is configured to acquire multimodal perception data at the current moment;
[0014] The semantic representation generation module is configured to generate the current semantic representation by performing semantic encoding processing on the multimodal perception data;
[0015] The semantic memory module is configured to store historical semantic information containing one or more historical semantic representations or their statistical features;
[0016] The semantic change degree calculation module is configured to determine the semantic change degree based on the current semantic representation and the pre-acquired historical semantic information, wherein the semantic change degree is used to quantify the degree of change of the current semantic state relative to the historical semantic state.
[0017] The keyframe determination module is configured to compare the semantic change degree with a preset change threshold, determine the current moment as a keyframe based on the comparison result, and trigger the generation of keyframe semantic description.
[0018] The semantic encoding and transmission module is configured to perform semantic encoding based on the generated keyframe semantic description and transmit it to the receiving end through a communication channel.
[0019] Thirdly, based on the same inventive concept, this application also provides a semantic communication system, including a sending end and a receiving end, wherein the sending end includes a keyframe generation device based on semantic variability as described in the second aspect.
[0020] Fourthly, based on the same inventive concept, this application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method as described in the first aspect.
[0021] Compared with existing technologies, the keyframe generation method, apparatus, and semantic communication system based on semantic variability described in this application have the following advantages:
[0022] The keyframe generation method based on semantic variability described in this application quantifies the degree of change of the current semantic state relative to the historical semantic state, and adaptively triggers the generation and transmission of keyframe semantic descriptions. This effectively reduces communication overhead and improves the stability and real-time performance of the semantic communication system in complex environments while ensuring the integrity of key semantic information. Attached Figure Description
[0023] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0024] Figure 1 This is a flowchart of a keyframe generation method based on semantic variability as described in an embodiment of this application;
[0025] Figure 2 This is a block diagram of the keyframe determination logic described in the embodiments of this application;
[0026] Figure 3 This is a block diagram of the multimodal semantic feature alignment and weighted fusion logic described in the embodiments of this application;
[0027] Figure 4 This is a schematic diagram of a keyframe generation device based on semantic variability, as described in an embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0029] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0030] The embodiments of this application are described in detail below with reference to the accompanying drawings.
[0031] Please see Figure 1 As shown, this embodiment provides a keyframe generation method based on semantic variability, applied to the sending end of a semantic communication system, specifically including the following steps:
[0032] Step S101: Obtain the multimodal sensing data at the current moment and perform semantic encoding processing on the multimodal sensing data to generate the current semantic representation; wherein, the multimodal sensing data includes one or more combinations of visual, audio stream, and sensor time-series signal data.
[0033] Specifically, in this embodiment, multimodal perception data at the current moment is acquired. The multimodal perception data includes visual, audio stream, and sensor time-series signal data. The multimodal perception data is semantically encoded to generate a current semantic representation. The semantic representation can be a continuous semantic vector, a structured semantic unit, or a scene semantic graph.
[0034] Step S102: Determine the semantic change degree based on the current semantic representation and the pre-acquired historical semantic information. The semantic change degree is used to quantify the degree of change of the current semantic state relative to the historical semantic state. The historical semantic information includes one or more historical semantic representations or their statistical features.
[0035] Specifically, in this embodiment, firstly, historical semantic information corresponding to the current perceived data is acquired. This historical semantic information serves as a benchmark for comparing the semantic state at the current moment, and its acquisition process specifically includes:
[0036] 1) Semantic consistency preprocessing and storage: During the perception process prior to the current moment, the system has processed the semantic representations generated at each historical moment through spatial alignment and dimensional unification, and stored them in the semantic memory module. Among them, the alignment process ensures that the semantic features of different moments and different modalities are mapped to the same standardized semantic metric space, thereby eliminating pseudo-semantic changes caused by device sampling differences or environmental fluctuations.
[0037] 2) Multi-scale retrieval of reference benchmarks: Based on the judgment requirements of the current task, historical reference information is dynamically extracted from the semantic memory module, including:
[0038] Instantaneous reference benchmark: retrieves the historical semantic representation immediately adjacent to the current moment to capture semantic continuity deviations;
[0039] Long-term reference benchmark: Retrieve long-term benchmarks (such as features from the previous keyframe or scene benchmark features) maintained in the semantic memory module to quantify the absolute drift amount at the current moment relative to the stable scene.
[0040] 3) Statistical integration of historical features: Temporal smoothing or feature integration processing is performed on the extracted historical semantic information from multiple frames.
[0041] In one embodiment of this invention, a time correlation operator is used to weight the retrieved historical semantic sequence to generate a statistical feature vector representing the historical semantic evolution trend. This processing aims to filter out random distortions in historical perceived data, providing a stable and reliable source of comparison for semantic variability calculation.
[0042] Specifically, in this embodiment, the retrieved length is... Multi-time-slice historical semantic sequence The temporal weights of each historical moment are calculated using a correlation operator with time decay characteristics. :
[0043] ;
[0044] In the formula, , Let represent the normalized weight coefficients, and . ; Indicates time-related factors, It is a real constant greater than 0, used to control the temporal sensitivity of the contribution of historical semantics to current memory, for example, in practical applications. The value of is usually dynamically determined based on the scenario, and the value range is usually between [0.1, 2.0]. The larger the value, the faster the system "forgets" distant history and the more it focuses on recent semantics; Represents the normalized index value of the currently processed semantic feature on the time axis, for example, in a sliding window of length . in the sequence, Values The corresponding number A historical fragment Values between; Representing historical semantic vectors With the current semantic vector Cosine similarity between them.
[0045] Subsequently, the historical semantic vectors are weighted and fused dimension by dimension according to temporal weights to obtain a statistical feature vector of fixed-dimensional historical semantic evolution trends. This vector has the same dimension as the single-time-slice semantic vector and is used to characterize temporal semantic drift and trends. Among them, the statistical feature vector... The formula is:
[0046] ;
[0047] In the formula, Indicates the first in the historical sequence The semantic feature vector at each time step.
[0048] Secondly, based on the current semantic representation and historical semantic information, the semantic change degree at the current moment is calculated to quantify the degree of change of the current semantic state relative to the historical semantic trend. The semantic change degree is calculated as follows: Figure 2 As shown.
[0049] In some implementations, semantic variability One or more of the following methods can be used for calculation:
[0050] 1) Calculation method based on semantic vector distance
[0051] When the semantic representation is expressed as a continuous semantic vector, the semantic variability is defined as the current semantic vector. With statistical eigenvectors Distance between: Among them, the distance metric function This includes Euclidean distance, cosine distance, or Mahalanobis distance.
[0052] 2) Calculation method based on semantic similarity variation
[0053] By calculating the change in similarity between the current semantic representation and the historical semantic representation, the degree of semantic change can be expressed as: The similarity index is defined as cosine similarity, inner product similarity, or normalized similarity.
[0054] 3) Calculation method based on semantic structure or target attribute changes
[0055] When semantics are represented as structured semantic units, semantic variability is calculated based on the degree of change in target category, number of targets, target attributes, or target spatial relationships.
[0056] 4) Multimodal semantic change degree fusion calculation method
[0057] For multimodal data such as vision, audio, and sensor data, a task-aware weight vector is introduced. The calculation process is as follows, and its logic block diagram is as follows. Figure 3 As shown:
[0058] Modal feature alignment: Mapping modal features of different dimensions to the same dimension through a linear mapping layer. ;
[0059] Attention weight allocation: Introducing task-aware weight vectors Calculate the attention weights for each modal semantic component. , Indicates the first Each mode in The attention weight at each moment is calculated using the following formula:
[0060] ;
[0061] in, Indicates the first Each mode in The characteristics of time, It is a task guidance vector determined by the type of downstream task (such as "obstacle avoidance" or "target tracking"); Represents the task guidance vector used for quantization. Modal feature representation or The semantic relationship between them, specifically, in this embodiment, The dot product operation is used, which measures the contribution of modal information to the current task by calculating the inner product of two vectors.
[0062] Weighted semantic change calculation: Calculate the sum of semantic changes at the current time relative to historical information by combining weights. :
[0063] ;
[0064] in, The distance metric function is defined by different metrics for different modalities: Euclidean distance for audio modality, cosine distance for visual modality, and Mahalanobis distance for sensor modality. Indicates the first Each mode in The historical characteristics of a moment, This represents the total number of modes.
[0065] 5) Calculation method of variability based on statistical characteristics of time window
[0066] When the historical semantic information is a set of semantic representations within a sliding time window, the degree of semantic change can be expressed as: ;in, and These represent the mean and dispersion of the semantic representation within the time window, respectively.
[0067] It should be noted that the above methods for calculating semantic change can be selected or combined according to specific application scenarios, perceptual modalities, or task requirements.
[0068] This embodiment introduces semantic variability to achieve adaptive control of keyframe generation, avoiding the frequent transmission of invalid or redundant semantic information; at the same time, it ensures reliable transmission of key semantic information in scenarios with limited bandwidth or poor channel conditions.
[0069] Step S103: Compare the semantic change degree with the preset change threshold, determine the current moment as a key frame based on the comparison result and trigger the generation of key frame semantic description, perform semantic encoding based on the generated key frame semantic description, and send it to the receiving end through the communication channel.
[0070] Specifically, in this embodiment, to address the problems of coarse judgment granularity and unstable key semantic triggering caused by relying solely on a single semantic change threshold for keyframe determination in existing semantic communication systems, this embodiment provides a keyframe determination algorithm based on multi-level semantic change intervals. The details are as follows:
[0071] 1) Construction of semantic change range
[0072] Specifically, during system initialization or operation, the sending end constructs at least two semantic change trigger thresholds based on historical semantic change statistics, namely the first change threshold. Second change threshold ,in .
[0073] First change threshold Second change threshold It can be determined through one or a combination of the following methods:
[0074] (1) Perform statistical analysis on the semantic change within the sliding time window, and set the mean, standard deviation or quantile;
[0075] (2) Dynamically adjust based on historical keyframe generation frequency or communication bandwidth constraints;
[0076] (3) Configure according to the current system load status or task requirements.
[0077] 2) Multi-level keyframe determination process
[0078] At the present moment The sending end obtains the current semantic change degree according to any of the aforementioned semantic change degree calculation methods. And perform keyframe determination according to the following steps:
[0079] (1) When Less than the first change threshold When the current semantic state is determined to be in the semantically stable range, it is considered that the current semantic representation has not changed significantly relative to the historical semantic information. At this time, no keyframe semantic description is generated, but the current semantic representation is written into the semantic memory module for subsequent semantic change degree calculation.
[0080] (2) When Greater than or equal to the first change threshold And less than the second change threshold When the current semantic state is determined to be in a moderate change range, a candidate keyframe semantic description is generated. This candidate keyframe semantic description can employ low-dimensional semantic vectors, simplified semantic tag sets, or summary-style semantic representations to reduce communication and computational overhead while ensuring semantic continuity.
[0081] (3) When Greater than or equal to the second change threshold When the current moment is determined to be a keyframe, the generation of a complete keyframe semantic description is triggered, and the keyframe semantic description is semantically encoded and sent to the receiving end.
[0082] This embodiment, through the aforementioned multi-level semantic change interval determination mechanism, can adaptively adjust the keyframe generation strategy according to the intensity of semantic change, significantly reduce the transmission of redundant semantic information during the semantically stable phase, and promptly trigger a complete semantic description when semantic mutations or significant changes occur, thereby effectively improving the overall efficiency and real-time performance of the semantic communication system.
[0083] In some implementations, considering that the semantic change degree calculated within a single time slice may be affected by transient noise, perceptual errors, or short-term anomalies, this embodiment further provides a keyframe determination algorithm based on semantic change trends. The specific details are as follows:
[0084] 1) Semantic change trend modeling
[0085] Specifically, the sending end calculates a semantic change degree sequence { based on the aforementioned semantic change degree calculation method over multiple consecutive time slices. , , ..., },in, The preset time window length.
[0086] The sending end performs trend analysis on the semantic change degree sequence, including but not limited to: calculating the rate of change of semantic change degree; calculating the offset of the current semantic change degree relative to the average change degree within the time window; and determining whether the semantic change degree shows trend characteristics such as continuous increase, sudden change, or increased fluctuation.
[0087] 2) Trend-driven keyframe determination logic
[0088] In this embodiment, keyframe determination is based on a comprehensive judgment combined with semantic change trends, specifically including:
[0089] (1) When the semantic change degree shows a continuous upward trend in multiple consecutive time slices, and the current semantic change degree exceeds the trend trigger threshold, the current moment is determined to be a key frame;
[0090] (2) When the current semantic change degree shows a significant sudden increase relative to the average change degree within the time window, and the sudden increase degree exceeds the preset ratio, the generation of key frame semantic description is triggered;
[0091] (3) When the rate of change of semantic change exceeds the preset rate of change threshold, it can be determined as a key frame even if the current change does not exceed the absolute threshold.
[0092] This embodiment introduces a semantic change trend determination mechanism, which can effectively suppress false triggering caused by instantaneous noise, while maintaining high sensitivity to continuous semantic evolution or sudden semantic changes, thereby improving the stability and robustness of keyframe determination.
[0093] In some implementations, semantic communication systems typically serve specific task scenarios, such as security monitoring, target tracking, or intelligent decision-making. To ensure that the keyframe determination results are consistent with the system's task objectives, this embodiment provides a task-aware adaptive keyframe determination algorithm.
[0094] 1) Obtaining task-related information
[0095] Specifically, when generating the current semantic representation, the sending end also parses the task-related semantic units contained in the semantic representation, such as key target categories, key event labels, or task priority identifiers.
[0096] 2) Task perception and judgment strategy
[0097] In this embodiment, the sending end introduces a task relevance factor based on the calculation of semantic variability to dynamically adjust the keyframe determination strategy:
[0098] (1) When the current semantic representation contains key semantic units that are highly relevant to the task, reduce the semantic change threshold to improve the sensitivity of keyframe generation;
[0099] (2) When the current semantic representation contains only semantic information that is not relevant to the task or is not critical, increase the semantic change threshold to reduce unnecessary keyframe triggering;
[0100] (3) When semantic information related to security risks, abnormal events or task priority is detected, the current moment can be directly determined as a key frame without relying on the degree of semantic change.
[0101] By introducing a task-aware mechanism, this embodiment enables the keyframe determination process to be consistent with the system application goals, ensuring that task-related semantic information is transmitted preferentially and reliably during communication.
[0102] In some implementations, in applications of intelligent driving or industrial unmanned systems, perception data typically includes video streams (vision), ambient sound (audio), and IMU / radar data (sensors). This embodiment achieves keyframe determination through the following process:
[0103] 1) Spatiotemporal semantic modeling
[0104] The semantic representation generation module utilizes a multimodal Transformer encoder to extract visual features. Audio features With sensor time-series characteristics .
[0105] 2) Introduce task guidance weights
[0106] Preset task semantic mask. If the current task is "obstacle avoidance", increase the weight of sensor features (distance change) and visual features (obstacle attributes); if the task is "environment monitoring", balance the weights of each modality.
[0107] 3) Dynamic distance measurement
[0108] For the visual modality, cosine similarity is used to measure changes in target distribution; for the audio modality, the offset between short-time energy distribution and spectral centroid is used to measure semantic jumps; and for the sensor modality, Mahalanobis distance is used to measure state anomalies.
[0109] 4) Fusion determination
[0110] The system calculates the degree of change after multimodal fusion. When visual changes are slow (such as a moving road scene) but the sensor detects a sudden vibration or the audio detects an abnormal collision sound, the attention mechanism compensates for the weight of the abnormal modality, even if the visual change is not significant. It will still quickly exceed the preset change threshold. This triggers the generation of the complete keyframe.
[0111] This embodiment achieves deep coupling between keyframe generation and actual task requirements by introducing an attention-weighted mechanism and a differentiated distance metric. In a multimodal environment, the system can automatically compensate for weak modalities (such as limited vision) and strengthen the weight of key features (such as abnormal sound or vibration sensor data), effectively avoiding the risk of missed detection caused by the limitations of single-modal perception. At the same time, this mechanism significantly reduces the communication load of perception data in the stable phase by adaptively filtering scene semantic redundancy. While ensuring the real-time performance of tasks such as target tracking and anomaly detection, it greatly improves the transmission efficiency of the semantic communication system and its environmental adaptability in complex scenarios.
[0112] In some implementations, the method also includes receiving semantic recovery quality feedback or downstream task performance indicators from the receiving end, dynamically adjusting the change threshold or keyframe determination strategy based on the feedback information, and updating historical semantic information.
[0113] Specifically, in this embodiment, during the semantic recovery or generation process, the receiving end calculates the semantic reconstruction error or downstream task performance indicators and sends the feedback information to the sending end. Based on the feedback information, the sending end dynamically adjusts the semantic change threshold or key frame determination strategy in subsequent time slices to improve the effectiveness of subsequent key frame generation.
[0114] This embodiment introduces a semantic memory feedback mechanism, which enables the key frame determination process and the semantic recovery effect at the receiving end to form a collaborative closed loop, further improving the overall performance of the semantic communication system.
[0115] This embodiment introduces semantic variability as the basis for keyframe generation, effectively avoiding the sensitivity of pixel- or motion feature-based methods to noise and local disturbances. It can highlight the real changes in scene semantic structure, target attributes, or task-related information, reduce semantic communication load, and improve the overall semantic transmission efficiency and robustness of the system. It is suitable for application scenarios such as unmanned systems, intelligent transportation, and multimodal perception.
[0116] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0117] Based on the same inventive concept, corresponding to any of the above embodiments, the embodiments of this application also provide a keyframe generation apparatus based on semantic variability.
[0118] like Figure 4 As shown, the keyframe generation device based on semantic variability includes:
[0119] The perception data acquisition module is configured to acquire multimodal perception data at the current moment; wherein, the historical semantic information includes one or more historical semantic representations or their statistical features;
[0120] The semantic representation generation module is configured to generate the current semantic representation by performing semantic encoding on multimodal sensing data;
[0121] The semantic memory module is configured to store historical semantic information containing one or more historical semantic representations or their statistical features;
[0122] The semantic change degree calculation module is configured to determine the semantic change degree based on the current semantic representation and the pre-acquired historical semantic information. The semantic change degree is used to quantify the degree of change of the current semantic state relative to the historical semantic state.
[0123] The keyframe determination module is configured to compare the semantic change degree with a preset change threshold, determine the current moment as a keyframe based on the comparison result, and trigger the generation of keyframe semantic description.
[0124] The semantic encoding and transmission module is configured to perform semantic encoding based on the generated keyframe semantic description and send it to the receiving end through the communication channel.
[0125] In some implementations, a feedback processing module is also included, configured to dynamically adjust the change threshold based on feedback information from the receiving end, and adjust the storage weight or update strategy of historical semantic information in the semantic memory module.
[0126] It also includes a semantic memory update module, which is configured to update the historical semantic information in the semantic memory module according to the current semantic representation or keyframe semantic description.
[0127] Specifically, in this embodiment, the feedback processing module is connected to the keyframe determination module and the semantic memory module, respectively, to construct a threshold adaptive closed-loop adjustment mechanism based on a PID (proportional-integral-derivative) controller. This module receives semantic recovery quality feedback (such as semantic reconstruction error or generation loss function value) and downstream task performance indicators from the receiving end in real time, and uses these as inputs to calculate a dynamic threshold offset. This offset is applied to the keyframe determination module in real time, and the semantic change trigger threshold is dynamically increased or decreased through the feedback compensation mechanism. At the same time, the semantic memory module optimizes the storage weight and update frequency of historical semantic features.
[0128] Furthermore, after completing semantic reconstruction, the receiving end calculates reconstruction quality metrics (such as the semantic similarity index SSIM or task accuracy metrics) and sends them back to the sending end. The sending end then performs the following adjustments based on the feedback information:
[0129] 1) Threshold dynamic compensation: If the reconstruction quality is continuous Frames below a preset threshold (This threshold represents the lower limit of semantic reconstruction quality, for example: the preset threshold in this embodiment) If set to 3), the current semantic change calculation is deemed too conservative, and the first and second change thresholds are automatically reduced. offset Improve keyframe trigger sensitivity.
[0130] 2) Memory weight correction: Adjust the update step size of historical features in the semantic memory module according to the reconstruction error, so as to achieve deep coupling between the key frame generation logic of the sending end and the reconstruction capability of the receiving end.
[0131] Through this closed-loop control, the device can automatically search for the optimal Pareto balance between "communication resource saving" and "perceived reconstruction accuracy" for real-time channel quality and back-end reconstruction effect, ensuring the robustness and efficiency of the semantic transmission system in complex channel environments.
[0132] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.
[0133] Based on the same inventive concept, corresponding to the apparatus of any of the above embodiments, embodiments of this application also provide a semantic communication system, including a sending end and a receiving end, wherein the sending end includes the keyframe generation apparatus based on semantic variability as described above.
[0134] The semantic communication system described above is used to implement the corresponding apparatus in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0135] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the methods described in any of the above embodiments.
[0136] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0137] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0138] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0139] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0140] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A keyframe generation method based on semantic variability, applied to the sending end of a semantic communication system, characterized in that, include: The system acquires multimodal sensing data at the current moment and performs semantic encoding processing on the multimodal sensing data to generate a current semantic representation; wherein the multimodal sensing data includes one or more combinations of visual, audio stream, and sensor time-series signal data. The semantic change degree is determined based on the current semantic representation and the pre-acquired historical semantic information. The semantic change degree is used to quantify the degree of change of the current semantic state relative to the historical semantic state. The historical semantic information includes one or more historical semantic representations or their statistical features. The semantic change degree is compared with a preset change threshold. Based on the comparison result, the current moment is determined to be a key frame and the generation of a key frame semantic description is triggered. Semantic encoding is performed based on the generated key frame semantic description, and the information is sent to the receiving end through the communication channel.
2. The keyframe generation method based on semantic variability according to claim 1, characterized in that, The method for obtaining the historical semantic information includes: Perform consistent preprocessing and storage of semantic representations generated at each historical moment; Based on the current task's judgment requirements, historical reference information is dynamically extracted at multiple scales based on the reference benchmark, including instantaneous reference benchmarks and long-term reference benchmarks. By performing temporal smoothing or feature integration on the extracted multi-frame historical semantic information, and using time correlation operators to perform weighted summation on the extracted historical semantic sequence, a statistical feature vector representing the historical semantic evolution trend is generated.
3. The keyframe generation method based on semantic variability according to claim 1, characterized in that, The semantic variability is calculated using one or more of the following methods, including: The degree of semantic change is determined based on the semantic vector distance between the current semantic representation and the historical semantic representation; The degree of semantic change is determined based on the change in semantic similarity between the current semantic representation and the historical semantic representation; Determine the degree of semantic change based on changes in semantic structure or target attributes; The semantic variability is determined based on multimodal semantic feature fusion processing, including aligning modal features of different dimensions, calculating the attention weight of each modal semantic component through the task-aware weight vector, and combining the attention weight to calculate the comprehensive semantic variability of the current semantic representation relative to historical semantic information. The degree of semantic variability is determined based on the statistical characteristics of the semantic representation within the sliding time window.
4. The keyframe generation method based on semantic variability according to claim 1, characterized in that: The semantic variability is compared with a preset variability threshold, which includes at least two variability thresholds to divide the semantic variability into a semantically stable interval, a moderate variability interval, and a significant variability interval. The keyframe determination strategy includes: In response to the semantic change degree being less than the first change threshold, it is determined that the current semantic state is in the semantically stable range, and no keyframe semantic description is generated; In response to the semantic change degree being greater than or equal to a first change threshold and less than a second change threshold, the current semantic state is determined to be in the medium change range, and a candidate keyframe semantic description is generated. In response to the semantic change degree being greater than or equal to the second change threshold, the current semantic state is determined to be in a significant change range, and a complete keyframe semantic description is generated.
5. The keyframe generation method based on semantic variability according to claim 4, characterized in that: If the current semantic representation contains semantic information related to security risks, abnormal events, or task priority upgrades, the current moment is directly determined to be a keyframe.
6. The keyframe generation method based on semantic variability according to claim 4, characterized in that, Also includes: It receives semantic recovery quality feedback or downstream task performance indicators from the receiving end, dynamically adjusts the change threshold or key frame determination strategy based on the feedback information, and updates historical semantic information.
7. A keyframe generation apparatus based on semantic variability, characterized in that, include: The sensing data acquisition module is configured to acquire multimodal sensing data at the current moment; wherein, the multimodal sensing data includes one or more combinations of visual, audio stream, and sensor time-series signal data. The semantic representation generation module is configured to generate the current semantic representation by performing semantic encoding processing on the multimodal perception data; The semantic memory module is configured to store historical semantic information containing one or more historical semantic representations or their statistical features; The semantic change degree calculation module is configured to determine the semantic change degree based on the current semantic representation and the pre-acquired historical semantic information, wherein the semantic change degree is used to quantify the degree of change of the current semantic state relative to the historical semantic state. The keyframe determination module is configured to compare the semantic change degree with a preset change threshold, determine the current moment as a keyframe based on the comparison result, and trigger the generation of keyframe semantic description. The semantic encoding and transmission module is configured to perform semantic encoding based on the generated keyframe semantic description and transmit it to the receiving end through a communication channel.
8. The keyframe generation apparatus based on semantic variability according to claim 7, characterized in that: It also includes a feedback processing module, which is configured to dynamically adjust the change threshold based on the feedback information from the receiving end, and adjust the storage weight or update strategy of historical semantic information in the semantic memory module.
9. A semantic communication system, characterized in that, It includes a sending end and a receiving end, wherein the sending end includes a keyframe generation device based on semantic variability as described in claim 7 or 8.
10. A non-transitory computer-readable storage medium, characterized in that, in, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute a keyframe generation method based on semantic variability as described in any one of claims 1-6.
Citation Information
Patent Citations
Video content similarity calculation method and device, equipment, medium and product
CN119600509A
IMU (Inertial Measurement Unit)-assisted deep SLAM (Simultaneous Localization and Mapping) method and system fusing language-vision multi-mode perception
CN120628058A
Dynamic SLAM method based on static semantic anchor points
CN121280528A
Real-time data analysis method and system based on multi-modal semantic mapping
CN121302211A
Artificial intelligence-driven multi-modal data analysis method and system
CN121479577A