Semantic communication method and system based on multi-modal perception

By introducing a semantic memory mechanism and a semantic restoration method based on context constraints, the bandwidth and latency issues of multimodal communication under complex channel conditions are solved, achieving robustness and continuity of semantic communication, which is suitable for applications such as unmanned systems, intelligent transportation, and safety monitoring.

CN121907408APending Publication Date: 2026-04-21TIANJIN 712 COMM & BROADCASTING CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN 712 COMM & BROADCASTING CO LTD
Filing Date
2026-03-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing multimodal communication solutions are highly dependent on communication bandwidth and latency under complex wireless channel conditions, making them difficult to adapt to real-world application scenarios such as low bandwidth, strong interference, or sudden packet loss. Furthermore, they lack effective memory and modeling of contextual semantic information across time and multiple frames, resulting in semantic discontinuity and insufficient scenario consistency.

Method used

A semantic memory mechanism and context-constrained semantic restoration method are introduced. By associating the current semantic representation of multimodal information with historical semantic information, a keyframe semantic description is generated. A multimodal model is used for semantic completion and restoration, and context-constrained conditions are combined for decision support.

Benefits of technology

While reducing the bandwidth requirements for multimodal data transmission, it improves the robustness and continuity of semantic communication in complex channel environments, making it suitable for applications such as unmanned systems, intelligent transportation, and safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121907408A_ABST
    Figure CN121907408A_ABST
Patent Text Reader

Abstract

The invention provides a semantic communication method and system based on multi-modal perception, and the method comprises the steps: obtaining multi-modal information, carrying out the semantic extraction of the multi-modal information, and generating a corresponding current semantic representation; performing semantic association on the current semantic representation and historical semantic information to obtain context enhanced semantic information; determining a semantic difference degree or a semantic change trend between the current semantic representation and historical semantic information based on the context enhanced semantic information to generate a key frame semantic description; semantic processing and semantic completion are carried out on the semantic description of the key frame; and according to a set context constraint condition, performing conditional semantic restoration on the complemented semantic description of the key frame by using a multi-modal model, and outputting auxiliary decision information or prompt information based on a restoration result. According to the method, the multi-modal data transmission bandwidth requirement is reduced, and meanwhile, the robustness and continuity of semantic communication in a complex channel environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of wireless communication technology, and in particular relates to a semantic communication method and system based on multimodal perception. Background Technology

[0002] With the rapid development of applications such as unmanned systems, intelligent transportation, and intelligent security, the scale of multimodal data such as images and videos that communication systems need to transmit continues to grow. Traditional communication systems aim for bit-level precision transmission, which can meet the needs in high-bandwidth, low-latency scenarios. However, under conditions of limited wireless channels, unstable links, or low power consumption, they often face problems such as insufficient bandwidth, increased latency, and decreased reliability, making it difficult to meet the application requirements with high real-time and reliability requirements.

[0003] To reduce the dependence of communication systems on bandwidth and channel quality, semantic communication, a novel communication paradigm, has been proposed in recent years. Semantic communication extracts and transmits the semantic content contained in information, rather than the complete raw multimodal data, thereby reducing the data transmission scale to some extent and improving the robustness of communication systems in complex channel environments.

[0004] However, existing multimodal communication solutions still mainly rely on the complete transmission of pixel-level or feature-level data. Even when combined with feature compression or coding optimization techniques, they still have a high dependence on communication bandwidth and latency under complex wireless channel conditions, making it difficult to adapt to practical application scenarios such as low bandwidth, strong interference, or sudden packet loss.

[0005] Furthermore, most existing semantic communication schemes focus on semantic extraction and reconstruction for single-frame images or independent semantic units, lacking effective memorization and modeling of contextual semantic information across time and multiple frames. In continuous scenes or dynamic environments, problems such as semantic discontinuity, insufficient scene consistency, or unstable semantic reconstruction results can easily occur, affecting the reliability of downstream tasks.

[0006] With the continuous improvement of the capabilities of multimodal large models in image understanding, video generation, and cross-modal reasoning, introducing them into semantic communication systems to achieve high-quality semantic reconstruction is of significant research value. However, in existing technologies, there is still a lack of mature and feasible technical solutions for effectively combining multimodal large models within a semantic communication framework and making full use of historical semantic context information to achieve continuous, stable, and low-bandwidth multimodal semantic communication. Summary of the Invention

[0007] In view of this, this application aims to propose a semantic communication method and system based on multimodal awareness to solve at least one of the above problems.

[0008] To achieve the above objectives, the technical solution of this application is implemented as follows: Firstly, this application provides a semantic communication method based on multimodal awareness, comprising: Acquire at least one type of multimodal information from images and videos, and perform semantic extraction on the multimodal information to generate a corresponding current semantic representation; The current semantic representation is semantically associated with the stored historical semantic information to obtain context-enhanced semantic information used to characterize cross-temporal contextual relationships; Based on the context-enhanced semantic information, determine the semantic difference or semantic change trend between the current semantic representation and the historical semantic information. In response to the semantic difference or semantic change trend meeting a preset threshold, generate a keyframe semantic description containing scene structure, target relationship or behavioral features. The semantic description of the keyframe is semantically processed, and the semantic description of the processed keyframe is semantically completed based on the historical semantic information. Based on the set context constraints, and using a multimodal model to perform conditional semantic restoration on the completed keyframe semantic description, auxiliary decision-making information or prompt information is output based on the restoration results.

[0009] Secondly, based on the same inventive concept, this application also provides an evasion control method, implemented based on a multimodal perception-based semantic communication method as described in the first aspect, comprising: Obtain keyframe or scene semantic information obtained through semantic information methods, wherein the keyframe or scene semantic information is a high-level semantic representation obtained through semantic extraction, semantic memory association and semantic restoration; Based on the keyframes or scene semantic representations, perform semantic-level analysis on the current environmental state to identify at least one potential risk target, risk area, or abnormal semantic event. Based on the potential risk targets, risk areas, or abnormal semantic events, and combined with preset semantic decision rules or decision models, corresponding avoidance decision information is generated. Based on the avoidance decision information, control instructions or prompts are output to guide the device to execute the avoidance suggestions.

[0010] Thirdly, based on the same inventive concept, this application also provides a semantic communication system based on multimodal awareness, the system comprising: The sending end and the receiving end exchange semantic information through a wireless communication channel: The transmitting end includes a multimodal perception module, a semantic extraction module, a first semantic memory module, a keyframe generation module, and a semantic encoding and transmission module; the receiving end includes a semantic receiving and decoding module, a second semantic memory module, a semantic restoration module, and an auxiliary decision-making module. The multimodal sensing module is configured to acquire at least one type of multimodal sensing data from images and videos, and to perform synchronization, cropping, or formatting preprocessing on the acquired data. The semantic extraction module is configured to extract semantics from the preprocessed multimodal data and generate the corresponding current semantic representation; The first semantic memory module and the second semantic memory module are configured to store and manage historical semantic information across time. The keyframe generation module is configured to generate a semantic description of the keyframe to be transmitted based on semantic changes. The semantic encoding and transmission module is configured to perform semantic encoding on the semantic description of the key frame and transmit it through a wireless channel; The semantic receiving and decoding module is configured to perform semantic decoding on the received semantic information and to perform semantic completion on the received semantic information in combination with the stored historical semantic information. The semantic restoration module is configured to perform conditional semantic restoration on the completed keyframe semantic description based on set context constraints and using a multimodal model. The auxiliary decision-making module is configured to output auxiliary decision-making information or prompt information based on the restoration result.

[0011] Fourthly, based on the same inventive concept, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect.

[0012] Fifthly, based on the same inventive concept, this application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method as described in the first aspect.

[0013] Compared with existing technologies, the semantic communication method and system based on multimodal awareness described in this application have the following advantages: The semantic communication method based on multimodal perception described in this application improves the robustness and continuity of semantic communication in complex channel environments by introducing a semantic memory mechanism and a semantic restoration method based on contextual constraints, thereby reducing the bandwidth requirements for multimodal data transmission. It is suitable for application scenarios such as unmanned systems, intelligent transportation, and safety monitoring. Attached Figure Description

[0014] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a semantic communication method based on multimodal awareness as described in an embodiment of this application; Figure 2 This is a schematic diagram of the semantic memory module structure described in an embodiment of this application; Figure 3 This is a flowchart illustrating the multimodal semantic restoration process described in an embodiment of this application. Figure 4 This is a flowchart of an evasion control method according to an embodiment of this application; Figure 5 This is a schematic diagram of a semantic communication system structure based on multimodal awareness, as described in an embodiment of this application. Figure 6 This is a schematic diagram of the hardware structure of the electronic device described in an embodiment of this application. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0016] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0017] The embodiments of this application are described in detail below with reference to the accompanying drawings.

[0018] Example 1 Please see Figure 1 As shown, this embodiment provides a semantic communication method based on multimodal awareness, which specifically includes the following steps: Step S11: Obtain at least one type of multimodal information from the image or video, and perform semantic extraction on the multimodal information to generate the corresponding current semantic representation.

[0019] Specifically, in this embodiment, the transmitting end periodically collects multimodal data, which includes image data, video data, or a combination thereof, and performs preprocessing operations such as denoising, normalization, or resolution adjustment on the multimodal data; the preprocessed multimodal data is input into the semantic extraction module, and the corresponding semantic representation is generated using a feature extraction network or a multimodal coding network.

[0020] This embodiment extracts and transmits only the semantic content of key frames in multimodal information such as images and videos, avoiding the bit-by-bit transmission of the original high-dimensional data. While ensuring the integrity of task-related semantic information, it significantly reduces the communication bandwidth usage of multimodal data and improves communication efficiency in bandwidth-constrained environments.

[0021] Step S12: Semantically associate the current semantic representation with the stored historical semantic information to obtain context-enhanced semantic information used to represent cross-temporal contextual relationships.

[0022] Specifically, in this embodiment, the current semantic representation is input into the semantic memory module and semantically associated with the historical semantic information stored in the semantic memory module to obtain context-enhanced semantic information that reflects cross-temporal contextual relationships. The semantic memory module is used to store, update, and participate in subsequent semantic communication processes for historical semantic information. Furthermore, such as Figure 2 As shown, the semantic memory module includes short-term semantic memory units and long-term semantic memory units, which are at least partially deployed at the transmitting end and / or the receiving end. The short-term semantic memory units store semantic information corresponding to adjacent time slices to reflect dynamic changes in the environment; the long-term semantic memory units store stable scene structure information, target attribute information, or prior environmental information.

[0023] In practical implementation, semantic information can be represented as semantic vectors, structured semantic units, or combinations thereof. When a new semantic representation is input into the semantic memory module, the semantic memory module updates the stored semantic information based on the semantic similarity, magnitude of change, and time decay factor between the current semantic representation and the stored historical semantic information.

[0024] For example, in some implementations, when the change in current semantic information within a preset time window is lower than a first change threshold and does not show a consistent change trend within consecutive time windows, the semantic information is determined to be short-term fluctuating semantic information, and only the short-term semantic memory unit is updated; when a certain semantic information shows a consistent change trend within multiple consecutive time windows, indicating that its semantic change is persistent and predictable, or when the change in semantic information within a consecutive time period is lower than a second stability threshold, indicating that its semantic state has tended to be stable, the semantic information is determined to have long-term validity, and it is written into the long-term semantic memory unit.

[0025] For semantic information that changes significantly but is not sustainable, the semantic memory module maintains it in the short-term semantic memory unit without triggering the update of the long-term semantic memory, so as to avoid interference from transient noise or abnormal fluctuations on long-term semantic modeling.

[0026] This embodiment introduces a semantic memory module into the semantic communication framework to perform correlation modeling on historical semantic information in the cross-time, multi-round communication process. This effectively alleviates the problems of semantic independence and context fragmentation in existing semantic communication methods, and improves the continuity, consistency and stability of semantic expression and the communication process.

[0027] Step S13: Determine the semantic difference or semantic change trend between the current semantic representation and historical semantic information based on context-enhanced semantic information, and generate keyframe semantic descriptions in response to the semantic difference or semantic change trend meeting a preset threshold.

[0028] Specifically, in this embodiment, after generating the semantic representation at the current moment, the sending end compares the semantic representation with the historical semantic information stored in the short-term semantic memory unit in the semantic memory module, and calculates the semantic difference degree between the current semantic and the semantic state in the most recent time window. The semantic difference degree is used to characterize the degree of change of the current multimodal information relative to the historical semantic state, and it can be calculated based on the distance between semantic vectors, the change in semantic similarity, the degree of change in target structure, or a combination thereof.

[0029] In some implementations, the semantic dissimilarity threshold is not a fixed value, but is adaptively set based on historical semantic change statistics. For example, the semantic dissimilarity values ​​of the most recent N time slices can be counted within a sliding time window, and the trigger threshold can be dynamically adjusted based on their mean and dispersion.

[0030] When the current semantic difference exceeds a dynamic threshold, it is determined that the current scene has undergone a significant semantic change relative to the short-term historical state, triggering the generation, encoding, and transmission of keyframe semantic descriptions. When the semantic difference does not exceed the threshold, only the short-term semantic memory unit is updated, without triggering the transmission of keyframe semantic descriptions. Furthermore, the semantic change trend used is similar to the semantic difference, used to characterize the cumulative change of the semantic change trend within consecutive time slices; similar descriptions will not be provided here.

[0031] In a further implementation, the trigger threshold can also be adjusted based on task relevance. For example, when semantic information related to security risks or critical objectives is detected, the threshold can be lowered to improve sensitivity to key semantic changes.

[0032] The keyframe semantic description includes at least one of the following: scene structure semantics; target relationship semantics; behavior or motion feature semantics; anomaly or risk-related semantics.

[0033] Step S14: Perform semantic processing on the semantic description of the keyframe and complete the semantic description of the keyframe after semantic processing based on historical semantic information.

[0034] Specifically, in this embodiment, the semantic description of the key frame is semantically encoded and sent to the receiving end through a communication channel. The receiving end performs semantic decoding on the received semantic information and performs semantic completion by combining the stored historical semantic information. The semantic encoding includes converting the semantic description of the key frame into discrete semantic symbols, low-dimensional semantic vectors or combinations thereof to adapt to different communication bandwidth conditions. Semantic completion involves inferring and completing missing, ambiguous, or incomplete semantic content based on historical semantic information stored in the semantic memory module of the receiving end.

[0035] Step S15: Based on the set context constraints, and using a multimodal model, perform conditional semantic restoration on the completed keyframe semantic description, and output auxiliary decision-making information or prompt information based on the restoration result.

[0036] Specifically, in this embodiment, such as Figure 3 As shown, at the receiving end, the semantic restoration module performs semantic restoration on the received and decoded semantic information under the contextual constraints provided by the semantic memory module. The semantic restoration module is executed by a multimodal model, which is a multimodal generation model that supports conditional generation. It can be an image generation model, a video generation model, or a cross-modal generation model, but is not limited to these.

[0037] In this embodiment, the multimodal model is not generated independently based solely on the currently received semantic information. Instead, historical semantic information from the semantic memory module of the receiving end is further introduced as a contextual constraint to guide the generated result to meet preset requirements in terms of temporal continuity and semantic consistency.

[0038] In specific implementations, historical semantic information can be represented as a sequence of semantic vectors, structured semantic units, or a combination thereof, and input as conditional information into the multimodal model. For example, in some implementations, historical semantic vectors can be embedded as contextual information and input together with current semantic information into the conditional input of the multimodal model; in other implementations, historical semantic information can be constructed as cue information or conditional control signals to constrain the generation space of the generative model.

[0039] In a further implementation, contextual constraints can be achieved through an attention mechanism, that is, using historical semantic vectors as reference information for the conditional attention module in the generative model, so that the generation process can prioritize content consistent with historical semantics, thereby reducing semantic drift or scene inconsistency issues and improving the stability and consistency of semantic restoration results.

[0040] Furthermore, the multimodal model in the semantic restoration module is a conditional multimodal generation model, which can be an image generation model, a video generation model, or a cross-modal generation model, but is not limited to these. When performing semantic restoration, the multimodal model does not generate independently based solely on the currently received semantic information, but further introduces historical semantic information from the semantic memory module as contextual constraints to guide the generation results to meet preset requirements in terms of temporal continuity and semantic consistency.

[0041] In specific implementations, historical semantic information can be represented as a sequence of semantic vectors, structured semantic units, or a combination thereof, and input as conditional information into the multimodal model. For example, in some implementations, historical semantic vectors can be embedded as contextual information and input together with current semantic information into the conditional input of the generative model; in other implementations, historical semantic information can be used as prompts or conditional control signals to constrain the generative space of the generative model.

[0042] In a further implementation, context constraints can also be achieved through an attention mechanism, that is, using historical semantic vectors as reference information for the conditional attention module in the generative model, so that the generation process can prioritize content consistent with historical semantics, thereby reducing semantic drift or scene inconsistency issues.

[0043] This embodiment combines a multimodal large model at the receiving end and performs semantic restoration under the contextual constraints provided by semantic memory. Compared with existing methods based on single-frame or local semantic reconstruction, it can effectively reduce semantic ambiguity and scene inconsistency, and improve the semantic consistency and scene rationality of the restoration results.

[0044] In this embodiment, the auxiliary decision-making information or prompting information includes at least one of the following: risk level assessment information, risk warning information, and avoidance action prompting information, thereby triggering corresponding behavioral decisions or avoidance instruction output.

[0045] In this embodiment, the above method performs semantic extraction on the acquired multimodal information such as images and videos at the sending end, and associates historical context information with the semantic memory module to generate keyframe semantic descriptions containing scene structure, target relationships, or behavioral features; the keyframe semantic descriptions are semantically encoded and sent to the receiving end through a communication channel; the received semantic information is decoded at the receiving end, and semantic completion is performed with the local semantic memory module, using a multimodal model to restore keyframe or scene information under contextual constraints; based on the restored multimodal semantic information, corresponding behavioral decisions or avoidance instructions are triggered and output.

[0046] This embodiment introduces a semantic memory mechanism and a semantic restoration method based on context constraints. This reduces the bandwidth requirements for multimodal data transmission while improving the robustness and continuity of semantic communication in complex channel environments. It is suitable for application scenarios such as unmanned systems, intelligent transportation, and safety monitoring.

[0047] Example 2 Based on the same inventive concept, corresponding to the methods of any of the above embodiments, embodiments of this application also provide a method for circumventing control, such as... Figure 4 As shown, the specific steps include the following: S21. Obtain keyframe or scene semantic information obtained through semantic information methods. The keyframe or scene semantic information is a high-level semantic representation obtained through semantic extraction, semantic memory association and semantic restoration. S22. Based on keyframes or scene semantic representation, perform semantic-level analysis on the current environmental state to identify at least one potential risk target, risk area or abnormal semantic event; S23. Based on potential risk targets, risk areas, or abnormal semantic events, and in conjunction with preset semantic decision rules or decision models, generate corresponding avoidance decision information. S24. Based on the avoidance decision information, output control instructions or prompts to guide the equipment to execute the avoidance suggestions.

[0048] Specifically, in this embodiment, after the receiving end completes the restoration of keyframes or scene semantic information, it analyzes the semantic information to determine whether there are potential risks or abnormal situations. When semantic information related to obstacles, dangerous areas, or abnormal behaviors is detected, corresponding avoidance action prompts or control commands are generated according to preset rules or strategy mapping relationships. Avoidance actions include, but are not limited to, deceleration, detour, stopping, alarm prompts, or path adjustment.

[0049] In multi-device collaborative scenarios, avoidance actions can also serve as high-priority semantic information, fed back to other collaborative devices via semantic communication, enabling collaborative avoidance and safety decision-making among multiple entities. It should be noted that the aforementioned avoidance actions or control commands can be executed after manual confirmation, or used as reference input signals for automatic control systems.

[0050] This method can be applied to at least one of unmanned mobile devices, autonomous driving systems, robotic systems, or intelligent monitoring systems. For example, in unmanned system applications, the transmitting end is deployed on the unmanned mobile device, and the receiving end is deployed on the control platform or other collaborative devices. The system shares key environmental semantic information through semantic communication, enabling collaborative perception and assisted decision-making under low bandwidth or complex channel conditions.

[0051] In this application scenario, the sending end can combine the key frame semantic judgment strategy to adaptively determine whether to trigger the generation and transmission of key frame semantic description based on changes in environmental semantics. This ensures timely transmission of key environmental semantic information while reducing unnecessary communication overhead and improving the operational stability and safety of unmanned systems in dynamic environments.

[0052] In intelligent transportation or safety monitoring applications, the transmitting end is deployed on roadside sensing units, vehicle-mounted terminals, or monitoring equipment, while the receiving end is deployed in traffic control centers or collaborative processing platforms. This embodiment effectively reduces the transmission pressure of multimodal data by transmitting only key semantic information, while ensuring the integrity of the core semantics.

[0053] In this scenario, the system can combine semantic encoding and compression methods to efficiently encode the semantic description of key frames, enabling reliable sharing of multi-source perception semantic information under conditions of limited bandwidth or concurrent communication of multiple devices, thereby improving the overall operating efficiency of the system and the ability to ensure traffic safety.

[0054] This embodiment further illustrates the semantic encoding and compression process of keyframe semantic description. Keyframe semantic description can be represented as structured semantic units, continuous semantic vectors, or a combination thereof.

[0055] During semantic encoding, the semantic encoding module converts semantic information into an encoded form suitable for communication transmission. In some implementations, semantic encoding includes mapping continuous semantic vectors to discrete semantic representations, for example, mapping semantic vectors to corresponding codeword index sequences through codebook-based vector quantization.

[0056] The discretized semantic representation can then be compressed using entropy coding to further reduce the number of communication bits. Under different communication bandwidth or channel conditions, the semantic coding module can adaptively adjust the coding precision, codebook size, or coding length to achieve a balance between semantic fidelity and communication overhead.

[0057] By employing the aforementioned semantic encoding and compression methods, this embodiment can effectively reduce the bandwidth consumption of multimodal semantic communication while ensuring the recoverability of key semantic information.

[0058] Based on the restored multimodal semantic information, this embodiment can directly provide decision support for downstream tasks and trigger corresponding avoidance or control command outputs in specific application scenarios, so that the semantic communication results can effectively serve practical applications such as unmanned systems, intelligent transportation and safety monitoring, and improve the overall intelligence and real-time response capability of the system.

[0059] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0060] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the embodiments of this application also provide a semantic communication system based on multimodal perception.

[0061] like Figure 5 As shown, the semantic communication system based on multimodal awareness includes: The sender and receiver exchange semantic information through a wireless communication channel. The sending end includes a multimodal perception module, a semantic extraction module, and a first semantic memory module (corresponding to...). Figure 5 The transmitting end includes a semantic memory module, a keyframe generation module, and a semantic encoding and transmission module; the receiving end includes a semantic reception and decoding module, and a second semantic memory module (corresponding to...). Figure 5 The receiver includes a semantic memory module, a semantic restoration module, and a decision support module. These modules can be deployed in an integrated or distributed manner. The multimodal sensing module is configured to acquire at least one type of multimodal sensing data from images and videos, and to perform synchronization, cropping, or formatting preprocessing on the acquired data; The semantic extraction module is configured to extract semantics from the preprocessed multimodal data and generate the corresponding current semantic representation; The first and second semantic memory modules are configured to store and manage historical semantic information across time. The keyframe generation module is configured to generate semantic descriptions of keyframes to be transmitted based on semantic changes. The semantic encoding and transmission module is configured to semantically encode the semantic description of the keyframe and transmit it via a wireless channel; The semantic receiving and decoding module is configured to perform semantic decoding on the received semantic information and to perform semantic completion on the received semantic information in combination with the stored historical semantic information. The semantic restoration module is configured to perform conditional semantic restoration on the completed keyframe semantic description based on the set context constraints and using a multimodal model. The decision support module is configured to output decision support information or prompts based on the restoration results.

[0062] It should be noted that in scenarios with poor channel quality, sudden interference, or limited communication bandwidth, the semantic communication system can adaptively reduce the transmission frequency of keyframe semantic descriptions or only send high-priority semantic information to ensure reliable transmission of core semantic content. After the channel quality is restored, the receiver can use the historical semantic information stored in the semantic memory module to complete the missing or delayed semantic content, thereby restoring a complete scene semantic representation.

[0063] For ease of description, the above system is described by dividing it into various modules based on their functions. Of course, in implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.

[0064] The system described in the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0065] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the above embodiments.

[0066] Figure 6This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0067] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0068] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0069] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0070] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0071] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0072] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0073] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0074] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the methods described in any of the above embodiments.

[0075] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0076] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0077] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.

[0078] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0079] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A semantic communication method based on multimodal awareness, characterized in that, include: Acquire at least one type of multimodal information from images and videos, and perform semantic extraction on the multimodal information to generate a corresponding current semantic representation; The current semantic representation is semantically associated with the stored historical semantic information to obtain context-enhanced semantic information used to characterize cross-temporal contextual relationships; Based on the context-enhanced semantic information, determine the semantic difference or semantic change trend between the current semantic representation and the historical semantic information, and generate a keyframe semantic description in response to the semantic difference or semantic change trend meeting a preset threshold. The semantic description of the keyframe is semantically processed, and the semantic description of the processed keyframe is semantically completed based on the historical semantic information. Based on the set context constraints, and using a multimodal model to perform conditional semantic restoration on the completed keyframe semantic description, auxiliary decision-making information or prompt information is output based on the restoration results.

2. The semantic communication method based on multimodal awareness according to claim 1, characterized in that, The method for obtaining the historical semantic information includes: Semantic information stored in the most recent time window based on short-term semantic memory units is used to characterize dynamic changes in the environment; The semantic information of scene structure, target attribute, or environment priors is stored in long-term semantic memory units and remains stable or exhibits a consistent trend of change over a continuous period of time.

3. The semantic communication method based on multimodal awareness according to claim 1, characterized in that: The stored historical semantic information is updated based on the semantic similarity, change magnitude, and time decay factor between the current semantic representation and the historical semantic information. The update method includes at least one of the following forms, specifically: Semantic updates based on time decay, replacement or merging updates based on semantic similarity, and priority updates based on task relevance.

4. The semantic communication method based on multimodal awareness according to claim 1, characterized in that, The keyframe semantic description includes at least one of the following semantic information: Scene structure semantics, target relationship semantics, behavior or motion feature semantics, and anomaly or risk-related semantics.

5. A semantic communication method based on multimodal awareness according to claim 1, characterized in that, The semantic processing of the keyframe semantic description includes: The semantic description of the key frame is semantically encoded and sent to the receiving end through the communication channel. The receiving end performs semantic decoding on the received semantic information and combines it with the stored historical semantic information to complete the semantic information. The semantic encoding includes converting the keyframe semantic description into discrete semantic symbols, low-dimensional semantic vectors, or combinations thereof to adapt to different communication bandwidth conditions.

6. The semantic communication method based on multimodal awareness according to claim 1, characterized in that, The semantic completion includes: Based on the stored historical semantic information, inferences are made to complete missing, ambiguous, or incomplete semantic content.

7. A semantic communication method based on multimodal awareness according to claim 1, characterized in that, The decision support information or prompts include at least one of the following: Risk level assessment information, risk warning information, and avoidance action prompts.

8. An evasion control method, implemented based on a semantic communication method based on multimodal awareness as described in any one of claims 1 to 7, characterized in that, include: Obtain keyframe or scene semantic information obtained through semantic information methods, wherein the keyframe or scene semantic information is a high-level semantic representation obtained through semantic extraction, semantic memory association and semantic restoration; Based on the keyframes or scene semantic representations, perform semantic-level analysis on the current environmental state to identify at least one potential risk target, risk area, or abnormal semantic event. Based on the potential risk targets, risk areas, or abnormal semantic events, and combined with preset semantic decision rules or decision models, corresponding avoidance decision information is generated. Based on the avoidance decision information, control instructions or prompts are output to guide the device to execute the avoidance suggestions.

9. A semantic communication system based on multimodal awareness, characterized in that, The system includes: The sending end and the receiving end exchange semantic information through a wireless communication channel: The transmitting end includes a multimodal perception module, a semantic extraction module, a first semantic memory module, a keyframe generation module, and a semantic encoding and transmission module; the receiving end includes a semantic receiving and decoding module, a second semantic memory module, a semantic restoration module, and an auxiliary decision-making module. The multimodal sensing module is configured to acquire at least one type of multimodal sensing data from images and videos, and to perform synchronization, cropping, or formatting preprocessing on the acquired data. The semantic extraction module is configured to extract semantics from the preprocessed multimodal data and generate the corresponding current semantic representation; The first semantic memory module and the second semantic memory module are configured to store and manage historical semantic information across time. The keyframe generation module is configured to generate a semantic description of the keyframe to be transmitted based on semantic changes. The semantic encoding and transmission module is configured to perform semantic encoding on the semantic description of the key frame and transmit it through a wireless channel; The semantic receiving and decoding module is configured to perform semantic decoding on the received semantic information and to perform semantic completion on the received semantic information in combination with the stored historical semantic information. The semantic restoration module is configured to perform conditional semantic restoration on the completed keyframe semantic description based on set context constraints and using a multimodal model. The auxiliary decision-making module is configured to output auxiliary decision-making information or prompt information based on the restoration result.

10. A non-transitory computer-readable storage medium, characterized in that, in, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute a semantic communication method based on multimodal awareness as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Generative multi-mode mutual benefit enhancement video semantic communication method

    CN116939320A

  • Multi-modal semantic communication method, system and equipment based on large model and medium

    CN118350416A

  • Semantic scene completion method, electronic equipment and storage medium

    CN118505990A

  • Social robot semantic understanding method and system based on knowledge graph, electronic equipment and storage medium

    CN121233842A

  • Dynamic semantic evolution tracking and concept drift adaptive updating method and system

    CN121257542A

Cited By

  • Semantic execution of motor tasks, apparatus and storage medium

    CN122343464A