Semantic information extraction and characterization method and system for coherent optical communication

CN122554012BActive Publication Date: 2026-09-15CHINESE PEOPLES LIBERATION ARMY INFORMATION SUPPORT CORPS ENGINEERING UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611039323.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-15
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

然而,现有相干光通信系统在传输海量遥感图像数据时面临诸多挑战

Benefits of technology

[0014] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the semantic information extraction and characterization method for coherent optical communication as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554012B_ABST
    Figure CN122554012B_ABST
Patent Text Reader

Abstract

The application provides a kind of semantic information extraction and characterization method and system for coherent optical communication, belong to communication technical field, the method includes: obtaining original image and task prompt word, obtains structured task demand by parsing;Based on task demand, target detection and semantic segmentation are carried out on image, and entity pixel-level segmentation mask is generated;Extract entity attribute information, analyze the relationship information between entities;Construct structured semantic scene graph;The node category, attribute and relationship information in semantic scene graph are respectively mapped to different physical degrees of freedom of coherent light carrier, and a coherent light signal modulated by semantics is generated.The application realizes the deep integration of semantic information and coherent optical physical layer, significantly improves the transmission efficiency and the explainability of semantic characterization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a method and system for semantic information extraction and representation in coherent optical communication. Background Technology

[0002] With the development of 6G and integrated space-air-ground communication, coherent optical communication has become a core enabling technology for satellite remote sensing image transmission due to its advantages such as high bandwidth and high security. However, existing coherent optical communication systems face many challenges when transmitting massive amounts of remote sensing image data. On the one hand, traditional coherent optical communication systems mainly adopt the syntax communication paradigm, aiming at the precise transmission of symbols without considering the semantic connotation of the image content. This results in serious waste of bandwidth resources and limited transmission efficiency when transmitting massive amounts of remote sensing data within a limited window period. On the other hand, existing semantic communication research mostly uses neural networks to extract latent feature vectors, which suffers from insufficient symbolization and structuring of semantic representation, poor interpretability, difficulty in accurately controlling the granularity of semantic information, and a lack of methods to explicitly represent image semantics as structured and symbolic information, making it difficult for the receiving end to utilize semantic priors for efficient processing. Furthermore, semantic information is characterized by high dimensionality, sparseness, and multiple types. However, the modulation formats of existing coherent optical communication are mainly designed for binary bit streams and lack a representation method that can efficiently map semantic information to the degrees of freedom such as optical carrier amplitude and phase. Moreover, existing semantic extraction methods are mostly general designs and cannot dynamically adjust the extraction strategy according to specific communication tasks, resulting in a large amount of redundancy in the transmitted information that is irrelevant to the task.

[0003] Therefore, how to achieve deep integration of semantic information and coherent optical modulation format to improve transmission efficiency is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This invention provides a method and system for semantic information extraction and representation in coherent optical communication, in order to overcome the deficiencies in the prior art.

[0005] In a first aspect, the present invention provides a method for semantic information extraction and representation in coherent optical communication, comprising: Obtain the original image and task prompts; Semantic analysis is performed on the task prompts to obtain structured task requirements; Based on the task requirements, target detection is performed on the original image to identify entities that match the task requirements, and semantic segmentation is performed on the identified entities to generate pixel-level segmentation masks for each entity. The attribute information of each entity is extracted based on the pixel-level segmentation mask; Based on the pixel-level segmentation mask and the spatial coordinates of each entity, analyze the spatial and / or logical relationships between each entity to obtain relationship information; Based on the identified entities, the attribute information, and the relationship information, a structured semantic scene graph is constructed. The semantic scene graph includes a node set and an edge set. The node set contains entity category information and attribute information, and the edge set contains relationship information. The entity category information, attribute information, and relationship information in the semantic scene graph are mapped to different physical degrees of freedom of the coherent optical carrier to generate semantically modulated coherent optical signals.

[0006] According to the semantic information extraction and representation method for coherent optical communication provided by the present invention, the step of semantically parsing the task prompt words to obtain structured task requirements includes: semantically encoding the task prompt words using a pre-trained language model to obtain encoding features; calculating the cosine similarity between the encoding features and each task template in a preset task knowledge base; and determining the task template with the highest similarity to the encoding features as the structured task requirements, wherein the structured task requirements include a set of target entity types, a set of attention attributes, and a set of attention relationship information.

[0007] According to the semantic information extraction and representation method for coherent optical communication provided by the present invention, based on the task requirements, target detection is performed on the original image to identify entities that match the task requirements, and semantic segmentation is performed on the identified entities to generate pixel-level segmentation masks for each entity. The method includes: using a preset target detection model to detect entities from the original image according to the target entity type set, and outputting text labels and bounding box coordinates for each entity; using a preset target segmentation model to perform semantic segmentation on each entity based on the bounding box coordinates to obtain the pixel-level segmentation mask.

[0008] According to the semantic information extraction and representation method for coherent optical communication provided by the present invention, the step of extracting attribute information of each entity based on the pixel-level segmentation mask includes: inputting the pixel-level segmentation mask and the set of interests into a multimodal visual language model; and outputting the attribute information of each entity from the multimodal visual language model.

[0009] According to the semantic information extraction and representation method for coherent optical communication provided by the present invention, the step of analyzing the spatial and / or logical relationships between the entities based on the pixel-level segmentation mask and the spatial coordinates of each entity to obtain relationship information includes: constructing a graph neural network, fusing the text labels and attribute information of each entity as the node features of the graph neural network; iteratively updating the node features through the graph neural network to analyze the spatial and / or logical relationships between entities, and outputting relationship information that matches the set of interest relationship information.

[0010] According to the semantic information extraction and representation method for coherent optical communication provided by the present invention, the step of mapping entity category information, attribute information, and relation information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier includes: mapping entity category information to amplitude modulation of the coherent optical carrier; mapping color attribute in attribute information to polarization state modulation of the coherent optical carrier; mapping material attribute in attribute information to phase modulation of the coherent optical carrier; and mapping size attribute or shape attribute in attribute information to orbital angular momentum modulation of the coherent optical carrier; and mapping relation information to wavelength modulation of the coherent optical carrier.

[0011] According to the semantic information extraction and characterization method for coherent optical communication provided by the present invention, the generation of semantically modulated coherent optical signal includes: using an IQ modulator, a polarization controller, an orbital angular momentum modulator and a tunable laser to synchronously and jointly load the amplitude modulation, polarization state modulation, phase modulation, orbital angular momentum modulation and wavelength modulation to obtain the semantically modulated coherent optical signal.

[0012] Secondly, the present invention also provides a semantic information extraction and representation system for coherent optical communication, comprising: The data acquisition module is used to acquire the original images and task prompts; The task parsing module is used to perform semantic parsing on the task prompt words to obtain structured task requirements; The object detection and semantic segmentation module is used to perform object detection on the original image based on the task requirements, identify entities that match the task requirements, perform semantic segmentation on the identified entities, and generate pixel-level segmentation masks for each entity. An attribute-aware module is used to extract attribute information of each entity based on the pixel-level segmentation mask; The relation reasoning module is used to analyze the spatial and / or logical relationships between the entities based on the pixel-level segmentation mask and the spatial coordinates of each entity, and to obtain relation information. The scene graph generation module is used to construct a structured semantic scene graph based on the identified entities, the attribute information, and the relationship information. The semantic scene graph includes a node set and an edge set. The node set contains entity category information and attribute information, and the edge set contains the relationship information. The coherent optical modulation module is used to map the entity category information, attribute information and relation information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier, respectively, to generate a semantically modulated coherent optical signal.

[0013] Thirdly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the semantic information extraction and characterization method for coherent optical communication as described above.

[0014] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the semantic information extraction and characterization method for coherent optical communication as described above.

[0015] This invention provides a method and system for semantic information extraction and representation in coherent optical communication. Through a task parsing module, natural language task prompts are transformed into specific task requirements, guiding the subsequent semantic extraction process. This achieves task-oriented semantic information extraction, effectively eliminating redundant information irrelevant to the task and significantly improving transmission efficiency. By constructing a structured semantic scene graph, implicit image features are transformed into explicit, symbolic semantic representations, enhancing the interpretability and granular controllability of semantic information. Furthermore, by mapping the node categories, node attributes, and relationship information of the semantic scene graph to multiple physical degrees of freedom of the coherent optical carrier, such as amplitude, polarization, phase, orbital angular momentum, and wavelength, a cross-layer mapping model of "semantic-optical modulation" is established. This achieves deep integration of semantic information and coherent optical modulation format, optimizes modulation efficiency, and increases the information carrying capacity of a single symbol, providing a complete solution for task-oriented, high-fidelity, and high-efficiency coherent optical semantic communication. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the semantic information extraction and characterization method for coherent optical communication provided by the present invention. Figure 2 This is a schematic diagram of the semantic information extraction and representation system for coherent optical communication provided by the present invention; Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0019] It should be noted that in the description of the embodiments of the present invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art can understand the specific meaning of the above terms in the present invention according to the specific circumstances. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0020] The following is combined with Figures 1-3 This invention describes a method and system for semantic information extraction and characterization for coherent optical communication, as provided in embodiments of the present invention.

[0021] Figure 1 This is a flowchart illustrating the semantic information extraction and representation method for coherent optical communication provided by the present invention, as shown below. Figure 1 As shown, including but not limited to the following steps: Step 101: Obtain the original image and task prompts.

[0022] This step serves as the starting point for the method's input, acquiring two types of data: first, the original image, which can be a satellite remote sensing image, a drone aerial image, or other visual data to be transmitted; second, task prompts, which are specific communication task requirements described in natural language, such as "identify vehicles in the image and their color and positional relationships." This step provides the data foundation for subsequent task-driven semantic extraction.

[0023] Step 102: Perform semantic parsing on the task prompt words to obtain structured task requirements.

[0024] Since the original task prompts are typically unstructured natural language text, computers struggle to understand them directly and use them to guide image processing. Therefore, this step utilizes semantic parsing technology to transform the natural language into machine-understandable structured data. This structured task requirement clarifies "what to look at (target entity)," "which features to look at (attributes of interest)," and "what relationships to focus on (relationships of interest)," providing precise guidance for the subsequent visual extraction process. This avoids the indiscriminate extraction of all information by general visual models, effectively reducing computational overhead and improving the relevance of information.

[0025] Optionally, step 102 includes: (1) Use a pre-trained language model to semantically encode the task prompt words to obtain encoding features.

[0026] The purpose of this step is to convert task cues in natural language form into numerical feature vectors that can be computed by the machine. Specifically, a pre-trained language model (such as BERT) is used as the encoder to semantically encode the input task cues T. The encoding process can be represented as follows: Encoded feature vector = BERT(T); Here, BERT() represents the forward propagation function of the pre-trained language model, whose output is a dense vector of fixed dimensions, which contains deep semantic information of the task prompts. By pre-training on a large-scale corpus, the BERT model has the ability to capture the semantic relationships between word contexts and can accurately understand the task intent and constraints implied in the task prompts.

[0027] (2) Calculate the cosine similarity between the encoded features and each task template in the preset task knowledge base.

[0028] This step compares the encoded task prompt word features with known task templates in a pre-defined task knowledge base using similarity matching. The pre-defined task knowledge base, denoted as K, stores multiple predefined task templates. Each task template corresponds to a standardized task requirement structure, and its semantic encoded feature vector is pre-calculated and stored in the same manner as in the previous steps.

[0029] For any task template T' in the knowledge base, the cosine similarity calculation formula is:

[0030] in, The cosine similarity function calculates the cosine of the angle between two vectors, with a value ranging from -1 to 1. A larger value indicates that the two vectors are closer in direction in the semantic space, meaning higher semantic similarity. By traversing all task templates in the knowledge base and calculating their similarity, the semantic matching degree between task prompts and each task template can be obtained.

[0031] (3) The task template with the highest similarity to the encoded features is determined as the structured task requirement, which includes a set of target entity types, a set of attention attributes, and a set of attention relationship information.

[0032] This step makes a decision based on the similarity calculation results, selecting the task template that is semantically closest to the task prompt words as the final task requirement. This selection process can be represented as:

[0033] Where argmax represents selecting the task template that maximizes the similarity function value. The finalized structured task requirements.

[0034] The determined structured task requirements It consists of three core subsets: Target entity type set (denoted as ): Defines the categories of entities that need to be detected and identified in the image that the task is interested in, such as "vehicles", "buildings", "ships", etc. The set of attributes to watch (denoted as) ): Defines the types of fine-grained attributes of the entities to be extracted, such as "color", "material", "size", "shape", etc.; Focus on the set of relational information (denoted as) ): Defines the relationship information between entities that need to be analyzed, such as spatial and / or logical relationships like "above", "near", "contains".

[0035] Step 103: Based on the task requirements, perform target detection on the original image, identify entities that match the task requirements, perform semantic segmentation on the identified entities, and generate pixel-level segmentation masks for each entity.

[0036] This step achieves visual localization and fine segmentation of entities from images. Specifically, the object detection process does not exhaustively identify all objects in the image, but is constrained by the task requirements obtained in step 102, detecting only entity categories relevant to the task. For example, when the task requirement is "identify vehicles," background elements such as trees and roads in the image will be ignored. After detecting the location of the entity, semantic segmentation technology is further used to generate a pixel-level segmentation mask. This mask can accurately separate entities from the background, providing clean pixel data for subsequent attribute analysis, solving the problem of traditional detection boxes containing a lot of background noise, and ensuring the purity of semantic information.

[0037] Optionally, step 103 includes: (1) Using a preset target detection model, detect entities from the original image according to the target entity type set, and output the text label and bounding box coordinates of each entity.

[0038] This step uses a pre-defined object detection model to perform task-oriented entity detection on the original image. The pre-defined object detection model is preferably the YOLO-World object detection network, which is an open-vocabulary object detection model capable of detecting corresponding entities in an image based on any text description, without requiring retraining for specific categories.

[0039] The inputs to the detection process include: the original image I (with dimensions H×W×3) and the set of target entity types obtained in step 102. The model output is a set of detection results, where the first... i The detected entities include: Text Label : Represents the category name (entity / category label) of the entity, which is related to It matches a certain entity type in the data; Bounding box coordinates :in , The x and y coordinates of the center point of the bounding box in the image. The width of the bounding box. This represents the height of the bounding box.

[0040] Training objective function of the detection network A weighted combination of classification loss and regression loss:

[0041] in, For classification loss functions, cross-entropy loss is typically used to constrain the prediction accuracy of entity categories; The bounding box regression loss function is typically the Smooth L1 loss, which is used to constrain the prediction accuracy of the bounding box coordinates. To balance the hyperparameters of the two loss weights.

[0042] This step only detects the set of target entity types that match the task requirements. Matching entities, background objects unrelated to the task, or entities of no interest are automatically ignored, thereby achieving task-oriented semantic information filtering and effectively reducing the amount of redundant data in subsequent processing.

[0043] (2) Using a preset target segmentation model, semantic segmentation is performed on each entity based on the bounding box coordinates to obtain the pixel-level segmentation mask.

[0044] Based on object detection, the next step performs fine pixel-level semantic segmentation on each detected entity. The preset object segmentation model is preferably the SAM2 model, which is a fundamental vision model with powerful zero-shot segmentation capabilities and can generate high-quality pixel-level segmentation masks based on given spatial cues (such as bounding boxes).

[0045] The inputs to the segmentation process include: the original image I and the bounding box coordinates generated from the previous sub-step. For the first i (As subscript) detected entities, in bounding boxes As a spatial cue, the SAM2 model outputs the pixel-level segmentation mask corresponding to the entity. :

[0046] in, It is a two-dimensional matrix with the same spatial resolution (H×W) as the original image I. Each element in the matrix takes a probability value in the interval [0,1], indicating that the pixel belongs to the [0,1]-th ... i The confidence level of an entity can be calculated; it can also be converted into a binary mask through thresholding, with a value of 0 or 1, representing whether the pixel does not belong to the entity or belongs to the entity.

[0047] The pixel-level segmentation mask of this invention provides the precise spatial boundary of an entity in an image. Compared to a bounding box that only contains a rectangular area, the mask can completely depict the irregular contour of the entity and effectively distinguish the pixels inside the entity from the background pixels. It provides a precise source of pixel data for the attribute extraction in the subsequent step 104, avoiding interference from background pixels on attribute recognition. It provides accurate pixel area and spatial coordinate information occupied by the entity for the spatial relationship analysis in step 105, making the judgment of spatial relationships such as "located above" and "nearby" more accurate and reliable.

[0048] Step 104: Extract the attribute information of each entity based on the pixel-level segmentation mask.

[0049] After obtaining the pixel-level mask of the entity, this step further mines the entity's fine-grained features. Attribute information includes, but is not limited to, visual features such as color, material, size, and shape. By mapping the masked area to the original image, only the pixel region where the entity is located is analyzed, thereby extracting attribute information describing the entity's inherent characteristics.

[0050] Optionally, the attribute information of each entity is extracted based on the pixel-level segmentation mask, including: (1) Input the pixel-level segmentation mask and the set of attention attributes into the multimodal visual language model.

[0051] Original Image I (dimensions: height H × width W × RGB three channels): Provides the original visual information of the complete scene in which the entity is located, and is the basic data source for attribute perception. The original image contains fine-grained visual cues such as the entity's texture, color, and edges, which are necessary for identifying attributes such as color, material, and shape.

[0052] Pixel-level segmentation mask (Dimension H×W): Generated from the previous object detection and semantic segmentation steps, accurately labeling the first... i The pixel area occupied by each entity in the original image. Masking. By performing pixel-by-pixel correspondence with the original image I, a local image region containing only the entity can be cropped or focused, thereby limiting the attention of attribute analysis to the entity itself, eliminating interference from background areas and other entities, and significantly improving the accuracy of attribute recognition.

[0053] Set of attributes to watch This set, generated during the task parsing step, defines the range of attribute types of interest to the current communication task, such as one or more of color, material, size, and shape. This set serves as a constraint input to the model, determining which attribute recognition branches the model needs to activate.

[0054] The multimodal visual language model (preferably the CoCa model or the LLaVa model) receives input: a pixel-level segmentation mask. and the set of attributes to focus on This type of model is trained on a large scale with image-text joint training and has cross-modal semantic alignment capabilities, enabling it to associate and map pixel features of visual regions with attribute semantics described in natural language.

[0055] (2) The attribute information of each entity is output by the multimodal visual language model.

[0056] The attribute information includes at least one of color attribute, material attribute, size attribute and shape attribute.

[0057] Multimodal visual language models perform semantic understanding on input entity pixel data and output the first... i Attribute information of an entity Its structure is a multidimensional vector:

[0058] The meanings of each component are as follows: (Color attribute): Describes the color characteristics of an entity. The value is a quantized value in the color category label or color space, such as red, gray, camouflage, etc. Its numerical range is usually [0,1], representing the confidence probability of the color attribute.

[0059] (Material Attributes): Describes the material type of the entity's surface, such as metal, plastic, wood, glass, etc., and is also output in the form of probability values ​​or category labels.

[0060] (Size attribute): Describes the scale information of an entity, such as large, medium, small, or expressed as an estimate of its actual physical size.

[0061] (Shape attribute): Describes the geometric outline features of an entity, such as rectangle, circle, ellipse, irregular shape, etc.

[0062] The model is based on the set of attention attributes in the input. This determines which attribute components are actually output. If a certain attribute type (such as color) is included in... If the model outputs the corresponding attribute value, then the model outputs the attribute value; if a certain attribute type is not listed... If this is achieved, the model can skip the calculation of that attribute, thus avoiding the extraction of redundant attribute information irrelevant to the task. This mechanism enables task-adaptive control of attribute extraction granularity, ensuring that the transmitted attribute information is highly relevant to the current communication task.

[0063] Optionally, the attribute perception of the multimodal visual language model adopts a multimodal contrastive learning framework, with the loss function being: in For real attribute tags, q Indicates the attribute category.

[0064] This invention clarifies the input (pixel-level segmentation mask, set of attributes of interest) and output (attribute information including at least one of color, material, size, and shape) of the multimodal visual language model in the attribute extraction step. This gives the attribute perception process clear input-output boundaries and task orientation, enabling the accurate extraction of fine-grained entity attributes related to the communication task from images. It avoids background interference and redundant attribute calculations, providing reliable and compact structured data support for the subsequent construction of semantic scene graphs and efficient multi-degree-of-freedom coherent optical modulation of semantic information.

[0065] Step 105: Based on the pixel-level segmentation mask and the spatial coordinates of each entity, analyze the spatial and / or logical relationships between the entities to obtain relationship information.

[0066] Spatial relationships describe the relative positions of entities in physical space, such as "above," "containing," etc.; logical relationships describe the semantic associations between entities, such as "supporting," "belonging to," etc. By combining masked regions and spatial coordinates, the system can construct the topological structure between entities, thereby understanding the overall semantic scene of the image, rather than viewing individual objects in isolation. This extraction of relational information provides crucial edge features for the subsequent construction of a structured scene graph.

[0067] Optionally, based on the pixel-level segmentation mask and the spatial coordinates of each entity, the spatial and / or logical relationships between the entities are analyzed to obtain relationship information, including: (1) Construct a graph neural network and fuse the text labels and attribute information of each entity as the node features of the graph neural network.

[0068] First, a graph neural network is constructed as the core computational framework for relational reasoning. Graph neural networks are naturally suitable for processing relational structures between entities, with nodes corresponding to the entities detected in the image and edges corresponding to the spatial and / or logical relationships to be analyzed.

[0069] For the i The node features of a detected entity are constructed as follows: the text label of the entity... (Representing entity categories, such as "vehicles" or "buildings") and attribute information (Including attribute components such as color, material, size, and shape) are fused. The fusion operation is implemented through a multilayer perceptron, represented as:

[0070] in, This indicates that the vector representation of the text label is concatenated with the attribute vector; The function represents a multilayer perceptron that maps the concatenated vectors to a fixed-dimensional latent space, generating node feature vectors. .

[0071] Through this fusion operation, each node of the graph neural network simultaneously contains the entity's category semantics and fine-grained attribute semantics, providing a rich semantic foundation for accurately determining the relationships between entities.

[0072] (2) The node features are iteratively updated through the graph neural network to analyze the spatial and / or logical relationships between entities and output relationship information that matches the set of interest relationship information.

[0073] This sub-step uses the message passing mechanism of a graph neural network to iteratively update the features of nodes and edges, ultimately inferring the specific relationship information between entities. The top-level function of this process is represented as:

[0074] in, This represents the functions of the relational reasoning engine. and The first j The entity and the first k Pixel-level segmentation mask for each entity Indicates the first j The specific types of relationships between the k-th entity and the k-th entity, and belong .

[0075] Within the relational reasoning engine, the graph neural network employs a multi-layered structure for iterative computation, reaching the [missing information - likely a specific step or stage]. l +1 layer, the update formula for edge features is:

[0076] in, and The first l Layer j The node and the first k Feature vectors of each node; For the first l Layer connection nodes j and nodes k The edge feature vector; [;;] indicates concatenating the three vectors; MLP() is a multilayer perceptron, which outputs the updated edge features. .

[0077] This update process enables edge features to simultaneously perceive the semantic information of both endpoints and the historical state of the edge itself. As the number of layers increases, graph neural networks can capture more complex and abstract relational patterns between entities. In addition to node features, the input information for relational reasoning also includes pixel-level segmentation masks for each entity. , The spatial information, including the nodes' spatial coordinates, is encoded as relative positional features between nodes, enabling the model to accurately determine spatial relationships such as "above" and "contains," as well as logical relationships such as "belongs to," "attacks," and "supports."

[0078] Finally, after multiple iterations, the graph neural network generates a relation classification result for each edge in the output layer, outputting a set of information related to the relationships of interest. Matching relationship information .

[0079] Step 106: Based on the identified entities, the attribute information, and the relationship information, construct a structured semantic scene graph. The semantic scene graph includes a node set and an edge set. The node set contains entity labels and attribute vectors, and the edge set contains relationship information.

[0080] Specifically, the semantic scene graph is defined as:

[0081] in, The symbol represents the semantic scene graph generated based on task requirements D (the subscript D emphasizes its task orientation); V is the node set, representing the set of all entity nodes in the scene graph; E is the edge set, representing the set of all relation edges in the scene graph.

[0082] Each node in node set V Corresponding to the detection and segmentation obtained in step 103, the first i An entity. A node. The complete representation can be written as:

[0083] in, For discrete category labels, It is a set of continuous or discrete attribute values. This structure allows each entity node to simultaneously possess both categorical semantics of "what it is" and fine-grained attribute semantics of "what features it has".

[0084] Each edge in edge set E Connecting the j-th entity and the k-th entity that have a mutual relationship is represented as:

[0085] The generated scene graph The graph neural network encoder further compresses the vector into a fixed-dimensional feature vector: g = GNN( V, E ) , where g is a d-dimensional real vector.

[0086] This compression method is beneficial for the numerical processing of subsequent modulation steps, but the structured node set and edge set of the scene graph itself are sufficient to support the multi-dimensional mapping of semantic information.

[0087] Step 107: Map the entity category information, attribute information and relationship information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier to generate semantically modulated coherent optical signals.

[0088] This step enables cross-layer conversion of semantic information to physical layer signals. Coherent optical carriers possess multiple physical degrees of freedom, including amplitude, phase, polarization state, orbital angular momentum (OAM), and wavelength. This invention creatively establishes mapping rules between semantic dimensions and physical degrees of freedom: for example, mapping node categories to amplitude, attribute features to polarization or phase, and relational information to wavelength. This mapping method fully utilizes the high-dimensional modulation capability of optical carriers, enabling a single optical signal symbol to carry multi-dimensional semantic information, significantly improving transmission efficiency. The resulting semantically modulated coherent optical signal not only carries bit information but also directly carries structured semantic connotations, achieving a deep integration of semantic communication and coherent optical communication.

[0089] Optionally, mapping the entity category information, attribute information, and relation information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier includes: (1) Map the entity category information to the amplitude modulation of the coherent optical carrier; Entity category information (i.e., text labels) Encoding is performed by adjusting the amplitude (intensity) of the coherent optical carrier. For entity category information... The amplitude modulation function is defined as:

[0090] in, This indicates the corresponding entity category. The complex amplitude of the coherent state; This corresponds to the average number of photons; The initial reference phase is indicated by the superscript. i It is the imaginary unit.

[0091] Determined through a preset category code lookup table:

[0092] The generated coherent states are:

[0093] That is, by adjusting the intensity (number of photons) of the optical carrier. This is used to distinguish different entity categories. For example, "vehicle" corresponds to photon number n1, and "building" corresponds to photon number n2. The receiver can recover the entity category information by detecting the intensity of the light signal.

[0094] (2) Map the color attribute in the attribute information to the polarization state modulation of the coherent optical carrier; The polarization state can be represented by the Stokes parameter: And satisfy ; in, These are the three components of the Stokes vector, each corresponding to a different polarization characteristic of the light wave.

[0095] Color feature vector (Such as RGB values) polarization states are obtained through nonlinear mapping:

[0096] in, It is a learnable mapping matrix.

[0097] Through this mapping, different colors are encoded as different polarization states of the optical carrier, and the receiver can recover the color attributes by measuring the polarization.

[0098] (3) Map the material properties in the attribute information to the phase modulation of the coherent optical carrier; Set material type The phase encoding is:

[0099] in, For the number of material types, Returns the material's index in the predefined dictionary. For the initial phase, Corresponding to material The phase modulation value.

[0100] This mapping employs phase keying to map different material types to different phase angles of optical carriers. Because phase modulation exhibits high sensitivity in coherent optical communication, this encoding method enables efficient transmission of material properties.

[0101] (4) Map the size or shape attributes (or other attributes) in the attribute information to the orbital angular momentum modulation of the coherent optical carrier;

[0102] in, A quantified value representing a size attribute, shape attribute (or other attribute); The lookup function is a preset orbital angular momentum encoding; is the topological charge value, which takes integer values ​​and represents the mode order of the orbital angular momentum.

[0103] The light field carrying orbital angular momentum can be represented as a superposition of multiple orbital angular momentum modes:

[0104] Among them, |l The topological charge is l The orbital angular momentum eigenstates, represents the complex amplitude coefficient of the corresponding mode.

[0105] Ultimately, the photonic quantum state of each entity is the direct product of the aforementioned degrees of freedom.

[0106] Indicates the first i The total photon states corresponding to each entity; It is an amplitude-modulated state, carrying entity category information; It is in a polarized state and carries color attribute information; It is a phase state, carrying material property information; It represents the orbital angular momentum state, carrying size or shape attribute information; This represents the tensor product operation, and represents the joint composition of states of various degrees of freedom.

[0107] (5) Map the relational information to the wavelength modulation of the coherent optical carrier.

[0108] For relational information The wavelength modulation function is defined as:

[0109] in: As the reference wavelength; The preset wavelength interval; Return relationship information The index in the preset relation dictionary takes values ​​in the range {0, 1, ..., R-1}, where R is the total number of relation information; To correspond to relational information The modulated carrier wavelength.

[0110] Different relational information corresponds to different carrier wavelengths, achieving wavelength domain multiplexing. This method can utilize wavelength division multiplexing (WDM) technology to transmit multiple sets of relational information in parallel within the same optical fiber, or to transmit different relational information at different time windows, thereby improving the overall throughput of the communication system.

[0111] Optionally, generating the semantically modulated coherent optical signal includes: synchronously and jointly loading the amplitude modulation, polarization state modulation, phase modulation, orbital angular momentum modulation, and wavelength modulation using an IQ modulator, a polarization controller, an orbital angular momentum modulator, and a tunable laser to obtain the semantically modulated coherent optical signal.

[0112] Specifically, the tunable laser is responsible for generating the optical carrier and adjusting the center wavelength of the output according to the relational information. Different relational types correspond to different wavelength offsets, and the corresponding offsets are superimposed on the reference wavelength to achieve information encoding in the wavelength dimension.

[0113] The optical carrier then enters the IQ modulator. In this invention, the IQ modulator simultaneously performs two modulation tasks: first, it converts entity category information into the intensity amplitude of the optical carrier, with different entity categories corresponding to different intensity levels; second, it loads the phase angle corresponding to the material properties onto the optical carrier, with different material types corresponding to different phase offsets. Thus, the IQ modulator outputs an optical signal that simultaneously carries both entity category and material properties.

[0114] The optical signal output from the IQ modulator enters the polarization controller. Based on the Stokes parameters corresponding to the color attribute, the polarization controller adjusts the polarization rotation angle and phase delay of the optical signal, adjusting the polarization state of the optical signal to match the target color. Different colors correspond to different polarization state configurations, and the receiving end can distinguish the color attribute through polarization measurement.

[0115] Finally, the optical signal passes through an orbital angular momentum modulator. This modulator, based on a spatial light modulator or a spiral phase plate, introduces a spiral phase distribution on the wavefront of the optical field, loading the orbital angular momentum topological charge corresponding to the size or shape attributes onto the optical field, generating a vortex beam carrying a specific orbital angular momentum mode. Different sizes or shapes correspond to different topological charge values, and different orbital angular momentum modes are orthogonal to each other, allowing for multiplexing and transmission within the same optical path.

[0116] The four devices described above operate in parallel under unified control timing, synchronously and jointly loading the five independent physical degrees of freedom of the optical carrier: amplitude, phase, polarization, orbital angular momentum, and wavelength. The modulation of each degree of freedom is orthogonal and does not interfere with each other. The final output semantically modulated coherent optical signal fully carries entity category information, multi-type attribute information, and inter-entity relationship information from the semantic scene graph, and can be transmitted to the receiving end via free space or fiber optic channels.

[0117] Figure 2 This is a schematic diagram of the semantic information extraction and representation system for coherent optical communication provided by the present invention, as shown below. Figure 2 As shown, the system includes: Data acquisition module 210 is used to acquire the original image and task prompt words; Task parsing module 220 is used to perform semantic parsing on the task prompt words to obtain structured task requirements; The object detection and semantic segmentation module 230 is used to perform object detection on the original image based on the task requirements, identify entities that match the task requirements, perform semantic segmentation on the identified entities, and generate pixel-level segmentation masks for each entity. The attribute-aware module 240 is used to extract attribute information of each entity based on the pixel-level segmentation mask; The relation reasoning module 250 is used to analyze the spatial and / or logical relationships between the entities based on the pixel-level segmentation mask and the spatial coordinates of each entity, and to obtain relation information. The scene graph generation module 260 is used to construct a structured semantic scene graph based on the identified entities, the attribute information and the relationship information. The semantic scene graph includes a node set and an edge set. The node set contains entity category information and attribute information, and the edge set contains the relationship information. The coherent optical modulation module 270 is used to map the entity category information, attribute information and relation information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier, respectively, to generate a semantically modulated coherent optical signal.

[0118] It should be noted that the semantic information extraction and characterization system for coherent optical communication provided in this embodiment of the invention can execute the semantic information extraction and characterization method for coherent optical communication described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0119] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 3 As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communications bus 340. The processor 310, communications interface 320, and memory 330 communicate with each other via the communications bus 340. The processor 310 can call logical instructions from the memory 330 to execute semantic information extraction and representation methods for coherent optical communication.

[0120] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0121] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the semantic information extraction and characterization method for coherent optical communication provided in the above embodiments.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for semantic information extraction and characterization for coherent optical communications, characterized in that, include: Obtain the original image and task prompts; Semantic analysis is performed on the task prompts to obtain structured task requirements; Based on the task requirements, target detection is performed on the original image to identify entities that match the task requirements, and semantic segmentation is performed on the identified entities to generate pixel-level segmentation masks for each entity. The attribute information of each entity is extracted based on the pixel-level segmentation mask; Based on the pixel-level segmentation mask and the spatial coordinates of each entity, analyze the spatial and / or logical relationships between each entity to obtain relationship information; Based on the identified entities, the attribute information, and the relationship information, a structured semantic scene graph is constructed. The semantic scene graph includes a node set and an edge set. The node set contains entity category information and attribute information, and the edge set contains relationship information. The entity category information, attribute information, and relationship information in the semantic scene graph are mapped to different physical degrees of freedom of the coherent optical carrier to generate semantically modulated coherent optical signals.

2. The method of claim 1, wherein, Semantic parsing of the task prompts yields structured task requirements, including: The task prompt words are semantically encoded using a pre-trained language model to obtain encoded features; Calculate the cosine similarity between the encoded features and each task template in the preset task knowledge base; The task template with the highest similarity to the encoded features is determined as the structured task requirement, which includes a set of target entity types, a set of attention attributes, and a set of attention relationship information.

3. The method of claim 2, wherein, Based on the task requirements, target detection is performed on the original image to identify entities matching the task requirements, and semantic segmentation is performed on the identified entities to generate pixel-level segmentation masks for each entity, including: Using a preset target detection model, entities are detected from the original image based on the target entity type set, and the text labels and bounding box coordinates of each entity are output. Using a preset target segmentation model, semantic segmentation is performed on each entity based on the bounding box coordinates to obtain the pixel-level segmentation mask. 4.The method of claim 2, wherein, The step of extracting attribute information of each entity based on the pixel-level segmentation mask includes: The pixel-level segmentation mask and the set of attention attributes are input into the multimodal visual language model; The multimodal visual language model outputs the attribute information of each entity.

5. The semantic information extraction and representation method for coherent optical communication according to claim 3, characterized in that, The step of analyzing the spatial and / or logical relationships between the entities based on the pixel-level segmentation mask and the spatial coordinates of each entity to obtain relationship information includes: A graph neural network is constructed, and the text labels and attribute information of each entity are fused together as the node features of the graph neural network. The graph neural network iteratively updates node features to analyze spatial and / or logical relationships between entities, and outputs relationship information that matches the set of interest relationship information.

6. The semantic information extraction and representation method for coherent optical communication according to claim 1, characterized in that, The step of mapping entity category information, attribute information, and relationship information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier includes: The entity category information is mapped to the amplitude modulation of the coherent optical carrier; The color attribute in the attribute information is mapped to the polarization state modulation of the coherent optical carrier; Mapping the material properties in the attribute information to the phase modulation of the coherent optical carrier; and, Map the size or shape attributes in the attribute information to the orbital angular momentum modulation of the coherent optical carrier; The relational information is mapped to the wavelength modulation of the coherent optical carrier.

7. The semantic information extraction and representation method for coherent optical communication according to claim 6, characterized in that, The generation of semantically modulated coherent optical signals includes: By using an IQ modulator, a polarization controller, an orbital angular momentum modulator, and a tunable laser, the amplitude modulation, polarization state modulation, phase modulation, orbital angular momentum modulation, and wavelength modulation are synchronously and jointly applied to obtain the semantically modulated coherent optical signal.

8. A semantic information extraction and representation system for coherent optical communication, characterized in that, include: The data acquisition module is used to acquire the original images and task prompts; The task parsing module is used to perform semantic parsing on the task prompts to obtain structured task requirements; The object detection and semantic segmentation module is used to perform object detection on the original image based on the task requirements, identify entities that match the task requirements, perform semantic segmentation on the identified entities, and generate pixel-level segmentation masks for each entity. An attribute-aware module is used to extract attribute information of each entity based on the pixel-level segmentation mask; The relation reasoning module is used to analyze the spatial and / or logical relationships between the entities based on the pixel-level segmentation mask and the spatial coordinates of each entity, and to obtain relation information. The scene graph generation module is used to construct a structured semantic scene graph based on the identified entities, the attribute information, and the relationship information. The semantic scene graph includes a node set and an edge set. The node set contains entity category information and attribute information, and the edge set contains the relationship information. The coherent optical modulation module is used to map the entity category information, attribute information and relation information in the semantic scene graph to different physical degrees of freedom of the coherent optical carrier, respectively, to generate a semantically modulated coherent optical signal.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the semantic information extraction and characterization method for coherent optical communication as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the semantic information extraction and characterization method for coherent optical communication as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Semantic coherent light communication method and system

    CN115842593A

  • Remote sensing scene graph guided semantic information reasoning method and device, equipment and medium

    CN120599616A