Underwater image transmission method and device based on multi-level semantic decoupling and diffusion reconstruction

CN122802628APending Publication Date: 2026-09-22XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610834944.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]第一,现有图像语义传输方法通常将水下图像编码为单一路径的隐变量特征,缺少对水下图像中不同层级语义信息的显式区分

Benefits of technology

[0015]根据本申请的基于多级语义解耦与扩散重构的水下图像传输方法,其有益效果在于:通过将水下图像解耦为场景、结构、内容三级语义并分别编码传输,在接收端经跨模态校验后以多维语义条件协同约束扩散重构,实现了在降低传输数据量的同时,提高重构图像的场景一致性、结构稳定性和内容语义可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802628A_ABST
    Figure CN122802628A_ABST
Patent Text Reader

Abstract

The application discloses a kind of underwater image transmission method and device based on multistage semantic decoupling and diffusion reconstruction.The method comprises: the original underwater image obtained is separated by semantics;Different coding networks are used to code the three types of semantics separated, to obtain scene semantic matrix, structure semantic vector and content semantic vector with different preset dimensions, and hierarchical transmission is carried out;The received semantic representation is processed by semantics at the receiving end, to obtain scene graph, structure edge graph and candidate text description;According to scene graph and candidate text description, cross-modal consistency verification is carried out, and according to the verification result, it is determined to trigger retransmission or enter reconstruction;After determining to enter reconstruction, scene graph is used as low-frequency scene priori, structure edge graph is used as spatial structure constraint, and candidate text description is used as text semantic condition, multidimensional semantic condition diffusion reconstruction is carried out, and underwater reconstructed image is output;Under the condition that underwater acoustic communication bandwidth is limited, high-reliability underwater image transmission can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater image transmission technology, and in particular to an underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction, an underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction, a computer-readable storage medium, and a computer device. Background Technology

[0002] With the increasing demands for ocean observation, seabed inspection, deep-sea resource exploration, and underwater robotic operations, the real-time and reliable transmission of underwater optical images has become a crucial foundation for underwater intelligent sensing and remote decision-making. Autonomous underwater vehicles (AUVs), remotely operated underwater vehicles (ROVs), and fixed underwater observation platforms typically need to transmit acquired underwater images to a receiving end for analysis and reconstruction when performing tasks such as target identification, environmental monitoring, and anomaly detection.

[0003] However, underwater wireless communication environments are complex, especially in underwater acoustic communication scenarios, where channels typically have limited available bandwidth, significant noise interference, large transmission delays, and rapidly changing channel states. Traditional image transmission methods usually compress the entire image into a unified bitstream before transmission. When bandwidth is limited or channel quality degrades, problems such as image blurring, structural loss, or content distortion can easily occur, making it difficult to meet the image availability requirements for underwater target observation and remote identification tasks.

[0004] In recent years, deep learning-based image compression, joint source-channel coding, and semantic communication methods have provided new technical approaches for underwater image transmission. These methods no longer simply pursue pixel-level bitstream transmission, but instead attempt to transmit semantic features that are more important for subsequent understanding and reconstruction. However, existing methods still have the following shortcomings when applied to underwater image transmission.

[0005] First, existing image semantic transmission methods typically encode underwater images into latent variable features along a single path, lacking explicit differentiation of semantic information at different levels within the underwater image. The overall background, color and lighting, target structure, and content category in an underwater image play different roles in the reconstruction process. Using a uniform encoding method to process all information can easily lead to unreasonable allocation of encoding resources, making it difficult to balance scene consistency, structural stability, and content accuracy with limited transmission resources.

[0006] Second, existing methods still have limited ability to express structural semantics. Underwater images often suffer from problems such as turbidity, low illumination, color shift, and local occlusion. Simple edge maps or single segmentation results are insufficient to stably represent target contours, boundary details, and spatial geometric layout. When structural information is disturbed during transmission or reconstruction, the image generated at the receiving end is prone to problems such as target shape deformation, edge breakage, or inaccurate spatial relationships.

[0007] Third, existing content semantic transmission methods lack reliable matching and verification mechanisms for underwater images. When performing generative reconstruction at the receiving end, if the content semantics are inconsistent with the scene semantics, the reconstruction result may have good visual clarity, but it may not match the target category, attributes, or scene meaning in the original image, thus affecting the reliability of subsequent recognition, detection, and manual interpretation.

[0008] Fourth, existing generative image reconstruction methods typically rely on single conditions or latent variables for image restoration, lacking collaborative constraints across scene, structure, and content semantics. For underwater images, relying solely on single path features makes it difficult to simultaneously maintain consistency between global background, target contours, and semantic content. Therefore, an underwater image transmission and reconstruction scheme capable of fusing multi-level semantic information and performing cross-modal verification before reconstruction is needed.

[0009] Therefore, how to establish hierarchical semantic representations for scene information, structural information, and content information of underwater images, how to perform multi-path coding and hierarchical transmission based on the information content and reconstruction effect of different semantics, and how to improve the stability and semantic consistency of image reconstruction at the receiving end through cross-modal semantic verification and multi-dimensional semantic conditional diffusion reconstruction are the technical problems that need to be solved in the field of underwater image semantic transmission. Summary of the Invention

[0010] This invention aims to at least partially solve one of the technical problems in the aforementioned technologies. To this end, one objective of this invention is to propose an underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction. By decoupling underwater images into three levels of semantics—scene, structure, and content—and encoding and transmitting them separately, and then performing cross-modal verification at the receiving end, diffusion reconstruction is performed under multi-dimensional semantic conditions. This achieves improved scene consistency, structural stability, and content semantic reliability of the reconstructed image while reducing the amount of transmitted data.

[0011] A second objective of this invention is to provide a computer-readable storage medium.

[0012] The third objective of this invention is to provide a computer device.

[0013] The fourth objective of this invention is to propose an underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction.

[0014] To achieve the above objectives, a first aspect of the present invention proposes an underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction, comprising the following steps: acquiring an original underwater image; performing semantic separation on the original underwater image to extract scene semantics, structural semantics, and content semantics; encoding the scene semantics, structural semantics, and content semantics using different coding networks to obtain scene semantic matrices, structural semantic vectors, and content semantic vectors with different preset dimensions, and performing hierarchical transmission; receiving the scene semantic matrix, structural semantic vector, and content semantic vector transmitted through an underwater acoustic channel, and performing semantic processing on the scene semantic matrix, structural semantic vector, and content semantic vector to obtain corresponding scene graphs, structural edge graphs, and candidate text descriptions; performing cross-modal consistency verification based on the scene graph and the candidate text descriptions to determine whether to trigger retransmission or enter reconstruction based on the verification result; after determining to enter reconstruction, performing multi-dimensional semantic conditional diffusion reconstruction based on the scene graph, structural edge graph, and candidate text descriptions to obtain an underwater reconstructed image.

[0015] The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction according to this application has the following advantages: by decoupling the underwater image into three levels of semantics—scene, structure, and content—and encoding and transmitting them separately, and then performing cross-modal verification at the receiving end, diffusion reconstruction is carried out under multi-dimensional semantic conditions, thereby reducing the amount of transmitted data while improving the scene consistency, structural stability, and content semantic reliability of the reconstructed image.

[0016] Optionally, a scene semantic coding network is used to semantically encode the background, color, lighting, and global environment information of the original underwater image to obtain a scene semantic matrix, wherein the scene semantic matrix includes a floating-point feature values.

[0017] Optionally, the original underwater image is subjected to multi-scale edge extraction using the HED soft edge extraction network to obtain a structural edge map, and the structural edge map is semantically encoded using a structural semantic convolutional encoder to obtain a structural semantic vector, wherein the structural semantic vector includes b floating-point feature values.

[0018] Optionally, the text description corresponding to the original underwater image is obtained, and the text description is semantically encoded using a text semantic encoder. The encoded result is then mapped to a content semantic vector using a multi-layer projector, wherein the content semantic vector includes c floating-point feature values. .

[0019] Optionally, semantic processing is performed on the scene semantic matrix, structural semantic vector, and content semantic vector to obtain the corresponding scene graph, structural edge graph, and candidate text description, including: The scene semantic matrix is ​​decoded using a scene semantic decoder to obtain a scene graph; The structural semantic vector is decoded using a structural semantic decoder to obtain a structural edge map; Based on the content semantic vector, nearest neighbor matching is performed in a preset image and text semantic prior library to obtain candidate text descriptions.

[0020] Optionally, cross-modal consistency verification is performed based on the scene graph and the candidate text description to determine whether to trigger retransmission or enter reconstruction based on the verification result, including: The scene graph is extracted using an image semantic encoder; A text semantic encoder is used to extract the text semantic features of the candidate text description; Calculate the cross-modal consistency score between the image semantic features and the text semantic features; If the cross-modal consistency score is less than the preset cross-modal consistency threshold, the corresponding verification result is determined to trigger retransmission; If the cross-modal consistency score is greater than or equal to the preset cross-modal consistency threshold, then the corresponding verification result is determined to proceed to reconstruction.

[0021] Optionally, when performing multidimensional semantic conditional diffusion reconstruction based on the scene graph, structural edge graph, and candidate text description, the scene graph is used as a low-frequency scene prior, the structural edge graph is used as a spatial structural constraint, and the candidate text description is used as a textual semantic condition. Diffusion reconstruction is performed under multidimensional semantic condition constraints to output an underwater reconstructed image.

[0022] To achieve the above objectives, a second aspect of the present invention provides a computer-readable storage medium storing an underwater image transmission program based on multi-level semantic decoupling and diffusion reconstruction. When executed by a processor, the underwater image transmission program based on multi-level semantic decoupling and diffusion reconstruction implements the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described above.

[0023] To achieve the above objectives, a third aspect of the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described above.

[0024] To achieve the above objectives, a fourth aspect of the present invention proposes an underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction, comprising: an acquisition module for acquiring original underwater images; a multi-level semantic decoupling module for performing semantic separation on the original underwater images to extract scene semantics, structural semantics, and content semantics; a multi-path semantic encoding and hierarchical transmission module for encoding the scene semantics, structural semantics, and content semantics using different encoding networks to obtain scene semantic matrices, structural semantic vectors, and content semantic vectors with different preset dimensions, and performing hierarchical transmission; and a semantic processing module for receiving... The underwater acoustic channel transmits scene semantic matrix, structural semantic vector, and content semantic vector, and performs semantic processing on these vectors to obtain corresponding scene graphs, structural edge graphs, and candidate text descriptions. A cross-modal semantic verification module performs cross-modal consistency verification based on the scene graph and candidate text descriptions to determine whether to trigger retransmission or enter reconstruction based on the verification result. A multi-dimensional semantic conditional diffusion reconstruction module performs multi-dimensional semantic conditional diffusion reconstruction based on the scene graph, structural edge graph, and candidate text descriptions after determining whether to enter reconstruction, to obtain an underwater reconstructed image. Attached Figure Description

[0025] Figure 1 A flowchart illustrating the underwater image transmission method based on multi-level semantic decoupling and multi-dimensional semantic conditional diffusion reconstruction provided in an embodiment of the present invention; Figure 2 A schematic diagram of the overall structure of an underwater image transmission system based on multi-level semantic decoupling and multi-dimensional semantic conditional diffusion reconstruction provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of multi-channel semantic coding and hierarchical transmission at the transmitter provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of cross-modal semantic verification and multi-dimensional semantic conditional diffusion reconstruction at the receiver provided in an embodiment of the present invention; Figure 5 This is a block diagram of an underwater image transmission device based on multi-level semantic decoupling and multi-dimensional semantic conditional diffusion reconstruction according to an embodiment of the present invention. Detailed Implementation

[0026] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0027] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the invention to those skilled in the art.

[0028] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0029] like Figure 2 As shown, the underwater image transmission method performs multi-level semantic decoupling and multi-path semantic coding on the input underwater image at the transmitting end, and cross-modal semantic verification and multi-dimensional semantic conditional diffusion reconstruction at the receiving end. The corresponding underwater image transmission method based on multi-level semantic decoupling and multi-dimensional semantic conditional diffusion reconstruction is as follows: Figure 1 As shown, it includes the following steps: S101, acquire the original underwater image, and perform semantic separation on the original underwater image to extract scene semantics, structural semantics and content semantics.

[0030] Specifically, the transmitting end acquires the input underwater image to be transmitted, denoted as X. The input underwater image X can be a single frame image or any frame image from an image sequence obtained by an underwater robot, remotely operated vehicle, underwater fixed observation platform, or other underwater image acquisition equipment.

[0031] The transmitter performs semantic separation on the input underwater image X through a multi-level semantic decoupling module to obtain scene semantics, structural semantics, and content semantics.

[0032] Scene semantics is used to characterize the overall background, color distribution, lighting conditions, and global environmental information of underwater images; structural semantics is used to characterize the edges, contours, shapes, and spatial geometric layout of underwater targets; and content semantics is used to characterize the category, attributes, and textual description information of core targets in the image.

[0033] Through the above semantic separation, the transmitting end no longer encodes the entire underwater image as a single path latent variable, but separates and processes semantic information of different levels and functions in the underwater image, providing a foundation for subsequent multi-path semantic coding and hierarchical transmission.

[0034] S102 uses different coding networks to encode scene semantics, structural semantics and content semantics, to obtain scene semantic matrices, structural semantic vectors and content semantic vectors with different preset dimensions, and then transmits them in a hierarchical manner.

[0035] As an example, a scene semantic coding network is used to semantically encode the background, color, lighting and global environment information of the original underwater image to obtain a scene semantic matrix, wherein the scene semantic matrix includes a floating-point feature values.

[0036] Specifically, for scene semantics, the transmitting end uses a scene semantic coding network to process the input underwater image. The overall background, color distribution, lighting conditions, and global environment information are encoded to obtain the scene semantic matrix. The process can be represented as follows:

[0037] in, This indicates the input of an underwater image; Represents a scene semantic encoding network; This represents the scene semantic matrix output by the scene semantic encoding network.

[0038] In a preferred embodiment, the scene semantic coding network employs a Swing Transformer-based visual coding network, extracting global scene representations of underwater images through local window attention and shifted window attention. For a size of... Input underwater image, scene semantic matrix The preferred feature is 4096 floating-point feature values.

[0039] Scene semantic matrix It is used to recover the background, color distribution, illumination status and global environment information of underwater images at the receiving end, and to provide low-frequency scene priors for subsequent multidimensional semantic conditional diffusion reconstruction.

[0040] As an example, the HED soft edge extraction network is used to perform multi-scale edge extraction on the original underwater image to obtain a structural edge map. The structural edge map is then semantically encoded by a structural semantic convolutional encoder to obtain a structural semantic vector, wherein the structural semantic vector includes b floating-point feature values.

[0041] Specifically, regarding structural semantics, the transmitter first processes the input underwater image through the HED soft edge extraction network. Multi-scale edge extraction is performed to obtain the structure edge map. .

[0042] The process is represented as follows:

[0043] in, This represents the structural edge map output by the HED soft edge extraction network; This represents the HED soft edge extraction network; The input underwater image is represented. The HED soft edge extraction network performs multi-level convolutional feature extraction on the input underwater image to obtain edge response maps at multiple scales. These edge response maps are then upsampled to the spatial dimensions corresponding to the input underwater image and fused to obtain the structural edge map. This structural edge map can characterize the edges, contours, shape, and spatial geometry of underwater targets. Subsequently, the launcher will transmit the structural edge map... Input the structural semantic convolutional encoder to obtain structural semantic vectors. The process can be represented as follows:

[0044] in, Represents a structural semantic vector; This represents a structural semantic convolutional encoder; This represents the structural edge map output by the HED soft edge extraction network.

[0045] In a preferred embodiment, the structural semantic vector Includes 1024 floating-point feature values. Structural semantic vector. Used to recover structural edge maps or structural feature maps at the receiving end, and as spatial structural constraints in multidimensional semantic conditional diffusion reconstruction.

[0046] As one embodiment, the text description corresponding to the original underwater image is obtained, and the text description is semantically encoded using a text semantic encoder. The encoded result is then mapped to a content semantic vector using a multi-layer projector. The content semantic vector includes c floating-point feature values, and... .

[0047] Specifically, regarding content semantics, the transmitting end acquires and inputs underwater images. The corresponding content text description is denoted as Content text description Used to characterize the category, attributes, and semantic relationships of core targets in an input underwater image. Content text description. It can come from a pre-set image and text annotation library, or it can be generated by a lightweight visual language model based on the input underwater image. Generation. The transmitting end uses a text semantic encoder to describe the content text. Semantic encoding is performed, and the encoded result is mapped to a content semantic vector through a multi-layer projector. The process can be represented as follows:

[0048] in, Represents a semantic vector of content; Indicates a multi-layer projector; This represents a text semantic encoder; This represents the text description corresponding to the input underwater image.

[0049] In a preferred embodiment, the text semantic encoder employs the CLIP text encoder, which uses content semantic vectors. Includes 512 floating-point feature values. Content semantic vector. It is used to obtain candidate text descriptions by matching them with a preset image and text semantic prior library at the receiving end, and to serve as text semantic conditions in multidimensional semantic condition diffusion reconstruction.

[0050] Specifically, in obtaining the scene semantic matrix Structural semantic vectors and content semantic vector Subsequently, the transmitting end uses a multi-channel semantic coding and hierarchical transmission module to transmit the three types of semantic representations in a hierarchical manner. Hierarchical transmission does not use the same coding resources for the three types of semantics, but rather uses different coding networks and different preset coding dimensions for transmission based on the differences in the amount of information of the three types of semantics and their different roles in image reconstruction.

[0051] In a preferred embodiment, the scene semantic matrix Includes 4096 floating-point feature values ​​to preserve overall background, color distribution, lighting conditions, and global environment information; structural semantic vectors Includes 1024 floating-point feature values ​​to preserve target edges, contours, shape, and spatial geometry; content semantic vectors. It includes 512 floating-point feature values, used to preserve the category, attributes, and textual semantic information of the core target.

[0052] Multi-path semantic representation can be expressed as:

[0053] in, This represents the set of multiple semantic representations to be transmitted; Represents the scene semantic matrix; Represents a structural semantic vector; This represents the content semantic vector. Through the hierarchical transmission method described above, this implementation ensures that scene semantics, structural semantics, and content semantics each occupy coding resources commensurate with their information content and reconstruction function, avoiding the redundant transmission problem caused by uniformly encoding the entire underwater image into a single path latent variable.

[0054] In other words, such as Figure 3 As shown, the image is split into three semantics, and different coding dimensions (4096 / 1024 / 512) are assigned to different coding networks.

[0055] S103 receives the scene semantic matrix, structural semantic vector, and content semantic vector transmitted through the underwater acoustic channel, and performs semantic processing on the scene semantic matrix, structural semantic vector, and content semantic vector to obtain the corresponding scene graph, structural edge graph, and candidate text description.

[0056] As one embodiment, semantic processing is performed on the scene semantic matrix, structural semantic vector, and content semantic vector to obtain the corresponding scene graph, structural edge graph, and candidate text description, including: decoding the scene semantic matrix using a scene semantic decoder to obtain the scene graph; decoding the structural semantic vector using a structural semantic decoder to obtain the structural edge graph; and performing nearest neighbor matching in a preset graph-text semantic prior library based on the content semantic vector to obtain the candidate text description.

[0057] Specifically, the receiving end receives multiple semantic representations transmitted by the transmitting end, and the multiple semantic representations include a scene semantic matrix. Structural semantic vectors and content semantic vector The receiving end first processes the scene semantic matrix through a scene semantic decoder. Decode the scene to obtain a scene map or low-frequency scene prior. The process is represented as follows:

[0058] in, This represents the scene graph or low-frequency scene prior obtained by decoding the scene semantic matrix; Represents a scene semantic decoder; This represents the scene semantic matrix received by the receiver. It can also represent a scene graph or low-frequency scene prior. Used to provide overall background, color distribution, lighting conditions, and global environmental information for underwater images.

[0059] The receiving end processes the structural semantic vector through a structural semantic decoder. Decoding is performed to obtain a structural edge map or structural feature map. The process is represented as follows:

[0060] in, This represents the structural edge map or structural feature map obtained by decoding the structural semantic vector; Represents a structural semantic decoder; This represents the structural semantic vector received by the receiver. (Structure edge map or structure feature map) Used to provide information on the target's edges, contours, shape, and spatial geometry.

[0061] The receiving end uses the content semantic vector In the pre-set image and text semantic prior library Nearest neighbor matching is performed to obtain candidate text descriptions. Preset image and text semantic prior library This includes underwater image samples, corresponding text descriptions, and their semantic encodings. The nearest neighbor matching process is represented as:

[0062] in, This represents the candidate text description obtained by content semantic vector matching; Represents a predefined semantic library for images and text. The first in A text description; This indicates a pre-defined semantic library for images and text; This represents the semantic vector of the content received by the receiving end. This represents a text semantic encoder; Indicates a multi-layer projector; Represents text description Vector representation in the content semantic coding space; This represents the calculation of cosine similarity. Through the above processing, the receiving end obtains three types of conditional information for reconstruction: scene image or low-frequency scene prior. Structural edge diagram or structural feature diagram and candidate text descriptions .

[0063] S104. Perform cross-modal consistency verification based on the scene graph and candidate text description, and determine whether to trigger retransmission or enter reconstruction based on the verification result.

[0064] As one embodiment, cross-modal consistency verification is performed based on the scene graph and candidate text descriptions to determine whether to trigger retransmission or enter reconstruction based on the verification result. This includes: extracting image semantic features of the scene graph using an image semantic encoder; extracting text semantic features of the candidate text descriptions using a text semantic encoder; calculating the cross-modal consistency score between the image semantic features and the text semantic features; if the cross-modal consistency score is less than a preset cross-modal consistency threshold, the corresponding verification result is determined to trigger retransmission; if the cross-modal consistency score is greater than or equal to the preset cross-modal consistency threshold, the corresponding verification result is determined to enter reconstruction.

[0065] It should be noted that, in order to reduce the probability of inconsistency between the generated result and the original image semantics, the receiver performs consistency verification on scene semantics and content semantics through a cross-modal semantic verification module before entering the multi-dimensional semantic conditional diffusion reconstruction.

[0066] Specifically, the cross-modal semantic verification module employs an image semantic encoder. Extract scene images or low-frequency scene priors Image semantic features were analyzed and a text semantic encoder was employed. Extract candidate text descriptions The text semantic features are then analyzed, and the cross-modal consistency score between the two is calculated. .

[0067] Cross-modal consistency score Represented as:

[0068] in, The cross-modal consistency score represents the relationship between scene semantics and content semantics. This represents an image semantic encoder; Represented by the scene semantic matrix The scene image or low-frequency scene prior obtained through decoding; This represents a text semantic encoder; Represented by content semantic vector In the pre-defined image and text semantic prior library Candidate text descriptions obtained through matching; This indicates the calculation of cosine similarity.

[0069] when When the cross-modal semantic verification module determines that the scene semantics and content semantics do not match, it indicates that there is a consistency deviation between the current scene semantic decoding result and the content semantic description. At this time, retransmission control is triggered, and the transmitter is requested to retransmit the scene semantic matrix first. It should be noted that text has higher credibility, so choosing to trust the text means the structure does not need to be retransmitted or verified. Therefore, retransmitting the semantic vector is the preferred option. If necessary, a retransmission of the content semantic vector can also be requested. .when At that time, the cross-modal semantic verification module determines that the scene semantics and content semantics meet the consistency requirements, allowing scene images or low-frequency scene priors. Structural edge diagram or structural feature diagram and candidate text descriptions Enter the multidimensional semantic conditional diffusion reconstruction module.

[0070] in, This indicates the preset cross-modal consistency threshold. The preset cross-modal consistency threshold can be set in advance according to different underwater mission scenarios, the size of the text and semantic prior library, or the reconstruction quality requirements.

[0071] S105, after determining to enter the reconstruction, multi-dimensional semantic conditional diffusion reconstruction is performed based on the scene map, structure edge map and candidate text description to obtain the underwater reconstructed image.

[0072] As an example, when performing multidimensional semantic conditional diffusion reconstruction based on scene graph, structural edge graph and candidate text description, the scene graph is used as a low-frequency scene prior, the structural edge graph is used as a spatial structural constraint, and the candidate text description is used as a text semantic condition. Diffusion reconstruction is performed under multidimensional semantic condition constraints to output an underwater reconstructed image.

[0073] Specifically, after the cross-modal semantic verification is passed, the receiving end reconstructs the underwater image using a multi-dimensional semantic conditional diffusion reconstruction module. This module simultaneously utilizes scene semantics, structural semantics, and content semantics to form multi-dimensional semantic conditional constraints.

[0074] Among them, scene images or low-frequency scene priors Used as a priori for low-frequency scenes to constrain the overall background, color distribution, illumination state, and global environment information of the reconstructed image; structure edge map or structure feature map. Used as a spatial structural constraint to constrain the edges, contours, shapes, and spatial geometric layout of objects in a reconstructed image; candidate text description. Used as a textual semantic condition to constrain the core target categories, attributes, and semantic relationships in the reconstructed image.

[0075] The multidimensional semantic conditional diffusion reconstruction process is represented as:

[0076] in, This represents the final output underwater reconstructed image; This represents a multidimensional semantic conditional diffusion reconstruction network; Represents scene images or low-frequency scene priors; Show the structural edge diagram or structural feature diagram; This indicates the candidate text description.

[0077] In a preferred embodiment, the multidimensional semantic conditional diffusion reconstruction module is based on scene images or low-frequency scene priors. The process of constructing the initial state or low-frequency guiding conditions for diffusion reconstruction is expressed as follows:

[0078] in, Indicates diffusion reconstruction at time step The initial state; Display scene images or low-frequency scene priors; Show in time step Introduced noise.

[0079] In a preferred embodiment, the multidimensional semantic conditional diffusion reconstruction module, during the training phase, bases the structure edge map or structure feature map on the structure edge map or structure feature map. Generate spatial weight graph We construct a weighted reconstruction loss based on the spatial weight map to improve the reconstruction fidelity of target edges, contours and structural regions.

[0080] The weighted reconstruction loss is expressed as:

[0081] in, Indicates the weighted reconstruction loss; This represents a structure edge map or structure feature map. The generated spatial weight map; It shows element-wise product; Represents an underwater reconstructed image; This indicates the input of an underwater image; This represents the square norm 2.

[0082] In other words, such as Figure 4 As shown, through the above-mentioned multidimensional semantic conditional diffusion reconstruction process, the receiving end can generate underwater reconstructed images under the joint constraints of scene semantics, structural semantics and content semantics, thereby improving the scene consistency, structural stability and content reliability of the reconstruction results.

[0083] In summary, the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction proposed in this application has the following beneficial effects: (1) Improved the efficiency of semantic expression under limited transmission resources.

[0084] This application decomposes underwater images into scene semantics, structural semantics, and content semantics. Based on the differences in information content and their respective roles in image reconstruction, different coding networks and preset coding dimensions are set for each type of semantics. Compared to encoding the entire image as a single latent variable feature, this application enables background, structural, and content information to be expressed in forms more suitable to their semantic attributes, thereby reducing redundant information transmission and improving semantic transmission efficiency under limited underwater communication conditions.

[0085] (2) It enhances the stability of the semantic expression of underwater image structure.

[0086] This invention obtains a structural edge map through a HED soft edge extraction network and compresses it into a structural semantic vector using a structural semantic convolutional encoder. Because the HED soft edge extraction network can fuse multi-scale edge responses, it is more effective at preserving the edges, contours, shapes, and spatial geometry of underwater targets compared to a single edge operator or segmentation result. After the receiver decodes the structural semantic vector into a structural edge map or structural feature map, it can serve as a spatial structural constraint during the diffusion reconstruction process, thereby reducing the probability of target shape deformation, edge loss, and spatial relationship deviations.

[0087] (3) Improved the semantic consistency discrimination capability of the receiver before reconstruction.

[0088] This invention includes a cross-modal semantic verification module at the receiving end, which decodes the received scene semantic matrix to obtain the scene image or low-frequency scene prior. And based on the received content semantic vector, candidate text descriptions are obtained by matching them in a preset image and text semantic prior library. Then, the cross-modal consistency score between the two is calculated. Through this verification mechanism, the system can determine whether the scene semantics and content semantics match before image reconstruction; when the consistency score is lower than a preset threshold, the scene semantics are retransmitted first, thereby reducing the probability of the generated image content being offset due to the damage to the scene semantics.

[0089] (4) Improved the scene consistency, structural stability and content reliability of generative reconstruction results.

[0090] The multidimensional semantic conditional diffusion reconstruction module of this invention utilizes three types of semantic conditions simultaneously for underwater image reconstruction: the scene semantic decoding result is used as a low-frequency scene prior to constrain the background, color, and lighting of the reconstructed image; the structural semantic decoding result is used as a spatial structure constraint to constrain target edges, contours, and geometric layout; and the candidate text description obtained from content semantic matching is used as a text semantic condition to constrain the core target categories and attributes in the image. Compared with reconstruction methods that rely solely on a single path latent variable or a single condition, this invention can simultaneously consider the global scene, spatial structure, and content semantics of the underwater image during the reconstruction process, which is beneficial for improving the overall usability and semantic consistency of the reconstructed image.

[0091] (5) Improved the applicability of the system to underwater application scenarios.

[0092] This invention does not rely on high-volume data transmission of the entire underwater image. Instead, it prioritizes the preservation of the most critical information for reconstruction through hierarchical expression and transmission of scene semantics, structural semantics, and content semantics. This approach is applicable to mission scenarios such as underwater robot inspection, marine target observation, deep-sea environmental monitoring, and underwater anomaly identification, providing a feasible technical path for remote sensing and reconstruction of underwater images under limited communication conditions.

[0093] In addition, the present invention also proposes a computer-readable storage medium storing an underwater image transmission program based on multi-level semantic decoupling and diffusion reconstruction. When the underwater image transmission program based on multi-level semantic decoupling and diffusion reconstruction is executed by a processor, it implements the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described above.

[0094] In addition, this invention also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described above.

[0095] To achieve the above embodiments, this application also proposes an underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction, such as... Figure 5 As shown, the underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction includes: an acquisition module 10, a multi-level semantic decoupling module 20, a multi-channel semantic encoding and hierarchical transmission module 30, a semantic processing module 40, a cross-modal semantic verification module 50, and a multi-dimensional semantic conditional diffusion reconstruction module 60.

[0096] The system comprises the following modules: an acquisition module 10 for acquiring raw underwater images; a multi-level semantic decoupling module 20 for semantic separation of the raw underwater images to extract scene semantics, structural semantics, and content semantics; a multi-path semantic coding and hierarchical transmission module 30 for encoding the scene semantics, structural semantics, and content semantics using different coding networks to obtain scene semantic matrices, structural semantic vectors, and content semantic vectors with different preset dimensions, and for hierarchical transmission; a semantic processing module 40 for receiving the scene semantic matrix, structural semantic vector, and content semantic vector transmitted through the underwater acoustic channel, and for performing semantic processing on the scene semantic matrix, structural semantic vector, and content semantic vector to obtain the corresponding scene graph, structural edge graph, and candidate text description; a cross-modal semantic verification module 50 for performing cross-modal consistency verification based on the scene graph and candidate text description to determine whether to trigger retransmission or enter reconstruction based on the verification result; and a multi-dimensional semantic conditional diffusion reconstruction module 60 for performing multi-dimensional semantic conditional diffusion reconstruction based on the scene graph, structural edge graph, and candidate text description after determining to enter reconstruction, to obtain the underwater reconstructed image.

[0097] It should be noted that the foregoing explanation of the embodiment of underwater image transmission based on multi-level semantic decoupling and diffusion reconstruction also applies to the underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction in this embodiment, and will not be repeated here.

[0098] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0099] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0100] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0101] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0102] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0103] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.

[0104] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

[0105] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0106] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0107] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0108] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0109] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. An underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction, characterized in that, Includes the following steps: Acquire raw underwater images; Semantic separation is performed on the original underwater image to extract scene semantics, structural semantics, and content semantics; The scene semantics, structural semantics, and content semantics are encoded using different encoding networks to obtain scene semantic matrices, structural semantic vectors, and content semantic vectors with different preset dimensions, and then transmitted in a hierarchical manner. Receive the scene semantic matrix, structural semantic vector and content semantic vector transmitted through the underwater acoustic channel, and perform semantic processing on the scene semantic matrix, structural semantic vector and content semantic vector to obtain the corresponding scene graph, structural edge graph and candidate text description; Cross-modal consistency verification is performed based on the scene diagram and the candidate text description, and the result of the verification determines whether to trigger retransmission or enter reconstruction. After determining whether to proceed with reconstruction, multidimensional semantic conditional diffusion reconstruction is performed based on the scene graph, structural edge graph, and candidate text description to obtain the underwater reconstructed image.

2. The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in claim 1, characterized in that, A scene semantic coding network is used to semantically encode the background, color, lighting, and global environment information of the original underwater image to obtain a scene semantic matrix, wherein the scene semantic matrix includes a floating-point feature values.

3. The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in claim 2, characterized in that, The original underwater image is subjected to multi-scale edge extraction using the HED soft edge extraction network to obtain a structural edge map. The structural edge map is then semantically encoded using a structural semantic convolutional encoder to obtain a structural semantic vector, wherein the structural semantic vector includes b floating-point feature values.

4. The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in claim 3, characterized in that, The text description corresponding to the original underwater image is obtained, and the text description is semantically encoded using a text semantic encoder. The encoded result is then mapped to a content semantic vector using a multi-layer projector. The content semantic vector includes c floating-point feature values. .

5. The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in claim 1, characterized in that, Semantic processing is performed on the scene semantic matrix, structural semantic vector, and content semantic vector to obtain the corresponding scene graph, structural edge graph, and candidate text description, including: The scene semantic matrix is ​​decoded using a scene semantic decoder to obtain a scene graph; The structural semantic vector is decoded using a structural semantic decoder to obtain a structural edge map; Based on the content semantic vector, nearest neighbor matching is performed in a preset image and text semantic prior library to obtain candidate text descriptions.

6. The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in claim 1, characterized in that, Cross-modal consistency verification is performed based on the scene diagram and the candidate text description, and the result of the verification determines whether to trigger retransmission or enter reconstruction, including: The scene graph is extracted using an image semantic encoder; A text semantic encoder is used to extract the text semantic features of the candidate text description; Calculate the cross-modal consistency score between the image semantic features and the text semantic features; If the cross-modal consistency score is less than the preset cross-modal consistency threshold, the corresponding verification result is determined to trigger retransmission; If the cross-modal consistency score is greater than or equal to the preset cross-modal consistency threshold, then the corresponding verification result is determined to proceed to reconstruction.

7. The underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in claim 1, characterized in that, When performing multidimensional semantic conditional diffusion reconstruction based on the scene graph, structural edge graph, and candidate text description, the scene graph is used as a low-frequency scene prior, the structural edge graph is used as a spatial structural constraint, and the candidate text description is used as a text semantic condition. Diffusion reconstruction is performed under multidimensional semantic condition constraints to output an underwater reconstructed image.

8. A computer-readable storage medium, characterized in that, It stores an underwater image transmission program based on multi-level semantic decoupling and diffusion reconstruction. When the underwater image transmission program based on multi-level semantic decoupling and diffusion reconstruction is executed by the processor, it implements the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in any one of claims 1-7.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the underwater image transmission method based on multi-level semantic decoupling and diffusion reconstruction as described in any one of claims 1-7.

10. An underwater image transmission device based on multi-level semantic decoupling and diffusion reconstruction, characterized in that, include: The acquisition module is used to acquire raw underwater images; A multi-level semantic decoupling module is used to perform semantic separation on the original underwater image in order to extract scene semantics, structural semantics and content semantics; The multi-channel semantic coding and hierarchical transmission module is used to encode the scene semantics, structural semantics and content semantics using different coding networks to obtain scene semantic matrices, structural semantic vectors and content semantic vectors with different preset dimensions, and to perform hierarchical transmission. The semantic processing module is used to receive the scene semantic matrix, structural semantic vector and content semantic vector transmitted through the underwater acoustic channel, and to perform semantic processing on the scene semantic matrix, structural semantic vector and content semantic vector to obtain the corresponding scene graph, structural edge graph and candidate text description. The cross-modal semantic verification module is used to perform cross-modal consistency verification based on the scene graph and the candidate text description, so as to determine whether to trigger retransmission or enter reconstruction based on the verification result; The multidimensional semantic conditional diffusion reconstruction module is used to perform multidimensional semantic conditional diffusion reconstruction based on the scene map, structure edge map and candidate text description after determining that reconstruction is to be entered, so as to obtain an underwater reconstructed image.