Semantic communication method and apparatus, system, splicing model training method and apparatus
By generating panoramic images at the receiving end using semantic communication methods, the problems of low transmission efficiency and insufficient channel resources in existing panoramic image technologies are solved, and efficient and reliable panoramic image data transmission is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies are inefficient and put a lot of pressure on channel resources when transmitting panoramic image data in real time, especially under poor channel conditions where there is a "cliff effect". Traditional source compression methods and channel coding cannot effectively solve this problem.
By employing a semantic communication method, semantic extraction information transmitted through the receiving channel is processed to perform received signal processing and multi-scale semantic feature extraction, generating panoramic images, reducing the pressure on channel data transmission and improving the reliability of image stitching.
Generating panoramic images directly at the receiving end reduces the amount of data transmitted through the channel, improves the transmission efficiency and reliability of panoramic images, and reduces data pressure during communication.
Smart Images

Figure CN119276423B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer science, specifically to the technical fields of large models, deep learning, and image processing, and in particular to a semantic communication method and apparatus, a semantic communication system, a splicing model training method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Currently, before panoramic image data stitched from multiple video streams can reach the user end, it must first be captured by a panoramic camera to obtain multiple video streams, which are then sent to a device with sufficient computing power for stitching and other processing, and finally transmitted via a communication link. This approach is inefficient for real-time transmission. Furthermore, due to the massive data volume of services corresponding to panoramic images, traditional source compression methods and channel coding still place significant pressure on channel resources. Summary of the Invention
[0003] This disclosure provides a semantic communication method and apparatus, a semantic communication system, a splicing model training method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] According to the first aspect, a semantic communication method is provided, the method comprising: receiving semantic extraction information of at least two images to be stitched transmitted through a channel; performing received signal processing on the semantic extraction information to obtain decoded information; performing multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and obtaining a panoramic image of the corresponding images to be stitched based on the semantic feature information.
[0005] According to the second aspect, another semantic communication method is provided, which includes: receiving at least two images to be stitched; inputting the images to be stitched into a semantic information extractor to obtain initial semantic information output by the semantic information extractor; performing transmission signal processing on the initial semantic information to obtain semantic extraction information; and sending the semantic extraction information to a receiving end through a channel so that the receiving end can perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the images to be stitched based on the extracted semantic feature information.
[0006] According to the third aspect, another semantic communication method is provided, which includes: a sending end obtaining semantic extraction information of at least two images to be stitched; the sending end sending the semantic extraction information to a receiving end through a channel; the receiving end receiving the semantic extraction information; the receiving end performing received signal processing on the semantic extraction information to obtain decoded information; the receiving end performing multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and the receiving end obtaining a panoramic image of the corresponding images to be stitched based on the semantic feature information.
[0007] According to the fourth aspect, a method for training a stitching model is provided. This method includes: selecting image samples from an image sample set; inputting the image samples into a semantic information extractor in an initial stitching model to obtain initial semantic information; performing transmit signal processing on the initial semantic information to obtain semantic extraction information; adding noise to the semantic extraction information and performing receive signal processing on the semantic extraction information to obtain decoded information; using a multi-scale semantic feature extractor in the stitching model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales; using a diffusion model in the stitching model to process the semantic feature information and image samples to obtain a panoramic image of the corresponding image samples; calculating the loss of the stitching model based on the panoramic image; and determining, based on the loss, that the stitching model meets the training completion conditions to obtain a trained stitching model.
[0008] According to a fifth aspect, a semantic communication apparatus is provided, comprising: an information receiving unit configured to receive semantic extraction information of at least two images to be stitched together transmitted via a channel; a processing unit configured to perform received signal processing on the semantic extraction information to obtain decoded information; a generation unit configured to perform multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and an image obtaining unit configured to obtain a panoramic image corresponding to the images to be stitched together based on the semantic feature information.
[0009] According to the sixth aspect, another semantic communication device is provided, comprising: an image receiving unit configured to receive at least two images to be stitched together; an input unit configured to input the images to be stitched together into a semantic information extractor to obtain initial semantic information output by the semantic information extractor; an information obtaining unit configured to perform transmission signal processing on the initial semantic information to obtain semantic extraction information; and a sending unit configured to send the semantic extraction information to a receiving end through a channel, so that the receiving end can perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the images to be stitched together based on the extracted semantic feature information.
[0010] According to a seventh aspect, a stitching model training apparatus is provided, comprising: a selection unit configured to select image samples from an image sample set; a training unit configured to input the image samples into a semantic information extractor in an initial stitching model to obtain initial semantic information; to perform transmit signal processing on the initial semantic information to obtain semantic extraction information; to add noise to the semantic extraction information and perform receive signal processing on the semantic extraction information to obtain decoded information; to perform multi-scale semantic feature extraction on the decoded information using a multi-scale semantic feature extractor in the stitching model to obtain multiple sets of semantic feature information at different scales; to process the semantic feature information and image samples using a diffusion model in the stitching model to obtain a panoramic image of the corresponding image samples; a calculation unit configured to calculate the loss of the stitching model based on the panoramic image; and a model obtaining unit configured to determine, based on the loss, that the stitching model meets the training completion conditions to obtain a trained stitching model.
[0011] According to the eighth aspect, a semantic communication system is provided, comprising: a receiver, a channel, and a transmitter; the transmitter is used to receive images to be stitched, including at least two images; input the images to be stitched into a semantic information extractor to obtain initial semantic information output by the semantic information extractor; perform transmission signal processing on the initial semantic information to obtain semantic extraction information; transmit the semantic extraction information to the receiver through the channel; the receiver is used to receive the semantic extraction information of at least two images to be stitched transmitted through the channel; perform reception signal processing on the semantic extraction information to obtain decoded information; perform multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and obtain a panoramic image of the corresponding images to be stitched based on the semantic feature information.
[0012] According to a ninth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first, second, or third aspect.
[0013] According to a tenth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the method described in any implementation of the first, second, or third aspect.
[0014] The semantic communication method and apparatus provided in the embodiments of this disclosure first receive semantic extraction information of at least two images to be stitched together transmitted through a channel; second, perform received signal processing on the semantic extraction information to obtain decoded information; third, perform multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and finally, obtain a panoramic image of the corresponding images to be stitched together based on the semantic feature information. Thus, by processing the semantic extraction information at the receiving end to generate a panoramic image of at least two images to be stitched together, the image stitching is placed at the receiving end, and the channel only needs to transmit semantic extraction information, reducing the channel data transmission pressure and improving the transmission efficiency of panoramic image data; by using multiple sets of semantic feature information at different scales to obtain a panoramic image, the reliability of obtaining the panoramic image is improved.
[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0016] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0017] Figure 1 This is a flowchart of an embodiment of the semantic communication method according to the present disclosure;
[0018] Figure 2 This is a schematic diagram of the system corresponding to the semantic communication method disclosed herein;
[0019] Figure 3 This is a flowchart of another embodiment of the semantic communication method according to the present disclosure;
[0020] Figure 4 This is a flowchart of an embodiment of the splicing model training method according to the present disclosure;
[0021] Figure 5 This is a schematic diagram of a structure of an embodiment of a semantic communication device according to the present disclosure;
[0022] Figure 6 This is a schematic diagram of another embodiment of the semantic communication device according to the present disclosure;
[0023] Figure 7 This is a schematic diagram of a structure of an embodiment of the splicing model training device according to the present disclosure;
[0024] Figure 8 This is a schematic diagram of a structure according to an embodiment of the semantic communication system of this disclosure;
[0025] Figure 9This is a block diagram of an electronic device used to implement the semantic communication method of the embodiments of this disclosure. Detailed Implementation
[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0027] With the advent of 6G, VR (Virtual Reality) is considered one of the most important high-traffic applications. Ubiquitous VR is one of the visions of 6G, but the massive data volume of VR and the need for data stitching and other processing mean that most VR applications are on-demand services, while live streaming requires enormous computing power.
[0028] Currently, before VR data reaches the user, it needs to be captured by a panoramic camera to obtain multiple video streams, which are then sent to a device with sufficient computing power for stitching and other processing. After source coding and channel coding, the data is transmitted through a communication link.
[0029] Such a solution is inefficient for real-time transmission. In addition, due to the huge amount of data in VR services, using traditional source compression methods and channel coding still puts significant pressure on channel resources, and will produce a "cliff effect" under harsh and variable channel conditions.
[0030] To address the aforementioned shortcomings, this disclosure proposes a semantic communication method that improves the efficiency of panoramic image data transmission. Figure 1 A flow 100 of an embodiment of a semantic communication method according to the present disclosure is shown, the semantic communication method comprising the following steps:
[0031] Step 101: Receive semantic extraction information of at least two images to be stitched from the channel transmission.
[0032] In this embodiment, a channel is a communication device that connects the transmitting end and the receiving end, transmitting signals from the transmitting end to the receiving end. Channels are divided into wireless channels and wired channels according to the transmission medium. Wireless channels use the propagation of electromagnetic waves in space to transmit signals, while wired channels use artificial light conduction or light signal transmission media to transmit signals. Optical fiber is currently the widely used transmission medium in wired optical communication systems.
[0033] In this embodiment, the at least two images to be stitched together can be images obtained by capturing an entity in the same scene using multiple different camera devices. Any two adjacent images to be stitched together may or may not have the same display area. For example, Figure 2 In this process, at least two images to be stitched (D) comprise two images to be stitched together, which have the same display area between them. It should be noted that the device for capturing the images to be stitched can be any type of camera with a different number of cameras, including fisheye lenses and pinhole lenses.
[0034] In this embodiment, the semantic extraction information is the semantic information obtained after semantic extraction of at least two images to be stitched together. The semantic extraction information is used to characterize the image meaning contained in the at least two images to be stitched together. For example, the semantic extraction information is the semantic feature map of at least two images to be stitched together.
[0035] In this embodiment, semantic extraction information can be obtained by identifying and extracting meaningful objects, scenes or concepts from at least two images to be stitched together through image processing and deep learning techniques. This process includes: performing contour extraction and semantic segmentation on at least two images to be stitched together to obtain semantic extraction information.
[0036] In this embodiment, the execution entity on which the semantic communication method runs can be, for example, such as... Figure 2 The receiver shown receives semantic extraction information transmitted in the channel. Figure 2 In this context, the channel between the transmitter and receiver can be a wireless channel.
[0037] The collection, storage, use, processing, transmission, provision, and disclosure of semantically extracted information in this technical solution are performed after authorization and comply with relevant laws and regulations. User-related information within the semantically extracted information is obtained with the user's permission.
[0038] Step 102: Process the received signal of the semantically extracted information to obtain the decoded information.
[0039] In this embodiment, received signal processing refers to processing the received channel data. Received signal processing includes decoding, demodulation, signal amplification, etc.
[0040] In this embodiment, the received signal processing is based on the processing performed by the semantic extraction information sending end on the transmitted signal. For example, if the sending end encodes the transmitted signal to obtain semantic extraction information, the received signal processing performed by the executing entity after receiving the semantic extraction information includes decoding the semantic extraction information. Alternatively, if the sending end adjusts the transmitted signal to obtain semantic extraction information, the received signal processing performed by the executing entity after receiving the semantic extraction information includes demodulating the semantic extraction information.
[0041] In one example, step 102 above includes: sequentially amplifying and decoding the semantically extracted information to obtain decoded information.
[0042] In this embodiment, the decoding information is information that characterizes the semantic features of at least two images to be stitched together. By processing the decoding information, at least two images to be stitched together can be obtained to obtain a panoramic image.
[0043] Step 103: Extract semantic features from the decoded information at multiple scales to generate multiple sets of semantic feature information at different scales.
[0044] In this embodiment, the scale of semantic features is a key concept in computer vision, especially in semantic segmentation tasks. The goal of semantic segmentation is to assign a semantic label to each pixel in an image, which requires the model not only to recognize objects in the image but also to accurately determine the boundaries of these objects. To achieve this goal, the model needs to process feature information at different scales.
[0045] In this embodiment, semantic feature information is the information obtained after extracting semantic features from the decoded information at different scales. For example, semantic feature information at different scales refers to semantic feature images at different scales.
[0046] Step 104: Based on semantic feature information, obtain the panoramic image of the corresponding image to be stitched.
[0047] In this embodiment, the panoramic image is the image obtained by stitching together at least two images to be stitched together.
[0048] In this embodiment, semantic feature information at different scales can describe the semantic information of the images to be stitched within different ranges. By processing the semantic feature information at different scales using a diffusion model for image stitching, panoramic images of at least two images to be stitched can be obtained. The diffusion model can be a diffusion model within a stitching model trained using the stitching model training method disclosed herein.
[0049] In this embodiment, the execution entity on which the semantic communication method runs can be, for example, Figure 2The receiver shown obtains semantic extraction information through a wireless channel, processes the semantic extraction information, and outputs a panoramic image Q.
[0050] The semantic communication method provided in this embodiment can be used to generate VR panoramic images. The receiving end directly extracts semantic information from at least two original images to be stitched by the panoramic camera and transmits them. The receiving end uses a diffusion model and semantic feature information to directly generate the stitched panoramic image, thereby integrating the panoramic image stitching process and the transmission process, reducing the amount of data transmitted during the communication process and improving end-to-end transmission efficiency.
[0051] The semantic communication method provided in this disclosure firstly receives semantic extraction information from at least two images to be stitched via a receiving channel; secondly, it performs received signal processing on the semantic extraction information to obtain decoded information; thirdly, it performs multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and finally, based on the semantic feature information, it obtains a panoramic image of the corresponding images to be stitched. Thus, by processing the semantic extraction information at the receiving end to generate a panoramic image of at least two images to be stitched, the image stitching is placed at the receiving end, and the channel only needs to transmit semantic extraction information, reducing the channel data transmission pressure and improving the transmission efficiency of panoramic image data; by using multiple sets of semantic feature information at different scales to obtain a panoramic image, the reliability of obtaining the panoramic image is improved.
[0052] In some optional implementations of this disclosure, the above-mentioned multi-scale semantic feature extraction of decoded information to generate multiple sets of semantic feature information at different scales includes: inputting the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer.
[0053] In this optional implementation method, such as Figure 2 As shown, the multi-scale semantic feature extractor includes at least two feature extractors. Therefore, the semantic feature information input to the diffusion model is multi-scale feature information.
[0054] In this optional implementation, the feature extractor generates semantic feature information based on the decoded information. The core idea of the residual network is to solve the training problem of deep networks by introducing residual blocks. Each residual block contains one or more convolutional layers and a skip connection, which directly adds the input to the output of the convolutional layer.
[0055] In this optional implementation, the transposed convolutional network is a type of deconvolutional network. The transposed convolutional network expands the size of the feature maps by inserting some padding values between the input feature maps, and image upsampling can be achieved through the transposed convolutional network.
[0056] In this optional implementation, the self-attention layer is a network layer that helps the model understand complex relationships in the image. The self-attention layer can help the feature extractor understand the distorted image captured by the panoramic camera.
[0057] In this optional implementation, the number and size of the semantic feature information extracted by each feature extractor in the multi-scale semantic feature extractor can be set based on development requirements.
[0058] The optional implementation provides a method for generating multiple sets of semantic feature information at different scales. By using a multi-scale semantic feature extractor to extract semantic feature information at multiple scales from the decoded information, the method can fully consider semantic feature information at different scales when generating panoramic images, thereby improving the reliability of panoramic image generation.
[0059] when Figure 2 The system shown processes VR data as follows: At least two images to be stitched are captured by a VR dual-fisheye camera. Initial semantic information is extracted from these images by a semantic information extractor. This initial semantic information is then encoded by an encoder into noise-robust coded information. The coded information is then quantized into finite-valued quantized information by a quantizer. The quantized information is modulated by an adjuster and transmitted wirelessly to the receiver, where it is demodulated to obtain demodulated information. The received and demodulated information is then decoded by a decoder to obtain decoded information. This decoded information is then input into a multi-scale semantic feature extractor to extract multiple sets of semantic features at different scales. Finally, the semantic features and the noisy image are input into a diffusion model to generate a panoramic image of the stitched image.
[0060] In some optional implementations of this disclosure, obtaining the panoramic image corresponding to the image to be stitched based on semantic feature information includes: inputting semantic feature information and random noise image into a trained diffusion model to obtain the stitched image output by the diffusion model, wherein the panoramic image is the image after stitching the image to be stitched.
[0061] In this optional implementation, the diffusion model can use the UNET architecture model, and the four sets of semantic feature information of different scales are passed to the co-diffusion model and connected to the corresponding scale layers of UNET.
[0062] In this optional implementation, the random noisy image is obtained by adding random image noise information to the image. The diffusion model is a model for stitching images. Inputting semantic feature information reflecting the semantic features of the panoramic image into the diffusion model allows the diffusion model to use the semantic feature information at different scales as input feature conditions for the panoramic image to be generated. Inputting the noisy image and input feature conditions into the diffusion model allows the diffusion model to denoise the noisy image; specifically, as follows... Figure 2 As shown, the diffusion model receives a noisy image ( Figure 2 (not shown in the image) and semantic feature information output by the multi-scale semantic feature extractor, outputting a panoramic image Q.
[0063] The optional implementation provides a method for obtaining panoramic images by inputting semantic feature information and random noisy images into a trained diffusion model. This allows the diffusion model to fully learn semantic feature information at different scales, reconstructing a panoramic image of the images to be stitched together, thus improving the accuracy and reliability of obtaining panoramic images.
[0064] In some optional implementations of this disclosure, the above-mentioned receiving signal processing of semantically extracted information to obtain decoded information includes: demodulating the semantically extracted information to obtain demodulated information; and decoding the demodulated information to obtain decoded information.
[0065] In this optional implementation, demodulation is the process of recovering the message from the modulated signal carrying information. In various information transmission or processing systems, the sending end modulates a carrier wave with the message to be transmitted, generating a signal carrying that message. The execution entity running on the semantic communication method recovers the transmitted message through demodulation.
[0066] In this optional implementation, the demodulation process varies depending on the modulation method, and includes sine wave amplitude demodulation, sine wave angle demodulation, pulse amplitude demodulation, and pulse phase demodulation. Sine wave demodulation can be further divided into amplitude demodulation, frequency demodulation, and phase demodulation. In addition, there are some variations such as single-sideband signal demodulation and vestigial sideband signal demodulation.
[0067] In this optional implementation, the demodulation information is the information obtained after demodulating the semantically extracted information, and it carries the semantic information of the images to be stitched together. The decoding information is the information obtained after decoding the demodulation information, and it carries the semantic information of the images to be stitched together.
[0068] In this optional implementation, decoding is the process of converting electrical pulse signals, optical signals, radio waves, etc., into the information and data they represent. Decoding is the process by which the receiver restores the received symbols or codes back to the original information.
[0069] In this optional implementation, the demodulation method for demodulating the semantically extracted information can be determined based on the modulation method of the transmitting end. The decoding method for decoding the demodulated information can be determined based on the encoding method of the transmitting end. Specifically, it can be determined through methods such as... Figure 2 The demodulator shown demodulates the semantically extracted information, through methods such as... Figure 2 The decoder shown implements the decoding process for demodulated information.
[0070] The optional implementation provides a method for obtaining decoded information by demodulating the semantically extracted information to obtain demodulated information; and then decoding the demodulated information to obtain decoded information. This effectively restores the semantically extracted information to the initial semantic information, is simple to operate, and improves the reliability of obtaining decoded information.
[0071] Optionally, the above-mentioned processing of the received signal to obtain decoded information from the semantically extracted information further includes: denoising the decoded information.
[0072] Optionally, the above-mentioned processing of the received signal to obtain decoded information from the semantically extracted information includes: amplifying the decoded information.
[0073] Figure 3 A flow 300 is shown according to another embodiment of the semantic communication method of this disclosure, the semantic communication method comprising the following steps:
[0074] Step 301: Receive at least two images to be stitched together.
[0075] In this embodiment, the at least two images to be stitched together can be images obtained by capturing an entity in the same scene using multiple different camera devices. Any two adjacent images to be stitched together may or may not have the same display area. For example, Figure 2 In this context, at least two images D to be stitched together include two images to be stitched together, which have the same display portion between them.
[0076] In this embodiment, the execution entity on which the semantic communication method of this disclosure runs can be the sending end. For example... Figure 2 As shown, the execution unit includes a semantic information extractor, an encoder, a quantizer, and a modulator.
[0077] Step 302: Input the image to be stitched into the semantic information extractor to obtain the initial semantic information output by the semantic information extractor.
[0078] In this embodiment, the semantic information extractor has the ability to extract the semantic features of an image. The initial semantic information is information that reflects the semantic features of the images to be stitched together, and the initial semantic information is information that has not undergone signal processing.
[0079] In this embodiment, the semantic information extractor includes four sets of residual networks and four downsampled convolutional networks. When the image to be stitched is input into the semantic information extractor, the residual network is processed first, and then the downsampled convolutional network is processed to obtain the initial semantic information.
[0080] Step 303: Process the initial semantic information by sending a signal to obtain semantic extraction information.
[0081] In this embodiment, the transmission signal processing refers to the signal processing of the data (initial semantic information) to be transmitted in the channel. The transmission signal processing includes: encoding processing, modulation processing, signal amplification processing, quantization processing, etc.
[0082] In this embodiment, step 303 includes: encoding the initial semantic information to obtain encoded information; and modulating the encoded information to obtain semantic extraction information. Encoding processing refers to the process of converting information from one form or format to another, which can be achieved through methods such as... Figure 2 The encoder shown implements information encoding processing. Modulation processing refers to the process of processing the information from the signal source and adding it to the carrier wave, making it suitable for channel transmission. It's the technique of changing the carrier wave according to the signal. Modulation methods can be divided into analog modulation and digital modulation according to the form of the modulating signal. This can be achieved through methods such as... Figure 2 The modulator shown modulates the information.
[0083] Step 304: Send the semantically extracted information to the receiving end through the channel.
[0084] In this embodiment, the receiving end performs multi-scale semantic feature extraction on the semantic extraction information, and obtains a panoramic image of the image to be stitched based on the extracted semantic feature information.
[0085] In this embodiment, the channel can be a wired channel or a wireless channel. For specific processing of semantic extraction information at the receiving end, please refer to [reference needed]. Figure 1 An example of a semantic communication method is shown.
[0086] The semantic communication method provided in this embodiment receives at least two images to be stitched; inputs the images to be stitched into a semantic information extractor to obtain initial semantic information output by the semantic information extractor; processes the initial semantic information for transmission signals to obtain semantic extraction information; and sends the semantic extraction information to the receiving end through a channel so that the receiving end can perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the images to be stitched based on the extracted semantic feature information. Thus, when a panoramic image of at least two images to be stitched is obtained, semantic information is extracted from at least two images to be stitched at the sending end, and transmission signal processing is performed. The obtained semantic extraction information is then sent to the receiving end through a channel. Compared with traditional image stitching methods, this method saves signal acquisition steps and reduces the information transmission pressure on the channel.
[0087] In some optional implementations of this disclosure, the above-mentioned processing of the initial semantic information to obtain semantic extraction information includes: encoding the initial semantic information to obtain encoded information; quantizing the encoded information to obtain quantized information; and modulating the quantized information to obtain semantic extraction information.
[0088] In this optional implementation, quantization refers to the process of approximating a signal's continuous values (or a large number of possible discrete values) to a finite number of discrete values. Quantization is mainly used in the conversion from continuous signals to digital signals. Continuous signals are sampled to become discrete signals, and discrete signals are quantized to become digital signals. Given the continuous nature of image data, quantization of the encoded information can reduce the amount of encoded data.
[0089] In this optional implementation, the encoded information is the information obtained by encoding the initial semantic information, and can be obtained by methods such as... Figure 2 The encoder shown encodes the initial semantic information. Quantization information is the information obtained after quantizing the encoded information, and can be obtained using methods such as... Figure 2 The quantizer shown quantizes the encoded information to obtain quantized information. This can be achieved using methods such as... Figure 2 The modulator shown modulates the quantized information to obtain semantically extracted information.
[0090] In this optional implementation, the encoder includes a set of residual networks and a downsampling convolution layer, the decoder includes a set of residual networks and a transposed convolution layer, and the quantizer can be selected to quantize to an integer from -7 to 8. The quantizer can quantize to an integer of the corresponding base according to the size of the channel resources and the required generation quality. The larger the base used, the more channel resources are used, and the better the generation quality.
[0091] The optional implementation provides a method for obtaining semantic extraction information by encoding the initial semantic information to obtain encoded information; quantizing the encoded information to obtain quantized information; and adjusting the quantized information to obtain semantic extraction information. This provides a reliable method for transmitting the initial semantic information in the channel and improves the reliability of semantic extraction information.
[0092] This disclosure proposes another embodiment of the semantic communication method, wherein: the sending end obtains semantic extraction information of at least two images to be stitched; the sending end sends the semantic extraction information to the receiving end through a channel; the receiving end receives the semantic extraction information; the receiving end performs received signal processing on the semantic extraction information to obtain decoded information; the receiving end performs multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and the receiving end obtains a panoramic image of the corresponding images to be stitched based on the semantic feature information.
[0093] Specifically, such as Figure 2 As shown, the transmitting end captures at least two images D to be stitched using multiple different cameras. At the transmitting end, an initial semantic information extractor extracts initial semantic information from the images D. This initial semantic information is then encoded into noise-robust coded information by an encoder. The coded information is quantized into finite-valued quantized information by a quantizer. This quantized information is then modulated by an adjuster to obtain semantically extracted information. This semantically extracted information is transmitted wirelessly to the receiving end. The receiving end demodulates the received semantically extracted information using a demodulator to obtain demodulated information. The receiving end then decodes the demodulated information using a decoder to obtain decoded information. The receiving end further extracts multiple sets of semantic features at different scales using a multi-scale semantic feature extractor. Finally, the semantic feature information and the noisy image are input into a diffusion model to generate a panoramic image Q of the stitched image.
[0094] The semantic communication method disclosed herein involves the following steps: The sending end obtains semantic extraction information from at least two images to be stitched; the sending end transmits the semantic extraction information to the receiving end via a channel; the receiving end receives the semantic extraction information; the receiving end performs received signal processing on the semantic extraction information to obtain decoded information; the receiving end performs multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and the receiving end obtains a panoramic image of the corresponding images to be stitched based on the semantic feature information. Thus, by directly extracting semantic extraction information from the images to be stitched captured by the camera at the sending end, and directly generating the stitched panoramic image at the receiving end using a diffusion generation model, the panoramic image stitching process and the information transmission process are integrated. This eliminates the need for direct image stitching at either the sending or receiving end, reducing the amount of data transmitted during communication and improving end-to-end transmission efficiency.
[0095] Figure 4A flowchart 400 is shown as an embodiment of the splicing model training method according to the present disclosure, which includes the following steps:
[0096] Step 401: Select image samples from the image sample set.
[0097] In this embodiment, the image sample set may include at least one image sample. Each image sample includes at least two sample images and a captioned image obtained by stitching the sample images together. The captioned image is a panoramic image obtained by stitching the sample images together. The sample images in each image sample may be images of objects in the same scene, and the sample images in each image sample may represent the same display portion.
[0098] The collection, storage, use, processing, transmission, provision, and disclosure of the image sample sets involved in this technical solution are performed after authorization and comply with relevant laws and regulations.
[0099] Step 402: Input the image sample into the semantic information extractor in the initial stitching model to obtain the initial semantic information.
[0100] In this embodiment, the semantic information extractor is used to extract the semantic information of the sample image in the image sample to obtain the initial semantic information, which is the initial semantic information directly extracted from the image sample.
[0101] In this embodiment, the semantic information extractor may include four sets of residual networks and four downsampled convolutional networks. The semantic information extractor can effectively obtain the initial semantic information of the sample images in the image samples.
[0102] Step 403: Process the initial semantic information by sending a signal to obtain semantic extraction information.
[0103] In this embodiment, signal processing refers to signal processing of the initial semantic information, which is the data to be transmitted in the channel. Signal processing includes encoding, modulation, signal amplification, quantization, etc.
[0104] Step 404: Add noise to the semantically extracted information and perform received signal processing on the semantically extracted information to obtain decoded information.
[0105] In this embodiment, the purpose of adding noise to the semantically extracted information is to simulate the transmission of the semantically extracted information in the channel. Specifically, noise with a preset signal-to-noise ratio can be added. When the splicing model is trained end-to-end, the wireless channel part will be replaced by the channel layer. The channel layer will add random noise of corresponding intensity to the transmitted information according to the set signal-to-noise ratio to simulate the real channel, so that the whole system has noise robustness. For example, the signal-to-noise ratio used by the splicing model during training is 10dB.
[0106] In this embodiment, received signal processing refers to signal processing of semantically extracted information with added noise. Received signal processing may include decoding, demodulation, signal amplification, etc.
[0107] In this embodiment, the received signal processing is based on the processing of the transmitted signal by the semantic extraction information sending end.
[0108] Step 405: Use the multi-scale semantic feature extractor in the splicing model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales.
[0109] In this embodiment, semantic feature information is the information obtained after extracting semantic features from the decoded information at different scales. For example, semantic feature information at different scales refers to semantic feature images at different scales.
[0110] Step 406: The diffusion model in the stitching model is used to process the semantic feature information and image samples to obtain the panoramic image of the corresponding image sample.
[0111] In this embodiment, the panoramic image corresponding to the image sample refers to the panoramic image obtained after the stitching model predicts the image sample. Compared with the labeled image of the sample image in the image sample, the panoramic image of the corresponding image sample may have a certain error. By adjusting the loss value of the stitching model during the training process, the error between the panoramic image of the corresponding image sample and the labeled image can be minimized. The loss value is calculated by the loss function.
[0112] In this embodiment, step 406 includes: adding Gaussian white noise to the labeled image in the image sample to obtain a noisy image; inputting the noisy image and semantic feature information into the diffusion model in the stitching model to obtain a panoramic image of the corresponding image sample output by the diffusion model.
[0113] Step 407: Calculate the loss of the stitching model based on the panoramic image.
[0114] In this embodiment, the above-mentioned calculation of the loss of the stitching model based on the panoramic image of the image sample includes: substituting the panoramic image and the labeled image of the image sample into the loss function to obtain the loss of the stitching model in each iteration of training.
[0115] In this embodiment, an end-to-end training method is used to train the stitching model, d p Let represent the panoramic image of the corresponding image sample. The stitching model will be trained with the goal of removing noise from the panoramic image, and will iterate a preset number of times (e.g., 200 times). The loss function of the stitching model is shown in Equation (1):
[0116] LOSS=(1-α)d1(x0,d p )+αd2(x0,d p (1)
[0117] In equation (1), d i Let x represent the i-th order Euclidean distance, x0 represent the output of the stitching model after a preset number of iterations, i.e., the labeled image, and d p This represents the panoramic image of the corresponding image sample. α is a weighting factor, the value of which can be set based on the training requirements of the stitching model. For example, α is 0.2.
[0118] Step 408: Based on the loss, determine if the stitching model meets the training completion conditions, and obtain the trained stitching model.
[0119] In this embodiment, the execution entity running on the stitching model can select image samples from step 401 and perform the training steps from 402 to 407 to complete one iteration of training of the stitching model. The selection method and number of image samples from the image sample set are not limited in this application, nor are the number of iterations of training the stitching model limited.
[0120] In this embodiment, the loss of the splicing model can be used to detect whether the splicing model meets the training completion condition. After the splicing model meets the training completion condition, the trained splicing model is obtained.
[0121] In this embodiment, the training completion condition includes: the loss of the splicing model is less than a first loss threshold. The first loss threshold can be determined based on specific training requirements; for example, the first loss threshold may be 0.01.
[0122] Optionally, in this embodiment, in response to the stitching model not meeting the training completion condition, the relevant parameters in the stitching model are adjusted to make the network loss value of the stitching model converge, and the above training steps 401-407 are continued based on the adjusted stitching model.
[0123] The stitching model training method provided in this embodiment selects image samples from an image sample set; inputs the image samples into the semantic information extractor in the initial stitching model to obtain initial semantic information; processes the initial semantic information for transmission signals to obtain semantic extraction information; adds noise to the semantic extraction information and processes the semantic extraction information for reception signals to obtain decoded information; uses a multi-scale semantic feature extractor in the stitching model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales; uses a diffusion model in the stitching model to process the semantic feature information and image samples to obtain panoramic images of the corresponding image samples; calculates the loss of the stitching model based on the panoramic image; and determines that the stitching model meets the training completion conditions based on the loss, thus obtaining a trained stitching model. Therefore, by using the semantic information extractor, multi-scale semantic feature extractor, diffusion model, and by adding noise to simulate a channel, semantic features are extracted from the unstitched original image captured by the panoramic camera at the transmitting end, and then processed at the receiving end and directly used to guide the diffusion model in generating the stitched panoramic image. End-to-end communication is achieved by integrating the stitching process into the transmission process.
[0124] Optionally, after the stitching model has been fully trained, it is input with a dual fisheye image, and the image input to the diffusion model is replaced with a random Gaussian noise map. The diffusion model is then iterated four times. The models at the transmitting and receiving ends are deployed separately and transmitted via a real wireless channel. After four iterations at the receiving end, the model generates a panoramic image corresponding to the input dual fisheye image.
[0125] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a semantic communication device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0126] like Figure 5 As shown, the semantic communication device 500 provided in this embodiment includes: an information receiving unit 501, a processing unit 502, a generation unit 503, and an image acquisition unit 504. The information receiving unit 501 can be configured to receive semantic extraction information from at least two images to be stitched together, transmitted via a channel. The processing unit 502 can be configured to process the received signal of the semantic extraction information to obtain decoded information. The generation unit 503 can be configured to extract multi-scale semantic features from the decoded information to generate multiple sets of semantic feature information at different scales. The image acquisition unit 504 can be configured to obtain a panoramic image of the corresponding images to be stitched together based on the semantic feature information.
[0127] In this embodiment, the specific processing of the information receiving unit 501, processing unit 502, generation unit 503, and image acquisition unit 504 in the semantic communication device 500, and the resulting technical effects, can be found in references to [reference needed]. Figure 1 The relevant descriptions of steps 101, 102, 103, and 104 in the corresponding embodiments will not be repeated here.
[0128] In some optional implementations of this embodiment, the generation unit 503 is configured to: input the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information of different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer.
[0129] In some optional implementations of this embodiment, the image obtaining unit 504 is configured to: input semantic feature information and random noise image into the trained diffusion model to obtain the stitched image output by the diffusion model, wherein the panoramic image is the image after stitching the image.
[0130] In some optional implementations of this embodiment, the processing unit 502 is configured to: demodulate the semantic extraction information to obtain demodulated information; and decode the demodulated information to obtain decoded information.
[0131] The semantic communication method provided in the embodiments of this disclosure firstly involves an information receiving unit 501 receiving semantic extraction information from at least two images to be stitched via a channel; secondly, a processing unit 502 performing received signal processing on the semantic extraction information to obtain decoded information; thirdly, a generation unit 503 performing multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales; and finally, an image obtaining unit 504 obtaining a panoramic image of the corresponding images to be stitched based on the semantic feature information. Thus, by processing the semantic extraction information at the receiving end to generate a panoramic image of at least two images to be stitched, the image stitching is placed at the receiving end, and the channel only needs to transmit semantic extraction information, reducing the channel data transmission pressure and improving the transmission efficiency of panoramic image data; the use of multiple sets of semantic feature information at different scales to obtain a panoramic image improves the reliability of the obtained panoramic image.
[0132] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0133] Further reference Figure 6As an implementation of the methods shown in the above figures, this disclosure provides another embodiment of a semantic communication device, which is similar to... Figure 3 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0134] like Figure 6 As shown, the semantic communication device 600 provided in this embodiment includes: an image receiving unit 601, an input unit 602, an information obtaining unit 603, and a sending unit 604. The image receiving unit 601 can be configured to receive at least two images to be stitched together. The input unit 602 can be configured to input the images to be stitched together into a semantic information extractor to obtain initial semantic information output by the semantic information extractor. The information obtaining unit 603 can be configured to perform transmission signal processing on the initial semantic information to obtain semantic extraction information. The sending unit 604 can be configured to send the semantic extraction information to a receiving end through a channel, so that the receiving end can perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the images to be stitched together based on the extracted semantic feature information.
[0135] In this embodiment, the specific processing of the image receiving unit 601, the input unit 602, the information obtaining unit 603, and the sending unit 604 in the semantic communication device 600, and the resulting technical effects, can be found in reference to [reference needed]. Figure 3 The relevant descriptions of steps 301, 302, 303, and 304 in the corresponding embodiments will not be repeated here.
[0136] In some optional implementations of this embodiment, the information obtaining unit 603 is configured to: encode the initial semantic information to obtain encoded information; quantize the encoded information to obtain quantized information; and modulate the quantized information to obtain semantic extraction information.
[0137] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides another embodiment of a semantic communication device, which is similar to... Figure 4 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0138] like Figure 7As shown, the semantic communication device 700 provided in this embodiment includes: a selection unit 701, a training unit 702, a calculation unit 703, and a model acquisition unit 704. The selection unit 701 can be configured to select image samples from an image sample set. The training unit 702 can be configured to input the image samples into the semantic information extractor in the initial stitching model to obtain initial semantic information; perform transmission signal processing on the initial semantic information to obtain semantic extraction information; add noise to the semantic extraction information and perform reception signal processing on the semantic extraction information to obtain decoded information; use a multi-scale semantic feature extractor in the stitching model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales; and use a diffusion model in the stitching model to process the semantic feature information and image samples to obtain a panoramic image of the corresponding image samples. The calculation unit 703 can be configured to calculate the loss of the stitching model based on the panoramic image. The model acquisition unit 704 can be configured to determine, based on the loss, that the stitching model meets the training completion conditions, and obtain the trained stitching model.
[0139] In this embodiment, the specific processing and technical effects of the semantic communication device 700, including the selection unit 701, training unit 702, calculation unit 703, and model acquisition unit 704, can be found in reference [reference needed]. Figure 4 The relevant descriptions of steps 401, 402, 403, and 404 in the corresponding embodiments will not be repeated here.
[0140] Further reference Figure 8 As an implementation of the methods shown in the above figures, this disclosure provides another embodiment of a semantic communication system, which is similar to... Figure 1 , Figure 3 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.
[0141] like Figure 8 As shown, the semantic communication system 800 provided in this embodiment includes: a transmitter 801, a channel 802, and a receiver 803. The transmitter 801 is used to receive at least two images to be stitched together; input the images to be stitched into a semantic information extractor to obtain initial semantic information output by the semantic information extractor; perform transmission signal processing on the initial semantic information to obtain semantic extraction information; and send the semantic extraction information to the receiver through the channel 802.
[0142] The aforementioned receiver 803 is used to receive semantic extraction information of at least two images to be stitched transmitted by channel 802; to process the received signal of the semantic extraction information to obtain decoding information; to extract multi-scale semantic features from the decoding information to generate multiple sets of semantic feature information at different scales; and to obtain a panoramic image of the corresponding images to be stitched based on the semantic feature information.
[0143] In this embodiment, the sending end 801 and the receiving end 803 can be either a client or a server.
[0144] In this embodiment, the specific processing of the sending end 801 in the semantic communication system 800 and the resulting technical effects can be referred to separately. Figure 1 For steps 101, 102, 103, and 104 in the corresponding embodiments, the specific processing of the receiving end 803 and its resulting technical effects can be found in the following references. Figure 3 The relevant descriptions of steps 301, 302, 303, and 304 in the corresponding embodiments will not be repeated here.
[0145] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0146] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0147] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded into random access memory (RAM) 903 from storage unit 908. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0148] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0149] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as semantic communication methods. For example, in some embodiments, the semantic communication method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the semantic communication method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform semantic communication methods by any other suitable means (e.g., by means of firmware).
[0150] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0151] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable semantic communication device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0152] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0154] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an information server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0155] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0156] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0157] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A semantic communication method, the method comprising: Semantic extraction information from at least two images to be stitched received from the receiving channel; The semantically extracted information is processed by receiving signals to obtain decoded information; Multi-scale semantic feature extraction is performed on the decoded information to generate multiple sets of semantic feature information at different scales; Based on the semantic feature information, a panoramic image corresponding to the image to be stitched is obtained; The step of extracting multi-scale semantic features from the decoded information to generate multiple sets of semantic feature information at different scales specifically includes: The decoded information is input into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer. The step of obtaining a panoramic image corresponding to the image to be stitched based on the semantic feature information specifically includes: The semantic feature information and random noisy images are input into the trained diffusion model to obtain the stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses the UNET architecture and four sets of semantic feature information at different scales are passed into the diffusion model and connected to the corresponding scale layers of UNET.
2. The method according to claim 1, wherein, The process of receiving signal processing from the semantically extracted information to obtain decoded information includes: The semantically extracted information is demodulated to obtain demodulated information; The demodulated information is decoded to obtain decoded information.
3. A semantic communication method, the method comprising: Receive at least two images to be stitched together; The image to be stitched is input into the semantic information extractor to obtain the initial semantic information output by the semantic information extractor; The initial semantic information is processed by sending signals to obtain semantic extraction information; The semantic extraction information is sent to the receiving end through a channel, so that the receiving end can perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the image to be stitched based on the extracted semantic feature information. The process of enabling the receiving end to perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the image to be stitched based on the extracted semantic feature information includes: inputting the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes a residual network, a transposed convolutional network, and a self-attention layer; inputting the semantic feature information and a random noisy image into a trained diffusion model to obtain a stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses a UNET architecture model, and four sets of semantic feature information at different scales are fed into the diffusion model and connected to the corresponding scale layers of the UNET.
4. The method according to claim 3, wherein processing the initial semantic information to obtain semantic extraction information includes: The initial semantic information is encoded to obtain encoded information; The encoded information is quantized to obtain quantized information; The quantized information is modulated to obtain semantic extraction information.
5. A semantic communication method, the method comprising: The sending end obtains semantic extraction information from at least two images to be stitched together; The sending end transmits the semantic extraction information to the receiving end through the channel; The receiving end receives the semantic extraction information; The receiving end processes the semantically extracted information to obtain decoded information. The receiving end performs multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales. Specifically, the multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales includes: inputting the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer. The receiving end obtains a panoramic image corresponding to the image to be stitched based on the semantic feature information. Specifically, obtaining the panoramic image corresponding to the image to be stitched based on the semantic feature information includes: inputting the semantic feature information and a random noise image into a trained diffusion model to obtain a stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses a UNET architecture model, and four sets of semantic feature information at different scales are passed to the diffusion model and connected to the corresponding scale layers of UNET.
6. A method for training a splicing model, the method comprising: Select image samples from the image sample set; The image sample is input into the semantic information extractor in the initial stitching model to obtain the initial semantic information; The initial semantic information is processed by transmitting signals to obtain semantic extraction information; noise is added to the semantic extraction information, and the semantic extraction information is processed by receiving signals to obtain decoded information; the multi-scale semantic feature extractor in the stitching model is used to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales; the diffusion model in the stitching model is used to process the semantic feature information and the image samples to obtain a panoramic image corresponding to the image samples. Based on the panoramic image, calculate the loss of the stitching model; The response determines that the stitching model meets the training completion condition based on the loss, and obtains the stitching model that has been trained. Specifically, the step of using the multi-scale semantic feature extractor in the splicing model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales includes: inputting the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer. The step of processing the semantic feature information and the image samples using the diffusion model in the stitching model to obtain the panoramic image corresponding to the image samples specifically includes: The semantic feature information and random noisy images are input into the trained diffusion model to obtain the stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses the UNET architecture and four sets of semantic feature information at different scales are passed into the diffusion model and connected to the corresponding scale layers of UNET.
7. A semantic communication device, the device comprising: The information receiving unit is configured to receive semantic extraction information from at least two images to be stitched together transmitted through the channel; The processing unit is configured to process the received signal of the semantic extraction information to obtain decoded information; The generation unit is configured to perform multi-scale semantic feature extraction on the decoded information to generate multiple sets of semantic feature information at different scales. Specifically, the generation unit is used to input the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer. The image acquisition unit is configured to obtain a panoramic image corresponding to the image to be stitched based on the semantic feature information. Specifically, the image acquisition unit is used to input the semantic feature information and random noise image into a trained diffusion model to obtain the stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses a UNET architecture model, and four sets of semantic feature information at different scales are passed to the diffusion model and connected to the corresponding scale layers of UNET.
8. A semantic communication device, the device comprising: The image receiving unit is configured to receive at least two images to be stitched together. The input unit is configured to input the image to be stitched into a semantic information extractor to obtain the initial semantic information output by the semantic information extractor; The information obtaining unit is configured to process the initial semantic information by sending a signal to obtain semantic extraction information; A sending unit is configured to send the semantic extraction information to a receiving end via a channel, so that the receiving end can perform multi-scale semantic feature extraction on the semantic extraction information and obtain a panoramic image of the image to be stitched based on the extracted semantic feature information. Specifically, the sending unit is used to input the decoded information into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information of different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes a residual network, a transposed convolutional network, and a self-attention layer. The sending unit also inputs the semantic feature information and a random noisy image into a trained diffusion model to obtain a stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses a UNET architecture model, and four sets of semantic feature information of different scales are passed into the diffusion model and connected to the corresponding scale layers of UNET.
9. A splicing model training device, the device comprising: The selection unit is configured to select image samples from the image sample set; The training unit is configured to input the image samples into the semantic information extractor in the initial stitching model to obtain initial semantic information; process the transmitted signal of the initial semantic information to obtain semantic extraction information; add noise to the semantic extraction information and process the received signal of the semantic extraction information to obtain decoded information; use the multi-scale semantic feature extractor in the stitching model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales; use the diffusion model in the stitching model to process the semantic feature information and the image samples to obtain a panoramic image corresponding to the image samples; wherein, the step of using the multi-scale semantic feature extractor in the stitching model to extract multi-scale semantic features from the decoded information to obtain multiple sets of semantic feature information at different scales specifically includes: processing the transmitted signal of the initial semantic information into the semantic extraction ... The code information is input into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes a residual network, a transposed convolutional network, and a self-attention layer. The step of using the diffusion model in the stitching model to process the semantic feature information and the image sample to obtain a panoramic image corresponding to the image sample specifically includes: inputting the semantic feature information and a random noisy image into the trained diffusion model to obtain a stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses a UNET architecture model, and four sets of semantic feature information at different scales are passed into the diffusion model and connected to the corresponding scale layers of UNET. A computing unit is configured to calculate the loss of the stitching model based on the panoramic image; The model obtains units, which are configured to respond based on the loss to determine that the stitched model meets the training completion condition, thus obtaining the trained stitched model.
10. A semantic communication system, the system comprising: Transmitter, channel, and receiver; The transmitting end is used to receive at least two images to be stitched together; The image to be stitched is input into a semantic information extractor to obtain initial semantic information output by the semantic information extractor; the initial semantic information is processed by a transmission signal to obtain semantic extraction information; the semantic extraction information is transmitted to the receiving end through the channel. The receiving end is used to receive semantic extraction information of at least two images to be stitched together transmitted through the channel; and to perform received signal processing on the semantic extraction information to obtain decoding information; Multi-scale semantic feature extraction is performed on the decoded information to generate multiple sets of semantic feature information at different scales; Based on the semantic feature information, a panoramic image corresponding to the image to be stitched is obtained; Specifically, the step of extracting multi-scale semantic features from the decoded information to generate multiple sets of semantic feature information at different scales includes: The decoded information is input into a pre-trained multi-scale semantic feature extractor to obtain multiple sets of semantic feature information at different scales output by the multi-scale conditional extractor. The multi-scale semantic feature extractor includes at least two feature extractors, each of which includes: a residual network, a transposed convolutional network, and a self-attention layer. The step of obtaining a panoramic image corresponding to the image to be stitched based on the semantic feature information specifically includes: The semantic feature information and random noisy images are input into the trained diffusion model to obtain the stitched image output by the diffusion model. The panoramic image is the image after stitching the image to be stitched. The diffusion model uses the UNET architecture and four sets of semantic feature information at different scales are passed into the diffusion model and connected to the corresponding scale layers of UNET.
11. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Diversity image restoration method based on multi-scale features and attention mechanism
CN117408920A
Semantic communication method based on multi-scale feature adaptive fusion
CN118585949A