Semantic Communication-Based Vehicle-Road Collaboration Image Information Transmission Method and Its Device

Through differential processing based on semantic communication and Swin Transformer architecture, combined with source channel joint encoding and decoding technology, the transmission efficiency problems of traditional communication systems under the shortage of bandwidth resources and changes in channel conditions are solved, and efficient and robust image transmission is achieved.

CN116844123BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310803014.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-07-29
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

The existing communication systems are limited by Shannon's limit when transmitting image data, and the transmission speed is slow and cannot meet the needs. The traditional source and channel independent encoding methods cannot dynamically adapt to changes in channel conditions, resulting in limited transmission performance in the case of shortage of bandwidth resources.

Method used

The vehicle-road collaborative image information transmission method based on semantic communication is adopted, and semantic analysis is performed through differential processing and Swin Transformer architecture to determine the weight value of the region of interest, and the source channel joint encoding and decoding technology is used to dynamically adjust the encoding method to adapt to different task requirements and channel conditions.

Benefits of technology

It improves the efficiency and robustness of image transmission, can maintain high-quality transmission performance when channel conditions change, avoids the "cliff effect" in traditional communication methods, and dynamically adjusts the compression ratio according to task requirements to optimize resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844123B_ABST
    Figure CN116844123B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle-road collaborative image information transmission method and device based on semantic communication, including: performing differential processing on the acquired original image, then dividing it into multiple units and combining them into a one-dimensional vector; determining the number of pixels in the region of interest in the unit and setting a weight value for the importance of the unit; performing semantic analysis on the unit including the pixels in the region of interest and extracting key semantic information to obtain the corresponding semantic space; encoding the one-dimensional vector according to the weight value of the unit to obtain a corresponding feature vector, transmitting it in a wireless channel, then using a source-channel joint decoder for decoding to obtain the decoded one-dimensional vector semantic information, and sequentially performing semantic fusion and semantic reconstruction to obtain the key semantic information for specific task requirements. The present invention realizes corresponding wireless encoding, transmission, and decoding according to different semantic tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and particularly relates to a vehicle-road collaborative image information transmission method based on semantic communication. Background Art

[0002] In an intelligent vehicle-road collaborative system, a moving intelligent vehicle not only has to complete a wide range of tasks but also needs to maintain information interaction with different participants in the traffic environment. According to statistics, more than 90% of the driving environment information can be obtained through images captured by cameras, making images the main carrier of traffic information. Furthermore, certain specific tasks can be completed using images, which is an important link to ensure the safety and reliability during the specific implementation of vehicle-road collaboration.

[0003] In the future 6G era, the physical layer dimension resources are all approaching a saturated state, and the shortage of bandwidth resources is a key problem faced by communication. Moreover, offloading a large number of complex computing tasks to edge servers will exacerbate the consumption of communication resources. In related technologies, based on Shannon's separation theorem, most current communication systems adopt a modular method of separate source-channel coding (SSCC) to transmit image data. However, traditional communication methods pursue lossless information transmission, but limited by the Shannon limit, their transmission volume is very limited, and the transmission speed (not fast) cannot meet the requirements, so using traditional communication methods is not very efficient.

[0004] Therefore, it is urgent to improve the defects existing in the prior art. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention provides a vehicle-road collaborative image information transmission method and device based on semantic communication. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] In a first aspect, the present invention provides a vehicle-road collaborative image information transmission method based on semantic communication, including:

[0007] Obtain an original image;

[0008] Perform differential processing on the original image, divide each differentially processed image into multiple units, and combine the multiple units of the same image into a one-dimensional vector, where some units in the one-dimensional vector include pixels of the region of interest and some units do not include pixels of the region of interest;

[0009] Facing the task requirements, determine the number of pixels of the region of interest included in the units in the one-dimensional vector, set weight values for the importance of the units in the one-dimensional vector, and guide semantic analysis, semantic information extraction, and source-channel joint coding operations;

[0010] Perform semantic analysis on each one-dimensional vector through the Swin Transformer architecture, extract key semantic information from the units including the pixels of the region of interest, assign different attentions to the semantic information according to the magnitude of the weight values of each unit, and obtain the corresponding semantic space;

[0011] Use a source-channel joint encoder to encode the semantic information in the semantic space, and guide the encoding process according to the magnitude of the weight values of each unit. When the difference in weight values of each unit is less than the first threshold, the same encoding operation is performed on each unit; when the difference in weight values is greater than the second threshold, a joint encoding operation is performed on the unit with a larger weight value to obtain the corresponding feature vector;

[0012] After transmitting the feature vector in the wireless channel, use a source-channel joint decoder to decode it to obtain the decoded semantic information;

[0013] Perform semantic fusion and semantic reconstruction on the decoded semantic information in sequence to obtain the key semantic information required for specific task requirements.

[0014] In a second aspect, the present application also provides a vehicle-road collaborative image information transmission device based on semantic communication, including:

[0015] A data acquisition module for acquiring the original image;

[0016] A preprocessing module for performing differential processing on the original image, dividing each image after differential processing into multiple units, and combining the multiple units of the same image into a one-dimensional vector, where some units in the one-dimensional vector include pixels of the region of interest and some units do not include pixels of the region of interest;

[0017] A semantic importance measurement module for determining, according to the task requirements, the number of pixels of the region of interest included in the units in the one-dimensional vector, setting weight values for the importance of the units in the one-dimensional vector, and guiding semantic analysis, semantic information extraction, and source-channel joint encoding operations;

[0018] A semantic analysis and semantic information extraction module for performing semantic analysis on each one-dimensional vector through the Swin Transformer architecture, extracting key semantic information from the units including the pixels of the region of interest, assigning different attentions to the semantic information according to the magnitude of the weight values of each unit, and obtaining the corresponding semantic space;

[0019] The source-channel joint coding module is used to encode the semantic information in the semantic space using a source-channel joint encoder, guide the coding process according to the magnitudes of the weight values of each unit, and when the difference in the weight values of each unit is less than the first threshold, perform the same coding operation on each unit; when the difference in the weight values is greater than the second threshold, perform a joint coding operation on the unit with a larger weight value to obtain the corresponding feature vector;

[0020] The source-channel joint decoding module is used to transmit the feature vector over a wireless channel and then decode it using a source-channel joint decoder to obtain the decoded semantic information;

[0021] The semantic fusion and reconstruction module is used to sequentially perform semantic fusion and semantic reconstruction on the decoded semantic information to obtain the key semantic information required for specific task requirements.

[0022] Advantages of the present invention:

[0023] A vehicle-road collaborative image information transmission method based on semantic communication provided by the present invention is experimentally analyzed using traffic scene images. First, the images are preprocessed. Secondly, a complete picture semantic transmission system model based on vehicle-road collaboration is given. Thirdly, the images are continuously trained to achieve good image transmission performance. In this embodiment, based on the method of semantic and source-channel joint coding and decoding, the corresponding wireless coding, transmission, and decoding methods can be organized according to the computing resources and communication resources available when performing different tasks.

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0025] Figure 1 is a flowchart of a vehicle-road collaborative image information transmission method based on semantic communication provided by an embodiment of the present invention;

[0026] Figure 2 is another flowchart of a vehicle-road collaborative image information transmission method based on semantic communication provided by an embodiment of the present invention. Detailed Embodiments

[0027] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0028] In the prior art, in an image transmission system, there are mainly three methods in terms of the compression processing and transmission method of the original image. First, the image is compressed by the JEPG compression method, and then a method independent of the source end is used to perform channel coding on the processed image; this method takes advantage of the insensitivity of the human eye to color and high-frequency information. On the premise of meeting the human visual requirements, some information is discarded accordingly to achieve the purpose of data compression. For example, downsampling, run-length encoding, and Huffman encoding techniques are used to further reduce data redundancy; however, such a compression method is relatively fixed and cannot dynamically adapt to the changes in the channel. The lossy compression method will reduce the data quality of the picture, and block effects will appear under high compression ratios, and key data information is easily lost. Second, the image is compressed by the JEPG2000 compression method, which improves some problems of JEPG based on wavelet transform, and then channel coding is performed in a manner independent of the information source; JEPG2000 supports lossy compression and lossless compression. In the lossy compression mode, JEPG2000 will not appear mosaic effects, and a larger compression ratio can also be obtained in the lossless compression mode; however, this compression method is relatively fixed and ignores the impact of changes in channel conditions on information transmission. In addition, there are also the inherent blurring and distortion problems of traditional compression methods. Third, the BPG compression method is used to achieve image compression. This is a new type of picture format that can be used when the file size and quality are limited, further reducing the spatial redundancy of the previous compression method; although this method overcomes some defects of the previous compression methods, there will also be inherent problems due to the compression technology used, such as common distortions such as color distortion and ringing effects. And based on the ideal condition assumption of the Shannon channel separation theory, the ideal state has not been achieved yet. Therefore, it can only be limited to using a fixed method to optimize the quality of image compression and restoration, without dynamically adapting to the channel, resulting in the "cliff effect" that appears in the source-channel separation coding.

[0029] In summary, although the above image compression and transmission schemes are widely used, when the channel conditions experienced are very different from the previously assumed situation, the performance of image transmission will be severely affected, that is, after assuming the channel state, the set image compression ratio cannot adapt to the changing channel conditions. Once there is a large fading in the channel, such lossy compression may cause a large amount of useful information to be lost. And the above methods only improve in terms of approaching the Shannon channel capacity infinitely, and the optimization space is very limited. In the current situation of scarce bandwidth resources, the spectrum resources occupied by traditional methods are also relatively fixed and inflexible. In addition, it is also impossible to perform controllable adjustment of the compression ratio according to the different downstream task requirements of the terminal.

[0030] In view of this, considering that complex and diverse tasks in the intelligent vehicle-road collaborative scenario require a large amount of data to complete, and specific tasks can still be completed even with slight data loss, the present invention provides a vehicle-road collaborative image information transmission method based on semantic communication, which has relatively low requirements for channel conditions during the image compression process, and the performance of the receiving end will not suddenly decline after the channel condition suddenly deteriorates, avoiding the "cliff effect" that appears in traditional communication methods.

[0031] Please refer to Figure 1 and Figure 2 as shown in Figure 1 is a flowchart of a vehicle-road collaborative image information transmission method based on semantic communication provided by an embodiment of the present invention. Figure 2 is another flowchart of a vehicle-road collaborative image information transmission method based on semantic communication provided by an embodiment of the present invention. A vehicle-road collaborative image information transmission method based on semantic communication provided by the present invention includes:

[0032] S101. Obtain the original image.

[0033] Specifically, in this embodiment, the original image is an image in the vehicle-road collaborative scenario.

[0034] S102. Perform differential processing on the original image, divide each image after differential processing into multiple units, and combine multiple units of the same image into a one-dimensional vector, where some units in the one-dimensional vector include pixels of the region of interest and some units do not include pixels of the region of interest.

[0035] Specifically, in this embodiment, preprocessing the obtained original image means performing a padding operation on the original image to facilitate subsequent semantic analysis, extraction, and encoding; in the preprocessing process, first perform differential processing on the original image, so that the changing objects in the original image are highlighted. For the objects in the vehicle-road collaborative scenario, using differential processing can make the pixels of the moving object main body be focused on later, and the attention of the unchanged background pixels is correspondingly reduced, that is, perform differential processing on each one-dimensional vector to obtain the pixels of the key moving objects in the unit, which are the pixels of the region of interest in the unit; then divide each image obtained by the difference into multiple units (patches) of a fixed size (patch-size×patch-size), and arrange the patches in an orderly manner to form multiple one-dimensional vectors.

[0036] S104. According to the task requirements, determine the number of pixels of the region of interest included in the units in the one-dimensional vector, set weight values for the importance of the units in the one-dimensional vector, and guide semantic analysis, semantic information extraction, and source-channel joint coding operations.

[0037] Specifically, in this embodiment, according to the task requirements, all units are assigned weight values, and the weight value range is from 0 to 1. It can be understood that the number of pixels of the region of interest included in the units in the one-dimensional vector is proportional to the weight value of the units in the one-dimensional image; the weight value range is from 0 to 1; if all pixels in some units are pixels of the region of interest, the weight value of this unit is set to 1; if some units do not have any pixels of the region of interest, the weight value of this unit is set to 0; and if some of the pixels in some units are pixels of the region of interest, the weight values of these units are set according to the ratio of the pixels of the region of interest to the pixels of the entire unit; through the screening method of assigning weight values to different units, the resource consumption in the image compression process can be effectively overcome; in this embodiment, a standard channel compression ratio is preset for the units with a weight value of 1, and for other units, according to the tasks to be actually processed, corresponding compression is performed according to the weight value ratio and the standard signal compression value. The range with a larger weight value has a larger compression value and a smaller compression degree to obtain a better recovery quality of the key part, and the range with a smaller weight value has a smaller compression value and a larger compression degree to reduce the useless consumption of bandwidth resources, and the resource utilization rate is improved by an unequal division method.

[0038] It should be noted that weight matching is performed for the region of interest, and the source-channel joint encoding and decoding are directly performed using the semantic information extracted from the region of interest. According to the number of pixels of the region of interest included in the pixels of each unit, the weight of each unit is allocated. When performing source-channel joint encoding and decoding, the attention mechanism is used to focus on and extract the semantic information in the units with high importance, such as key information such as cars, pedestrians, and traffic lights, to improve the transmission quality of this region. For the region of interest with lower importance, the compression degree is increased.

[0039] S103. Perform semantic analysis on each one-dimensional vector through the Swin Transformer architecture, extract key semantic information from the units including the pixels of the region of interest, and assign different attentions to the semantic information according to the weight values of each unit to obtain the corresponding semantic space.

[0040] Specifically, in this embodiment, first, a fixed window is set; secondly, the fixed window slides on the units including the pixels of the region of interest, and the attention mechanism of the sliding window is used to connect the information between adjacent windows to expand the receptive field range, and different attentions are allocated according to the size of the semantic weight value. Greater attention is given to the units with a large weight, and the attention of the units with a small weight is correspondingly reduced. Through the downsampling in each stage and the encoding effect of the Swin Transformer Block, the semantic space is obtained. It can be understood that using the Swin Transformer Block architecture can effectively implement the image compression process.

[0041] It should be noted that the process of differentiating the units in a one-dimensional vector is used to obtain the pixels of the key moving objects in the unit, which are the pixels of the region of interest (ROI) in the unit; then, the roadside device is used to segment the image after the differentiation process, that is, some of the units included in the image after the differentiation process include the pixels of the region of interest, some do not include the pixels of the region of interest, and the number of pixels of the region of interest included in each unit is also different.

[0042] It should be noted that in this embodiment, in the preprocessing process, the differential processing method is used to perform a differential operation on the image pixel points at two consecutive time points at the roadside end in the vehicle-road collaborative scenario, highlighting the changing and moving parts, and then giving greater attention to them using the attention mechanism in the subsequent process, while reducing the attention to the static and unchanged parts accordingly.

[0043] It should be noted that for the image transmission system for agents, different regions of interest are divided according to different communication tasks in the vehicle-road collaborative scenario, such as image classification, object detection, etc. Since the vehicle-road collaborative tasks focus on different objects in each frame of the image, the regions of interest extracted from the same image may not be exactly the same. For the images in the traffic scenario, they are divided into several parts such as people, vehicles, traffic lights, and the background, and these specific object contents are regarded as semantics. According to the particularity of the downstream tasks, different semantic information is extracted.

[0044] S105. Use a source-channel joint encoder to encode the semantic information in the semantic space, and guide the encoding process according to the magnitudes of the weight values of each unit. When the difference in the weight values of each unit is less than the first threshold, the same encoding operation is performed on each unit; when the difference in the weight values is greater than the second threshold, a joint encoding operation is performed on the unit with a larger weight value to obtain the corresponding feature vector.

[0045] Specifically, in this embodiment, the semantic importance weight is used to guide the encoding process of the source-channel joint encoder, and the weight value is input into the convolutional neural network for joint encoding; the weight value of the unit in the one-dimensional vector is related to the operation of the source-channel joint encoder. When the difference in the weight values of each unit is less than the first threshold, the same encoding operation is performed on each unit, that is, when the difference in the weight values of each unit is small, the same encoding operation is used; when the difference in the weight values of each unit is greater than the second threshold, a joint encoding operation is performed on the unit with a larger weight value, that is, when the difference in the weight values of each unit is large, a joint encoding is performed on the unit with a larger weight value. In this way, according to the magnitude of the weight value corresponding to the unit, the corresponding encoding method is used, and the image compression process can be effectively realized.

[0046] S106. After transmitting the feature vector in the wireless channel, use a joint source-channel decoder to decode it to obtain the decoded one-dimensional vector semantic information.

[0047] Specifically, in this embodiment, the feature vector is transmitted in the wireless channel. If OFDM modulation is used, resource blocks (RBs) are allocated for it, and it is specified which resource blocks are used for transmission. Information crucial for maintaining the source semantics is allocated to resource blocks with greater channel gains. In this mode, the channel transmission will have stronger anti-fading ability in a harsh wireless channel, and at the same time, the source semantic information processing will also have better robustness. After transmission in the wireless channel, the amplitude of the signal will be attenuated, and signals of different components will be accompanied by different degrees of distortion.

[0048] In this embodiment, at the receiving end, corresponding operations need to be performed on the information received from the wireless channel. The receiving end will use a joint source-channel decoder to recover the distorted signal, and use multi-layer deconvolution and Relu activation functions in the deep convolutional neural network for step-by-step upsampling to coarsely recover the distorted and distorted information.

[0049] S107. Perform semantic fusion and semantic reconstruction on the decoded one-dimensional vector semantic information in sequence to obtain the key semantic information required for specific task requirements.

[0050] Specifically, in this embodiment, the Swin Transformer framework is used to complete the operations of semantic recovery and reconstruction. Among them, semantic fusion is to more accurately recover semantic information, and semantic reconstruction is to recover specific semantic information according to task requirements. This operation process is opposite and symmetric to the structure of the semantic encoder; after predicting the corresponding result through the forward propagation process in the neural network, the predicted result and the original result are further processed and used as the loss function. Through the backpropagation process, the loss function is continuously trained and optimized. Specifically, the source-channel joint encoder will uniformly use the MSE loss function for joint training, with the goal of minimizing the semantic distance between the source and terminal information, and finally transmit the key semantic information required for specific tasks to the corresponding terminal.

[0051] It should be noted that the vehicle-road collaborative image information transmission method based on semantic communication in this embodiment realizes the functions of image processing and prediction with the assistance of an NVIDIA RTX4090 server.

[0052] Specifically, in this embodiment, traffic scene images are used for experimental analysis. First, the images are preprocessed. Second, a complete picture semantic transmission system model based on vehicle-road cooperation is given. Third, the images are continuously trained to achieve good image transmission performance. In this embodiment, based on the method of joint source-channel coding and decoding, the corresponding wireless coding, transmission, and decoding methods can be organized according to the computing resources and communication resources available when performing different tasks.

[0053] Based on the same inventive concept, the present invention also provides a vehicle-road cooperation image information transmission device based on semantic communication, which is applied to the vehicle-road cooperation image information transmission method based on semantic communication provided in the above embodiment. For the content of the method, please refer to the above embodiment, and the repeated parts will not be elaborated. The device includes:

[0054] A data acquisition module 201, configured to acquire an original image;

[0055] A preprocessing module 202, configured to perform differential processing on the original image, divide each image after differential processing into multiple units, and combine the multiple units of the same image into a one-dimensional vector. Among them, some units in the one-dimensional vector include pixels of the region of interest, and some units do not include pixels of the region of interest;

[0056] A semantic importance measurement module 203, configured to determine the number of pixels of the region of interest included in the units in the one-dimensional vector according to task requirements, set weight values for the importance of the units in the one-dimensional vector, and guide semantic analysis, semantic information extraction, and joint source-channel coding operations;

[0057] A semantic analysis and semantic information extraction module 204, configured to perform semantic analysis on each one-dimensional vector through the Swin Transformer architecture, extract key semantic information from the units including pixels of the region of interest, assign different attentions to the semantic information according to the magnitude of the weight values of each unit, and obtain a corresponding semantic space;

[0058] A joint source-channel coding module 205, configured to encode the semantic information in the semantic space using a joint source-channel encoder, guide the encoding process according to the magnitude of the weight values of each unit, and when the difference between the weight values of each unit is less than a first threshold, perform the same encoding operation on each unit; when the weight value difference is greater than a second threshold, perform a joint encoding operation on the unit with a large weight value to obtain a corresponding feature vector;

[0059] A joint source-channel decoding module 206, configured to transmit the feature vector in a wireless channel and then decode it using a joint source-channel decoder to obtain the decoded semantic information;

[0060] The semantic fusion and reconstruction module 207 is configured to perform semantic fusion and semantic reconstruction on the decoded semantic information in sequence to obtain the key semantic information for specific task requirements.

[0061] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements includes not only those elements but also other elements not expressly listed. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the article or device comprising said element. "Connection" or "coupling" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "up", "down", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention.

[0062] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0063] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A vehicle-road collaborative image information transmission method based on semantic communication, characterized in that, Including: Obtain the original image; Perform differential processing on the original image, divide each differentially processed image into multiple units, and combine the multiple units of the same image into a one-dimensional vector, where some units in the one-dimensional vector include pixels of the region of interest and some units do not include pixels of the region of interest; Facing the task requirements, determine the number of pixels of the region of interest included in the units in the one-dimensional vector, set weight values for the importance of the units in the one-dimensional vector, and guide semantic analysis, semantic information extraction, and source-channel joint coding operations; Perform semantic analysis on each one-dimensional vector through the Swin Transformer architecture, extract key semantic information from the units including pixels of the region of interest, assign different attentions to the semantic information according to the magnitude of the weight values of each unit, and obtain the corresponding semantic space; Use the source-channel joint encoder to encode the semantic information in the semantic space, guide the encoding process according to the magnitude of the weight values of each unit. When the difference in weight values of each unit is less than the first threshold, perform the same encoding operation on each unit; when the weight value difference is greater than the second threshold, perform the joint encoding operation on the unit with a larger weight value to obtain the corresponding feature vector; After transmitting the feature vector in the wireless channel, use the source-channel joint decoder to decode it to obtain the decoded semantic information; Perform semantic fusion and semantic reconstruction on the decoded semantic information in sequence to obtain the key semantic information for specific task requirements.

2. The vehicle-road collaborative image information transmission method based on semantic communication according to claim 1, wherein, The extracting key semantic information from the units including pixels of the region of interest, assigning different attentions to the semantic information according to the magnitude of the weight values of each unit, and obtaining the semantic space includes: Set a fixed window; Slide the fixed window on the units including pixels of the region of interest, and use the attention mechanism of the sliding window to connect the information between adjacent windows to obtain the corresponding semantic space.

3. The method for transmitting vehicle-road collaborative image information based on semantic communication according to claim 1, wherein The number of pixels of the region of interest included in the units in each one-dimensional vector is proportional to the weight value of the units in the one-dimensional vector.

4. The vehicle-road collaborative image information transmission method based on semantic communication according to claim 1, characterized in that, The weight value of the units in the one-dimensional vector is related to the operation of the source-channel joint encoder.

5. The method for vehicle-road collaborative image information transmission based on semantic communication according to claim 1, characterized in that, During the transmission of the feature vector in the wireless channel, use OFDM for transmission and allocate resource blocks, and allocate the information crucial for maintaining the source semantics to the resource blocks with larger signal gains.

6. The method for vehicle-road collaborative image information transmission based on semantic communication according to claim 1, wherein, The using the source-channel joint decoder to decode includes: Use the multi-layer deconvolution and Relu activation function in the deep convolutional neural network to recover the feature vector transmitted through the wireless channel to obtain the semantic information with distortion and aberration.

7. The method for transmitting vehicle-road cooperative image information based on semantic communication according to claim 1, characterized in that: The process of performing semantic fusion and semantic reconstruction on the decoded semantic information uses the Swin Transformer architecture.

8. An image information transmission device for vehicle-road cooperation based on semantic communication, characterized in that, Including: A data acquisition module for obtaining the original image; A preprocessing module for performing differential processing on the original image, dividing each differentially processed image into multiple units, and combining the multiple units of the same image into a one-dimensional vector, where some units in the one-dimensional vector include pixels of the region of interest and some units do not include pixels of the region of interest; A semantic importance measurement module, which is used to determine the number of pixels in the region of interest included in the units of the one-dimensional vector according to the task requirements, set weight values for the importance of the units in the one-dimensional vector, and guide semantic analysis, semantic information extraction, and source-channel joint coding operations; A semantic analysis and semantic information extraction module, which is used to perform semantic analysis on each one-dimensional vector through the Swin Transformer architecture, extract key semantic information from the units including pixels in the region of interest, assign different attentions to the semantic information according to the magnitude of the weight values of each unit, and obtain the corresponding semantic space; A source-channel joint coding module, which is used to encode the semantic information in the semantic space using a source-channel joint encoder, guide the coding process according to the magnitude of the weight values of each unit, and when the difference in the weight values of each unit is less than the first threshold, perform the same coding operation on each unit; when the difference in weight values is greater than the second threshold, perform a joint coding operation on the unit with a large weight value to obtain the corresponding feature vector; A source-channel joint decoding module, which is used to transmit the feature vector in the wireless channel and then decode it using a source-channel joint decoder to obtain the decoded semantic information; A semantic fusion and reconstruction module, which is used to sequentially perform semantic fusion and semantic reconstruction on the decoded semantic information to obtain the key semantic information for specific task requirements.