Camera data sending method, camera data receiving method, camera data transmission method, camera data transmission system and camera data transmission device
By introducing lightweight visual Transformer network and semantic encoder of large language models into the camera system, the transmission efficiency problem of traditional cameras in low bandwidth, low signal-to-noise ratio environments is solved, and low-cost and high-precision image transmission and recovery is achieved, suitable for scenarios such as intelligent monitoring and autonomous driving.
Patent Information
- Application Number
- CN202510786100.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional camera systems are difficult to efficiently transmit video image information in complex environments with limited bandwidth and low signal-to-noise ratio, resulting in high communication overhead, serious information redundancy and slow system response, especially in real-time decision-making scenarios.
The lightweight visual Transformer network is used to extract image semantic features, and the semantic encoder based on large language models is used for joint encoding, combining signal modulation and decoding with channel adaptive capabilities to achieve efficient semantic information transmission.
It significantly reduces the amount of data transmission, is suitable for low bandwidth and low signal-to-noise ratio environments, realizes low-cost and high-precision image recovery and intelligent analysis tasks, and improves system response speed.
Smart Images

Figure CN120455832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data transmission, and in particular to a camera data sending method, receiving method, transmission method, system and device. Background Art
[0002] As scenarios like intelligent surveillance, remote perception, and autonomous driving continue to demand ever-increasing data real-time and efficient transmission, traditional camera systems face increasing challenges. Efficiently transmitting video image information is a key issue, particularly in complex environments with limited bandwidth and low signal-to-noise ratios. Traditional cameras capture enormous amounts of video data, and transmission relies on high bandwidth and stable channels, resulting in high communication overhead and a sharp performance degradation in weak signal or congested networks.
[0003] In addition, in the traditional perception-communication decoupling architecture, the camera first completes image acquisition and processing, and then transmits the image or analysis results as a whole to the receiving end. This approach has some drawbacks:
[0004] (1) High transmission overhead: High-definition images and video frames have large data volumes and consume a large amount of bandwidth, resulting in a significant increase in transmission delay, making it difficult to meet the requirements of low-latency applications;
[0005] (2) Serious information redundancy: The original image contains a large amount of redundant data that is irrelevant to downstream tasks, and blind transmission causes a waste of resources;
[0006] (3) Slow system response: The serial process of perception and communication results in longer end-to-end processing time, especially in scenarios requiring real-time decision-making (such as fire warning and anomaly detection), where efficiency is difficult to guarantee. Summary of the Invention
[0007] The purpose of the present invention is to design a camera data sending method, receiving method, transmission method, system and device in order to solve the above problems.
[0008] The present invention achieves the above-mentioned purpose through the following technical solutions:
[0009] The camera data sending method includes:
[0010] S1, the camera acquires the image of the target area;
[0011] S2, extracting semantic features of the image;
[0012] S3, using a semantic encoder to jointly encode communication parameters and semantic features;
[0013] S4. Perform signal modulation on the semantic code to obtain a modulated complex signal , expressed as , represents the complex signal after semantic coding corresponding to the n-th frame image, H represents the channel gain between the user and the base station, and N represents additive white Gaussian noise;
[0014] S5, the modulated complex signal It is sent as a transmission signal over the channel to the receiving end.
[0015] The camera data receiving method includes:
[0016] (1) receiving a transmission signal sent by a transmitter on a channel, wherein the transmission signal is a complex signal modulated by the transmitter using the above-mentioned camera data transmission method for image processing;
[0017] (2) Use the semantic decoder to restore the original image information.
[0018] An efficient camera data transmission method, comprising:
[0019] 1) The sending end processes the image of the target area using the above-mentioned camera data to obtain a transmission signal, and sends the transmission signal to the receiving end;
[0020] 2) The receiving end receives the transmission signal and uses the above-mentioned camera data receiving method to restore the original image information of the transmission signal.
[0021] An efficient camera data transmission system, comprising:
[0022] The transmitting end is used to obtain an image of the target area, and use the above-mentioned efficient camera data transmission to process it, obtain a transmission signal, and send the transmission signal to the receiving end;
[0023] The receiving end is used to receive the transmission signal and restore the original image information of the transmission signal using the above-mentioned efficient camera data receiving method.
[0024] A device comprising:
[0025] Storage; the storage stores at least one instruction, at least one program, code set or instruction set;
[0026] Processor; the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement any one of the above-mentioned efficient camera data sending method, the above-mentioned efficient camera data receiving method and the above-mentioned one efficient camera data transmission method.
[0027] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement any one of the above-mentioned efficient camera data sending method, the above-mentioned efficient camera data receiving method, and the above-mentioned efficient camera data transmission method.
[0028] The beneficial effect of the present invention is that it combines semantic communication with camera hardware and software to achieve low-cost, high-precision data transmission. First, at the transmitting end (camera), a lightweight visual Transformer network is designed to extract semantic features from the image. Then, a semantic encoder based on a large language model (LLM) is proposed to jointly process the communication parameters of the text modality and the semantic features of the image to realize a semantic encoding process with channel adaptability. At the receiving end (such as the base station), a visual Transformer with a decoder-only architecture is used as a semantic decoder to achieve high-quality image restoration. This method provides key algorithmic support and theoretical foundation for the system implementation of the "semantic camera". BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of an efficient camera data transmission method of the present invention;
[0030] Figure 2 It is a schematic diagram of an efficient camera data transmission system of the present invention. DETAILED DESCRIPTION
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0032] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0033] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not require further definition or explanation in subsequent drawings.
[0034] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the inventive product is conventionally placed when in use, or are the orientations or positional relationships conventionally understood by those skilled in the art. These are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0035] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.
[0036] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, terms such as "disposed" and "connected" should be understood in a broad sense. For example, "connected" can mean a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can also mean internal communication between two components. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0037] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0038] The camera data sending method includes:
[0039] S1. The camera acquires the image of the target area .
[0040] S2. Use the lightweight visual Transformer network - BiFormer as the backbone network to extract images The semantic features of ; specifically include:
[0041] S21, the PatchEmbedding layer of the lightweight visual Transformer network will Divided into non-overlapping regions, each region contains feature vectors, expressed as: ; Apply the linear projection layer to obtain the query, key, and value tensors, denoted as Q, K, and , 、 、 ;in 、 and Represent the projection weights of query, key, and value respectively; by averaging the query and key tensors of each region, the region-level query and key are calculated, which are expressed as and ;
[0042] S22. Calculate the adjacency matrix , quantifying the correlation between different regions, expressed as: , where T represents the transpose operation;
[0043] S23, retain the first h feature connections of each region to prune the adjacency matrix and obtain the routing index matrix , expressed as ,in, Represents the operation function of selecting the first h row vectors with the largest value, The i-th row of contains the indices of the h regions that are most correlated with the i-th region;
[0044] S24, area-to-area routing index matrix , by collecting the corresponding key and value tensors, we can achieve fine-grained token-to-token attention, expressed as: 、 ,in, and are the collected key and value tensors;
[0045] S25. Apply the attention mechanism to the collected key and value tensors to obtain the output tensor , expressed as: 、 ,in, is the scaling factor, represents the local context enhancement constraint;
[0046] S26. Apply linear projection layer To the output tensor, get the semantic features of the image, expressed as: .
[0047] S3. Jointly encode the communication parameters and semantic features using a semantic encoder; specifically, including:
[0048] S31. Use GPT-2 word segmenter to adjust communication parameters Perform word segmentation and generate input word segmentation ID and the corresponding attention mask , using the GPT-2 word embedding layer Mapping the tokenized input to the embedding space is represented as: 、 ,in, is the length of the text embedding;
[0049] S32. Semantic features of images and text embedding Connect along the sequence dimension to form the fusion input, which is expressed as: 、 ,in, ;
[0050] S33. Constructing fused attention mask , expressed as: ,in Initialize to a matrix of all 1s;
[0051] S34, fusion input and fused attention mask Pass the pre-trained GPT-2 model to generate semantic encoding , expressed as: 、 .
[0052] S4. Perform signal modulation on the semantic code to obtain a modulated complex signal , expressed as , Indicates the The complex signal after semantic coding corresponding to the frame image, H represents the channel gain between the user and the base station, and N represents additive white Gaussian noise;
[0053] S5, the modulated complex signal It is sent as a transmission signal over the channel to the receiving end.
[0054] This method eliminates the need for cameras to transmit complete pixel-level images. Instead, it extracts and encodes high-level semantic information from the image (such as object category, location, and behavior). The receiving end (such as a base station) can use this semantic information to reconstruct the image content or perform corresponding intelligent analysis tasks. This method significantly reduces the amount of data required for transmission and is particularly suitable for demanding communication environments such as low bandwidth and low signal-to-noise ratios. It has strong practicality and potential for widespread adoption.
[0055] The camera data receiving method includes:
[0056] (1) Receive the transmission signal sent by the transmitter on the channel. The transmission signal is a complex signal modulated by the transmitter using the camera data transmission method on the image processing.
[0057] (2) Use the semantic decoder to restore the original image information; specifically including:
[0058] (201), encode the received semantics Reshape into a spatial representation compatible with convolution operations , expressed as: 、 ,in, The convolution operation rearranges it into The size of , is the channel dimension of the backbone network output;
[0059] (202) Using the ViT network with Decoder Only architecture to restore images and obtain spatial information ; Specifically: space representation Upsample and then add position embedding , thus preserving spatial information in the Transformer-based processing, expressed as: 、 ,in is the upsampling operation;
[0060] (203) The Transformer module in the ViT decoder performs global feature aggregation and maps the transformed features back to the pixel space, which is expressed as: ,in Represents multiple Transformer modules, Represents the operation of converting a patch into an image.
[0061] An efficient camera data transmission method, comprising:
[0062] 1) The sending end processes the image of the target area using the above-mentioned efficient camera data transmission method to obtain a transmission signal, and sends the transmission signal to the receiving end; specifically:
[0063] 101) The camera obtains an image of the target area;
[0064] 102) Use the lightweight visual Transformer network - BiFormer as the backbone network to extract images semantic features of
[0065] 103) Jointly encode communication parameters and semantic features using an LLM-based semantic encoder;
[0066] 104) Perform signal modulation on the semantic code to obtain a modulated complex signal , expressed as , Indicates the The complex signal after semantic coding corresponding to the frame image, H represents the channel gain between the user and the base station, and N represents additive white Gaussian noise;
[0067] 105) The modulated complex signal It is wirelessly transmitted on the channel as a transmission signal to the receiving end.
[0068] 2) The receiving end receives the transmission signal and uses the above-mentioned efficient camera data receiving method to restore the original image information of the transmission signal; specifically:
[0069] 201), receiving a transmission signal sent by a transmitter on a channel, wherein the transmission signal is a complex signal modulated by an image processing method using an efficient camera data transmission method on the transmitter.
[0070] 202) Use the semantic decoder to restore the original image information.
[0071] This framework integrates semantic communication with cameras through hardware and software to achieve low-cost, high-precision data transmission. First, a lightweight visual Transformer network is designed at the transmitting end (camera) to extract semantic features from images. A semantic encoder based on a large language model (LLM) is then proposed to jointly process communication parameters of the text modality and image semantic features, achieving a semantic encoding process with channel adaptability. At the receiving end (e.g., base station), a visual Transformer with a DecoderOnly architecture is used as a semantic decoder to achieve high-quality image restoration. This approach provides key algorithmic support and theoretical foundation for the system implementation of the "semantic camera."
[0072] An efficient camera data transmission system, comprising:
[0073] The transmitting end is used to obtain an image of the target area, and use the above-mentioned efficient camera data transmission method to process it to obtain a transmission signal, and send the transmission signal to the receiving end;
[0074] The receiving end is used to receive the transmission signal and restore the original image information of the transmission signal using the above-mentioned efficient camera data receiving method.
[0075] A device comprising:
[0076] Storage; the storage stores at least one instruction, at least one program, code set or instruction set;
[0077] Processor; the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement any one of the above-mentioned efficient camera data sending method, the above-mentioned efficient camera data receiving method and the above-mentioned one efficient camera data transmission method.
[0078] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement any one of the above-mentioned efficient camera data sending method, the above-mentioned efficient camera data receiving method, and the above-mentioned efficient camera data transmission method.
[0079] The technical solution of the present invention is not limited to the above-mentioned specific embodiments. Any technical variations made according to the technical solution of the present invention fall within the protection scope of the present invention.
Claims
1. A camera data sending method, characterized in that: include: S1, the camera acquires the image of the target area; S2, extracting semantic features of the image; S3, using a semantic encoder to jointly encode communication parameters and semantic features; S4. Perform signal modulation on the semantic code to obtain a modulated complex signal , expressed as , represents the complex signal after semantic coding corresponding to the n-th frame image, H represents the channel gain between the user and the base station, and N represents additive white Gaussian noise; S5, the modulated complex signal It is sent as a transmission signal over the channel to the receiving end.
2. The camera data sending method according to claim 1, wherein: In S2, a lightweight visual Transformer network, BiFormer, is used as the backbone network to capture images from the camera. Extract semantic features from Specifically include: S21, the PatchEmbedding layer of the lightweight visual Transformer network will Divided into non-overlapping regions, each containing feature vectors, expressed as: ; Apply the linear projection layer to obtain the query, key, and value tensors, denoted as Q, K, and , 、 、 ;in 、 and Represent the projection weights of query, key, and value respectively; by averaging the query and key tensors of each region, the region-level query and key are calculated, which are expressed as and ; S22. Calculate the adjacency matrix , quantifying the correlation between different regions, expressed as: , where T represents the transpose operation; S23, retain the first h feature connections of each region to prune the adjacency matrix and obtain the routing index matrix , expressed as ,in, Represents the operation function of selecting the first h row vectors with the largest value, The i-th row of contains the indices of the h regions that are most correlated with the i-th region; S24, area-to-area routing index matrix , by collecting the corresponding key and value tensors, we can achieve fine-grained token-to-token attention, expressed as: 、 ,in, and are the collected key and value tensors; S25. Apply the attention mechanism to the collected key and value tensors to obtain the output tensor , expressed as: 、 ,in, is the scaling factor, represents the local context enhancement constraint; S26. Apply linear projection layer To the output tensor, get the semantic features of the image, expressed as: .
3. The camera data sending method according to claim 1, wherein: Included in S3: S31. Use GPT-2 word segmenter to adjust communication parameters Perform word segmentation and generate input word segmentation ID and the corresponding attention mask , using the GPT-2 word embedding layer Mapping the tokenized input to the embedding space is represented as: 、 ,in, is the length of the text embedding; S32. Semantic features of images and text embedding Connect along the sequence dimension to form the fusion input, which is expressed as: 、 ,in, ; S33. Constructing fused attention mask , expressed as: ,in Initialize to a matrix of all 1s; S34, fusion input and fused attention mask Pass the pre-trained GPT-2 model to generate semantic encoding , expressed as: 、 .
4. A camera data receiving method, characterized in that: include: (1) receiving a transmission signal sent by a transmitter on a channel, wherein the transmission signal is a complex signal modulated by the transmitter using the camera data transmission method according to any one of claims 1 to 3 for image processing; (2) Use the semantic decoder to restore the original image information.
5. The camera data receiving method according to claim 4, wherein: In (2) include: (201), encode the received semantics Reshape into a spatial representation compatible with convolution operations , expressed as: 、 ,in, The convolution operation rearranges it into The size of , is the channel dimension of the backbone network output; (202) Using the ViT network with Decoder Only architecture to restore images and obtain spatial information ; (203) The Transformer module in the ViT decoder performs global feature aggregation and maps the transformed features back to the pixel space, which is expressed as: ,in Represents multiple Transformer modules, Represents the operation of converting a patch into an image.
6. The camera data receiving method according to claim 5, wherein: Specifically included in (202): the spatial representation Upsample and then add position embedding , expressed as: 、 ,in is the upsampling operation.
7. A camera data transmission method, characterized in that: include: 1) The transmitting end processes the image of the target area using the camera data transmission method according to any one of claims 1 to 3 to obtain a transmission signal, and sends the transmission signal to the receiving end; 2) The receiving end receives the transmission signal and uses the camera data receiving method described in any one of claims 4 to 6 to restore the original image information of the transmission signal.
8. A camera data transmission system, characterized in that: include: The transmitting end is used to obtain an image of the target area, and process the image using the camera data transmission method according to any one of claims 1-3 to obtain a transmission signal, and transmit the transmission signal to the receiving end; Receiving end; the receiving end is used to receive the transmission signal and restore the original image information of the transmission signal using the camera data receiving method described in any one of claims 4-6.
9. A device, characterized in that: include: Storage; The memory stores at least one instruction, at least one program, a code set or an instruction set; Processor; the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement any one of the camera data sending method according to any one of claims 1 to 3, the camera data receiving method according to any one of claims 4 to 6, and the camera data transmission method according to any one of claim 7.
10. A computer-readable storage medium, characterized in that The readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement any one of the camera data sending method according to any one of claims 1 to 3, the camera data receiving method according to any one of claims 4 to 6, and the camera data transmission method according to any one of claim 7.