Low-bandwidth-based end-side collaborative visual detection method and system
Through lightweight semantic segmentation and semantic database completion, the problem of missing and reconstruction distortion of target area information under low bandwidth conditions is solved, the detection accuracy and delay performance are improved, and it is suitable for visual detection in low bandwidth environments.
Patent Information
- Application Number
- CN202510649206.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-02
AI Technical Summary
The lack of semantic perception in existing low-bandwidth end-edge collaborative visual detection leads to the loss of target area information, reconstruction distortion and fluctuation in detection accuracy, especially in high-precision scenarios such as medical images and autonomous driving.
The lightweight semantic segmentation model is used to identify the target bounding box and divide it into semantic key blocks. Feature completion and texture filling are performed through high-reliable channel transmission and combined with the semantic database of the edge server, and the semantic enhancement target detection model is used to improve detection accuracy.
Under low bandwidth conditions, the target area reconstruction error is reduced by more than 40%, the detection accuracy is improved by 15-20%, and the delay is reduced by more than 60%, achieving an optimized balance between accuracy and delay.
Smart Images

Figure CN120580428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual inspection technology, and in particular to a method and system for low-bandwidth end-edge collaborative visual inspection. Background Art
[0002] In the field of IoT visual inspection, the edge-end collaborative architecture has become the mainstream solution due to the limited computing power of IoT devices. However, it faces the contradiction between low-bandwidth communication and detection accuracy. For example, in an existing published patent (publication number: CN118552490A), in low-power wide-area networks (LPWANs) such as LoRa, image transmission latency can be as long as several minutes, and high-frequency packet loss leads to severe degradation of image quality and a significant decrease in target detection accuracy. While this patent reduces latency by prioritizing the transmission of image blocks with dense feature points, this method lacks semantic perception and cannot distinguish between target areas and background (for example, key blocks containing vehicles may be misclassified as background blocks and discarded). In addition, the generative visual model reconstruction relies on a complete feature point distribution, which is prone to reconstruction distortion in feature-sparse scenes (such as low-texture targets and complex lighting), leading to fluctuating detection results. Furthermore, this patent does not introduce target semantic information to guide transmission and reconstruction, making it difficult to simultaneously ensure target area integrity and detection accuracy under low bandwidth. This makes it particularly unsuitable for high-precision scenarios such as medical imaging and autonomous driving. Summary of the Invention
[0003] Technical problems solved
[0004] In response to the shortcomings of the existing technology, the present invention provides a method and system based on low-bandwidth end-edge collaborative visual detection, which solves the problems of target area information loss, reconstruction distortion and detection accuracy fluctuation caused by the traditional scheme in existing low-bandwidth end-edge collaborative visual detection due to the lack of semantic perception and reliance on feature point distribution.
[0005] Technical Solution
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: Based on a low-bandwidth end-edge collaborative visual detection method, the method includes the following steps:
[0007] S1. Use a lightweight semantic segmentation model to process raw images collected by IoT devices, identify target bounding boxes and category labels, and divide the area covered by the bounding box into semantic key blocks and the area outside the bounding box into background blocks.
[0008] S2. Prioritize the transmission of semantically critical blocks over highly reliable communication channels, employing automatic repeat request (ARQ) to ensure complete reception of critical blocks. Compress and transmit background blocks over low-power communication channels, ensuring that the background block data loss rate does not exceed a preset threshold.
[0009] S3. The edge server performs integrity checks on the semantic key blocks. If any missing blocks are detected, the IoT device is triggered to retransmit the key blocks. For completely received semantic key blocks, prior features are retrieved from the built-in semantic database in combination with their category labels to complete the key block details. For background blocks with a loss rate ≤ a preset threshold, a semantically guided interpolation algorithm is used to repair the missing segments. For background blocks with a loss rate greater than the preset threshold, similar scene textures are retrieved from the semantic database to fill in the missing segments.
[0010] S4. Semantically enhanced target detection: The reconstructed image and semantic mask (with key block locations and categories marked) are input into the target detection model, the detection confidence of the key block area is enhanced, and the detection result integrating the semantic information is output.
[0011] Preferably, the lightweight semantic segmentation model is MobileNet-SSD or YOLOv5n-seg, which outputs triplet information including target category, bounding box coordinates and pixel-level semantic mask.
[0012] Preferably, the semantic database storage content includes:
[0013] Target category prior feature library: contains feature templates such as contour, texture, color distribution, etc. of each target category;
[0014] Scene background texture library: Contains a set of high-resolution texture images of common backgrounds such as roads, farmland, and buildings.
[0015] Preferably, the semantically guided interpolation algorithm is specifically:
[0016] Based on the semantic labels of background blocks, similar texture samples are matched from the semantic database, and the structure and texture of missing fragments are reconstructed through a sparse representation algorithm.
[0017] Preferably, in the semantically enhanced target detection step, the detection accuracy is improved by:
[0018] For semantic key block areas, the classification confidence threshold of the object detection model is reduced by 30%;
[0019] For the background block area, the non-maximum suppression (NMS) algorithm is introduced to filter out false detection frames, and the suppression threshold is increased to 0.7.
[0020] Based on low-bandwidth edge-to-edge collaborative visual inspection system, including:
[0021] Semantic perception module: deploys a lightweight semantic segmentation model to identify semantic key blocks and background blocks in the original image and generate semantic masks;
[0022] Layered transport module:
[0023] Key block transmission unit: transmits semantic key blocks through a highly reliable channel and supports ARQ retransmission mechanism;
[0024] Background block transmission unit: compresses and transmits background blocks through low-power channels, and supports dynamic adjustment of compression ratio;
[0025] End-edge collaborative reconstruction module:
[0026] Key block check unit: verifies the integrity of semantic key blocks and triggers retransmission logic;
[0027] Semantic completion unit: performs feature completion and texture filling for key blocks and background blocks based on the semantic database;
[0028] Semantic Enhanced Detection Module: Integrates the target detection model, combines the semantic mask to detect the reconstructed image, and outputs the result with semantic labels.
[0029] Preferably, the semantic perception module and the layered transmission module are deployed on the IoT device side, and the end-edge collaborative reconstruction module and the semantic enhancement detection module are deployed on the edge server side, and the two realize semantic information interaction through a cross-layer communication protocol.
[0030] Preferably, the semantic database supports online updating, and automatically optimizes the target feature template and background texture library through historical detection data of the edge server.
[0031] Beneficial effects
[0032] The present invention provides a method and system for low-bandwidth edge-to-edge collaborative visual inspection. It has the following beneficial effects:
[0033] 1. The present invention provides a low-bandwidth end-edge collaborative visual detection method and system, which uses a lightweight semantic segmentation model to identify the target bounding box, divide it into semantic key blocks, and force complete transmission through a highly reliable channel (such as LTE-M) to avoid target information loss due to sparse feature points. Combined with the prior feature completion of the edge server semantic database (such as vehicle contours and pedestrian textures), the target area reconstruction error is reduced by more than 40%, and the average detection accuracy (mAP) is improved by 15-20%. For example, in low-contrast crop disease and pest detection scenarios, the priority transmission of key blocks can increase the detail retention rate of the diseased area from 65% of the original solution to 92%, and the missed detection rate is reduced by 35%.
[0034] 2. This invention provides a low-bandwidth, end-to-end collaborative visual detection method and system. Background blocks are compressed and transmitted over low-power channels (such as LoRa), allowing for 30% data loss. Combining a semantically guided interpolation algorithm with background texture library padding, this method reduces invalid data transmission by over 50% while ensuring basic fidelity for background reconstruction. While key blocks are reliably transmitted (including ARQ retransmissions), they only account for 20-30% of the total data volume. Overall end-to-end latency is reduced by over 60% compared to traditional full-image transmission solutions and by ≤20% compared to the original feature point solution, achieving an optimal balance between accuracy and latency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0037] like Figure 1 As shown, an embodiment of the present invention provides a low-bandwidth end-edge collaborative visual detection method, including the following steps:
[0038] S1. Use a lightweight semantic segmentation model to process raw images collected by IoT devices, identify target bounding boxes and category labels, and divide the area covered by the bounding box into semantic key blocks and the area outside the bounding box into background blocks.
[0039] S2. Prioritize the transmission of semantic key blocks over a highly reliable communication channel, using automatic retransmission request (ARQ) to ensure complete reception of key blocks. Background blocks are compressed and transmitted over a low-power communication channel, allowing the background block data loss rate to not exceed a preset threshold. A lightweight semantic segmentation model, such as MobileNet-SSD or YOLOv5n-seg, is used to output a triplet of object category, bounding box coordinates, and pixel-level semantic mask.
[0040] S3. The edge server performs integrity checks on the semantic key blocks. If a missing block is detected, the IoT device is triggered to retransmit the key block. For completely received semantic key blocks, the server uses the category label to retrieve prior features from the built-in semantic database to complete the key block details. For background blocks with a loss rate ≤ a preset threshold, the server uses a semantically guided interpolation algorithm to repair the missing fragments. For background blocks with a loss rate greater than the preset threshold, the server retrieves similar scene textures from the semantic database to fill in the missing fragments. The semantic database stores the following content:
[0041] Target category prior feature library: contains feature templates such as outline, texture, color distribution, etc. of each target category (such as vehicles and pedestrians);
[0042] Scene background texture library: Contains a set of high-resolution texture images of common backgrounds such as roads, farmland, and buildings;
[0043] The semantic-guided interpolation algorithm is specifically:
[0044] Based on the semantic labels of background blocks (such as "grass" and "wall"), similar texture samples are matched from the semantic database, and the structure and texture of the missing fragments are reconstructed using a sparse representation algorithm.
[0045] S4. Semantically enhanced object detection: The reconstructed image and the semantic mask (with the key block locations and categories marked) are input into the object detection model, the detection confidence of the key block area is enhanced, and the detection result integrating the semantic information is output. In the semantically enhanced object detection step, the detection accuracy is improved by the following methods:
[0046] For semantic key block areas, the classification confidence threshold of the object detection model is reduced by 30%;
[0047] For the background block area, the non-maximum suppression (NMS) algorithm is introduced to filter out false detection frames, and the suppression threshold is increased to 0.7.
[0048] Specific embodiment: Agricultural pest leaf detection based on low-bandwidth edge-edge collaboration
[0049] Application scenarios and parameter configuration
[0050] IoT device: A low-power camera deployed in farmland with a resolution of 1024×1024 pixels and equipped with the lightweight semantic segmentation model MobileNet-SSD (parameter size 1.44MB).
[0051] Communication channel:
[0052] Key block transmission: LTE-M (bandwidth 1Mbps, bit error rate 10 -5 );
[0053] Background block transmission: LoRa (bandwidth 50Kbps, bit error rate 10 -3 ).
[0054] Semantic database: pre-stored healthy / diseased leaf feature templates of crops such as rice and wheat (including 20 types of pest and disease textures), as well as a farmland background (soil, weeds) texture library.
[0055] Specific implementation steps
[0056] Step 1: Semantic-aware key block extraction
[0057] Input image: collected rice leaf image (including rice blast lesions), size 1024×1024×3.
[0058] Semantic Segmentation:
[0059] MobileNet-SSD outputs the target category "rice blast spot", with bounding box coordinates of (200, 300, 700, 600) (upper left corner x, y = 200, 300; lower right corner x, y = 700, 600), and semantic mask matrix
[0060] M∈R 1024 × 1024 (The key block area value is 1, and the background is 0).
[0061] Divide the image into (1024 / 16) according to the generative model input size (16×16 pixels) 2 = 4096 image patches, where the number of key patches covered by the bounding box is:
[0062]
[0063] The key block set is recorded as K = {K1, K2, ..., K 558}, the background block set is
[0064] Step 2: Hierarchical Priority Transmission
[0065] Key block transfer:
[0066] Each key block K i It is divided into four 512-byte sub-segments (16×16×3 pixels ≈ 768 bytes, compressed to 512 bytes), transmitted via LTE-M, and retransmitted using ARQ until all are correctly received.
[0067] Background block transfer:
[0068] The background blocks are compressed by WebP (compression ratio 8:1), and the average size of each block is about 150 bytes. They are transmitted via LoRa, and the allowed loss rate threshold is set to 30% (i.e., a single background block can be lost at most fragments, rounded to 0, i.e. the background block must be received completely).
[0069] Step 3: End-to-end collaborative semantic completion
[0070] Key block verification and feature completion:
[0071] After receiving the key block, the edge server calculates the checksum CRC32(K i ), and compare it with the checksum sent by the device. If they are inconsistent, retransmission is triggered.
[0072] Key block feature completion: The key block category label of the lesion is "rice blast spot", and the typical texture feature matrix T∈R of this category is retrieved from the semantic database 16×16×3 , enhancing details through weighted fusion: (α=0.7 is the fusion coefficient) If there is a missing area in the key block (caused by transmission error), the mask matrix M mask Locate missing pixels and perform fusion only on the missing regions:
[0073] It is a binary mask (1 means missing and needs to be completed).
[0074] Background block reconstruction:
[0075] The semantic label of the background block is "rice leaf background". If it is received completely, it is directly retained; if the loss rate is ≤30% (in this embodiment, the LoRa bit error rate is low, assuming no loss), similar textures are retrieved from the database based on the semantic label and repaired using the semantic-guided interpolation algorithm:
[0076]
[0077] Among them, Φ is the observation matrix (corresponding to the retained fragments), y is the observation value, D is the "rice leaf background" texture dictionary, λ=0.1 is the regularization parameter, and the missing fragments are reconstructed through sparse representation.
[0078] Step 4: Semantic Enhanced Object Detection
[0079] Reconstruct image stitching: According to the semantic mask M, the key blocks With background block B j Stitching into a complete image
[0080] Detection confidence adjustment:
[0081] The confidence threshold for key block area detection is reduced from the default 0.5 to 0.35, and the formula is:
[0082]
[0083] The non-maximum suppression (NMS) threshold of the background block area is increased from 0.5 to 0.7 to filter out overlapping false detection boxes.
[0084] Output result: Rice blast spots were detected with a confidence level of 0.91 and coordinates (210, 310, 690, 590), which is a significant improvement over the original solution (confidence level of 0.78).
[0085] Comparative experimental data
[0086]
[0087] Finally, in the agricultural pest and disease leaf detection embodiment, the IoT device locates the key blocks of rice blast spots through a lightweight semantic segmentation model. After reliable transmission via LTE-M, the edge server completes the details of the blast spots in combination with the prior features of the semantic database. The background blocks are compressed and transmitted via LoRa and reconstructed through semantic-guided interpolation. The test results show that the mAP of blast spot detection reaches 0.89, an improvement of 20.3% over the original feature point solution, and the end-to-end delay is 1.6 minutes, an increase of only 6.7%. In scenarios such as corn stalk rot and night orchards, the detection accuracy and reconstruction effect of the present invention in low-texture and complex lighting environments are significantly better than the existing technology, verifying the effectiveness of semantic-driven transmission and end-edge collaborative completion in improving detection accuracy and delay control.
[0088] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. Based on low-bandwidth edge-to-edge collaborative visual inspection method, characterized by: The following steps are involved: S1. Use a lightweight semantic segmentation model to process raw images collected by IoT devices, identify target bounding boxes and category labels, and divide the area covered by the bounding box into semantic key blocks and the area outside the bounding box into background blocks. S2. Prioritize the transmission of semantically critical blocks over highly reliable communication channels, employing automatic repeat request (ARQ) to ensure complete reception of critical blocks. Compress and transmit background blocks over low-power communication channels, ensuring that the background block data loss rate does not exceed a preset threshold. S3. The edge server performs integrity checks on the semantic key blocks. If any missing blocks are detected, the IoT device is triggered to retransmit the key blocks. For completely received semantic key blocks, prior features are retrieved from the built-in semantic database in combination with their category labels to complete the key block details. For background blocks with a loss rate ≤ a preset threshold, a semantically guided interpolation algorithm is used to repair the missing segments. For background blocks with a loss rate greater than the preset threshold, similar scene textures are retrieved from the semantic database to fill in the missing segments. S4. Semantically enhanced target detection: The reconstructed image and semantic mask (with key block locations and categories marked) are input into the target detection model, the detection confidence of the key block area is enhanced, and the detection result integrating the semantic information is output.
2. The low-bandwidth edge-to-edge collaborative visual inspection method according to claim 1 is characterized in that : The lightweight semantic segmentation model is MobileNet-SSD or YOLOv5n-seg, which outputs triplet information including target category, bounding box coordinates and pixel-level semantic mask.
3. The low-bandwidth edge-to-edge collaborative visual inspection method according to claim 1 is characterized in that : The semantic database storage content includes: Target category prior feature library: contains feature templates such as contour, texture, color distribution, etc. of each target category; Scene background texture library: Contains a set of high-resolution texture images of common backgrounds such as roads, farmland, and buildings.
4. The low-bandwidth edge-to-edge collaborative visual inspection method according to claim 1 is characterized in that : The semantic guided interpolation algorithm is specifically as follows: Based on the semantic labels of background blocks, similar texture samples are matched from the semantic database, and the structure and texture of missing fragments are reconstructed through a sparse representation algorithm.
5. The low-bandwidth edge-to-edge collaborative visual inspection method according to claim 1 is characterized in that In the semantic enhancement target detection step, the detection accuracy is improved by: For semantic key block areas, the classification confidence threshold of the object detection model is reduced by 30%; For the background block area, the non-maximum suppression (NMS) algorithm is introduced to filter out false detection frames, and the suppression threshold is increased to 0.
7.
6. Based on low-bandwidth edge-to-edge collaborative visual inspection system, it is characterized by: include: Semantic perception module: deploys a lightweight semantic segmentation model to identify semantic key blocks and background blocks in the original image and generate semantic masks; Layered transport module: Key block transmission unit: transmits semantic key blocks through a highly reliable channel and supports ARQ retransmission mechanism; Background block transmission unit: compresses and transmits background blocks through low-power channels, and supports dynamic adjustment of compression ratio; End-edge collaborative reconstruction module: Key block check unit: verifies the integrity of semantic key blocks and triggers retransmission logic; Semantic completion unit: performs feature completion and texture filling for key blocks and background blocks based on the semantic database; Semantic Enhanced Detection Module: Integrates the target detection model, combines the semantic mask to detect the reconstructed image, and outputs the result with semantic labels.
7. The low-bandwidth edge-to-edge collaborative visual inspection system according to claim 6, characterized in that: The semantic perception module and the layered transmission module are deployed on the IoT device side, and the end-edge collaborative reconstruction module and the semantic enhancement detection module are deployed on the edge server side. The two realize semantic information interaction through a cross-layer communication protocol.
8. The low-bandwidth edge-to-edge collaborative visual inspection system according to claim 6, characterized in that: The semantic database supports online updating and automatically optimizes the target feature template and background texture library through historical detection data of the edge server.
Citation Information
Patent Citations
Low-bandwidth end-side collaborative detection method and system based on generative visual model
CN118552490A