A video transmission method, device and equipment for bridge detection and a storage medium

By segmenting, encapsulating, transmitting, and parsing video data and defect identification data in bridge inspection, the problem of limited wireless communication bandwidth is solved, achieving low latency in video streams and high reliability in defect identification data, thereby improving the accuracy and efficiency of bridge inspection.

CN121418550BActive Publication Date: 2026-04-28SHAANXI DEXIN INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHAANXI DEXIN INTELLIGENT TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In bridge inspection, the wireless communication bandwidth between drones and ground stations is limited and unstable, resulting in video stream delays and loss of defect identification data, which affects the real-time performance and accuracy of remote diagnosis.

Method used

By segmenting and encapsulating video data to generate multiple segments, and serializing and encapsulating disease identification data, and transmitting it in conjunction with User Datagram Protocol (UDP) and Transmission Control Protocol (TCP), low latency of the video stream and high reliability of disease identification data are achieved. After transmission, the data is parsed, frame-identified, grouped and reassembled, and time-series synchronization calibration and overlay rendering are performed to generate an enhanced video stream.

Benefits of technology

It improves the anti-packet loss capability and synchronization accuracy of bridge inspection, builds an end-to-end closed-loop transmission system, ensures the smoothness of video stream and the accuracy of defect identification data, and improves the accuracy and efficiency of remote inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121418550B_ABST
    Figure CN121418550B_ABST
Patent Text Reader

Abstract

The application discloses a bridge detection video transmission method, device and equipment and a storage medium. The method comprises the following steps: acquiring video data collected by a UAV and disease identification data comprising information such as disease type, position and confidence; performing segmentation processing on the video data to obtain a plurality of segments, performing encapsulation processing on the segments and the disease identification data respectively, obtaining video stream data packets and disease identification data packets and transmitting the video stream data packets and the disease identification data packets; performing analysis and packet recombination on the video stream data packets to obtain recombined video frames, performing analysis on the disease identification data packets to determine boundaries to obtain identification results; performing time sequence synchronization calibration to obtain cooperative detection data, superimposing and rendering the disease information and the video frames to generate an enhanced video stream comprising a disease marking box, type text and a confidence identifier. The application improves the anti-packet loss capability and synchronization accuracy, constructs an end-to-end closed-loop transmission system, and effectively improves the accuracy and efficiency of bridge remote detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) video transmission technology, and in particular to a video transmission method, apparatus, equipment, and storage medium for bridge inspection. Background Technology

[0002] With the continuous advancement of infrastructure construction, bridges, as key nodes in transportation networks, face an increasingly urgent need for safety inspection. Drones, with their flexibility and efficiency, have been widely applied in bridge inspection scenarios. Equipped with high-definition cameras and edge computing devices, they can acquire continuous high-definition video streams of the bridge surface in real time and simultaneously perform AI-based defect identification, generating structured defect identification data that includes defect type, quantity, location, confidence level, and associated image tags. This data needs to be transmitted back to a ground station or cloud control center in real time to provide core evidence for remote diagnosis.

[0003] However, the wireless communication link between the UAV and the ground station is greatly affected by the environment, and its bandwidth is usually limited and unstable, often restricted to 4MB / s or even lower. Current bridge inspection data transmission schemes mostly use a single protocol: if the transmission control protocol is used entirely, the video stream will experience severe delays due to the protocol's retransmission and congestion control mechanisms, resulting in video stuttering and affecting real-time monitoring; if the user datagram protocol is used entirely, the defect identification data may be lost due to packet loss, causing missed defect information and reducing detection accuracy.

[0004] Therefore, under strict bandwidth constraints, designing an efficient transmission method that can simultaneously meet the requirements of low latency and smoothness of video streams and high reliability of disease identification data, and ensure the real-time performance and accuracy of remote diagnosis, is a technical problem that urgently needs to be solved. Summary of the Invention

[0005] In view of this, the video transmission method, apparatus, device, and storage medium for bridge inspection provided in this application can improve packet loss resistance and synchronization accuracy, construct an end-to-end closed-loop transmission system, and effectively improve the accuracy and efficiency of remote bridge inspection. The video transmission method, apparatus, device, and storage medium for bridge inspection provided in this application are implemented as follows:

[0006] This application provides a video transmission method for bridge detection, comprising:

[0007] Acquire video data and disease identification data, wherein the disease identification data includes at least one of disease type, quantity, location, confidence level, and associated image identifiers;

[0008] The video data is segmented to obtain multiple segments, each segment including a frame identifier, segment index, total number of segments, data size, and total frame size;

[0009] Multiple segments are encapsulated to obtain video stream data packets, and the disease identification data is serialized and encapsulated to obtain disease identification data packets.

[0010] The video stream data packets and the disease identification data packets are transmitted. The transmitted video stream data packets are parsed, frame identifiers are grouped and reassembled to obtain reassembled video frames. The transmitted disease identification data packets are parsed and boundary determination is performed to obtain disease identification results.

[0011] The reconstructed video frames and the disease identification results are subjected to time-series synchronization calibration to obtain collaborative detection data;

[0012] The disease identification data in the collaborative detection data is overlaid with the corresponding video frames to obtain an enhanced video stream, which includes disease annotation boxes, type text, and confidence level indicators.

[0013] In some embodiments, the step of performing time-series synchronization calibration processing on the reconstructed video frames and the disease identification results to obtain collaborative detection data includes:

[0014] The timestamps of the reconstructed video frames and the disease identification results are obtained to get multiple sets of timestamp data.

[0015] The multiple sets of timestamp data are processed by difference calculation to obtain delay difference data;

[0016] The timestamp of the disease identification result is calibrated based on the delay difference data to obtain the calibrated timestamp.

[0017] The associated image identifiers in the reconstructed video frames are matched with the associated image identifiers in the disease identification results to obtain initial matching data.

[0018] A second matching and verification process is performed on the timestamps of the reconstructed video frames in the initial matching data and the calibrated timestamps to obtain collaborative detection data.

[0019] In some embodiments, the step of overlaying and rendering the disease identification data in the collaborative detection data with the corresponding video frames to obtain an enhanced video stream includes:

[0020] The disease location information in the collaborative detection data is subjected to coordinate transformation to obtain pixel-level disease locations;

[0021] The pixel-level disease location, disease type, and confidence level are rendered to obtain rendering configuration information;

[0022] The corresponding video frames are drawn according to the rendering configuration information to obtain enhanced video frames;

[0023] The enhanced video frames are spliced ​​and integrated to obtain an enhanced video stream.

[0024] In some embodiments, the segmentation of the video data to obtain multiple segments includes:

[0025] The maximum transmission unit of the network is detected to obtain the maximum data packet size;

[0026] Based on the maximum data packet size, the video data is adapted to obtain the segmentation rules;

[0027] Based on the segmentation rules, the video data is segmented frame by frame to obtain multiple basic video segments.

[0028] After adding header information to each basic video segment, the basic video segments are integrated to obtain multiple segments.

[0029] In some embodiments, the process of parsing, grouping, and reassembling the transmitted video stream data packets to obtain reassembled video frames includes:

[0030] The transmitted video stream data packets are parsed to obtain the parsing results;

[0031] Based on the frame identifiers in the parsing results, the video stream data packets are grouped to obtain a set of frame packets.

[0032] Based on the fragmentation index in the parsing result, the video stream data packets are sequentially arranged to obtain an ordered data packet sequence;

[0033] The ordered data packet sequence is subjected to timeout monitoring, discarding, and buffer release processing to obtain valid frame packets;

[0034] The video stream data packets in the effective frame group are spliced ​​and integrated to obtain reconstructed video frames.

[0035] In some embodiments, the parsing and boundary determination processing of the transmitted disease identification data packet to obtain the disease identification result includes:

[0036] The transmitted disease identification data packet is read and processed to obtain the original data stream;

[0037] The original data stream is parsed to obtain the protocol header parsing result;

[0038] Based on the protocol header parsing results, the original data stream is separated to obtain the target data segment;

[0039] The target data segment is deserialized to obtain initial disease identification data;

[0040] The initial disease identification data is verified to obtain the disease identification results.

[0041] In some embodiments, the header information includes at least one of an identifier, a fragment index, a total number of fragments, a data size, and a total frame size.

[0042] This application provides a video transmission device for bridge inspection, comprising:

[0043] The acquisition module is used to acquire video data and disease identification data, wherein the disease identification data includes at least one of disease type, quantity, location, confidence level, and associated image identifiers;

[0044] The processing module is used to segment the video data to obtain multiple segments, each segment including a frame identifier, segment index, total number of segments, data size, and total frame size;

[0045] The processing module is also used to encapsulate multiple segments to obtain video stream data packets and to serialize and encapsulate the disease identification data to obtain disease identification data packets.

[0046] The transmission module is used to transmit the video stream data packets and the disease identification data packets, and to parse, frame identify group and reassemble the transmitted video stream data packets to obtain reassembled video frames, and to parse and boundary determine the transmitted disease identification data packets to obtain disease identification results.

[0047] The processing module is also used to perform time-series synchronization calibration processing on the reconstructed video frames and the disease identification results to obtain collaborative detection data;

[0048] The processing module is also used to overlay and render the disease identification data in the collaborative detection data with the corresponding video frames to obtain an enhanced video stream, which includes disease annotation boxes, type text, and confidence level indicators.

[0049] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.

[0050] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method described in this application embodiment.

[0051] This application provides a video transmission method, apparatus, device, and storage medium for bridge inspection. It acquires video data collected by a drone and defect identification data including defect type, location, and confidence level. The video data is segmented into multiple segments, and each segment and defect identification data is encapsulated to obtain video stream data packets and defect identification data packets, which are then transmitted. The video stream data packets are parsed, grouped, and reassembled to obtain reconstructed video frames. The defect identification data packets are parsed to determine boundaries and obtain identification results. Cooperative detection data is obtained through time-series synchronization calibration. Defect information and video frames are overlaid and rendered to generate an enhanced video stream containing defect annotation boxes, type text, and confidence level indicators. This improves packet loss resistance and synchronization accuracy, constructs an end-to-end closed-loop transmission system, effectively improves the accuracy and efficiency of remote bridge inspection, and solves the technical problems mentioned in the background art. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A schematic diagram illustrating the implementation process of a video transmission method for bridge detection provided in this application embodiment;

[0054] Figure 2 This is a schematic diagram illustrating an implementation process for acquiring collaborative detection data, provided in an embodiment of this application.

[0055] Figure 3 This is a schematic diagram of the structure of a video transmission device for bridge detection provided in an embodiment of this application. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0057] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0058] Figure 1 This is a schematic diagram illustrating the implementation flow of a video transmission method for bridge detection provided in an embodiment of this application, including steps 101 to 106. Wherein, Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order of a video transmission method for bridge detection. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.

[0059] Step 101: Obtain video data and disease identification data.

[0060] In this embodiment, the bridge inspection drone is equipped with a high-definition camera and an edge computing device. The high-definition camera collects video data of the bridge surface in real time, covering image information of key inspection areas such as bridge beams, piers, and bridge deck. The edge computing device runs a pre-trained deep learning defect recognition model to analyze each frame of video data in real time, identify bridge defects such as cracks, concrete spalling, and steel corrosion, and generate corresponding defect recognition data.

[0061] The disease identification data is structured information, including at least the disease type (such as cracks, spalling, etc.), the number of diseases, the specific location of the disease in the video frame, the identification confidence level (reflecting the reliability of disease identification), and the associated image identifier. The associated image identifier is consistent with the frame identifier of the corresponding video frame to ensure that a precise correspondence can be established between the two. The UAV synchronously acquires the video data and the disease identification data through a data interface.

[0062] Step 102: The video data is segmented to obtain multiple segments.

[0063] In this embodiment, to adapt to bandwidth-constrained transmission environments and avoid transmission delays or packet loss due to excessively large single-frame video data, video data needs to be segmented. First, the maximum transmission unit (MTB) value of the current wireless communication link is obtained. This MTB is then used to perform adaptation calculations based on the single-frame size of the video data to determine reasonable segmentation rules. This ensures that the data packets formed after encapsulation of a single segment do not exceed the MTB limit, while also considering bandwidth utilization efficiency.

[0064] According to the segmentation rules, each frame of video data is divided frame by frame to obtain multiple fixed-length or variable-length basic video segments. The size of the last segment can be flexibly adjusted according to the actual amount of remaining data. Subsequently, header information is added to each basic video segment. This header information includes a frame identifier (used to distinguish different video frames), a segment index (used to identify the order of segments within the same frame), a total number of segments (the total number of segments in the same video frame), a data size (the amount of data in the current segment), and a total frame size (the total amount of data corresponding to the original video frame), ultimately forming multiple segments that meet the transmission requirements.

[0065] Step 103: Encapsulate multiple segments to obtain video stream data packets and serialize and encapsulate disease identification data to obtain disease identification data packets.

[0066] In this embodiment, the fragments with header information are encapsulated. During encapsulation, the transmission protocol identifier of the data packet is explicitly defined to ensure the receiving end can accurately identify the data type. Simultaneously, the encapsulated data packets are configured with transmission priority, setting the video stream data packets to high priority to avoid latency accumulation due to low priority during transmission. Pre-processing for transmission is completed using a User Datagram Protocol (UDP) socket, forming a video stream data packet that can be transmitted via a wireless link.

[0067] To address the high reliability requirements of disease identification data, the acquired data is first structured and organized, clarifying the hierarchical relationships and logical connections between data fields to ensure a standardized and consistent data format. Subsequently, the organized disease identification data is serialized into a compact structured data format for easy network transmission and parsing, resulting in serialized data.

[0068] To address the packet fragmentation issue in Transmission Control Protocol (TCP) data streams, a protocol header is added to the serialized data. This header includes a data type identifier and the data length. The TCP then encapsulates the serialized data with the header and performs reliable transmission preprocessing via TCP sockets, resulting in a transmittable disease identification data packet.

[0069] Step 104: Transmit the video stream data packet and the disease identification data packet. Then, parse, group, and reassemble the transmitted video stream data packet to obtain the reassembled video frame. Finally, parse and determine the boundary of the transmitted disease identification data packet to obtain the disease identification result.

[0070] In this embodiment, the UAV synchronously transmits video stream data packets and disease identification data packets to the ground station receiver via a wireless communication link. During transmission, the video stream data packets are sent in a low-latency manner based on the connectionless nature of the User Datagram Protocol (UDP); the disease identification data packets are delivered completely through reliable transmission mechanisms based on the Transmission Control Protocol (TCP), including acknowledgment retransmission and flow control.

[0071] The ground station receiver receives video stream data packets via User Datagram Protocol (UDP) sockets. First, it parses the header information of each data packet to extract key information such as frame identifier, fragment index, and total number of fragments. Based on the frame identifier, data packets belonging to the same video frame are grouped and integrated to obtain a frame group set. Then, based on the fragment index, the data packets in each frame group set are arranged sequentially to obtain an ordered data packet sequence.

[0072] Simultaneously, a frame-level timeout monitoring mechanism is initiated for each frame packet set, with a reasonable timeout period set. If a frame packet set fails to collect all fragments within the timeout period, it is determined to be a transmission failure, the data packets corresponding to that frame packet set are discarded, and the occupied buffer resources are released; if all fragments are collected within the timeout period, the data in the ordered data packet sequence is spliced ​​and integrated to restore the complete video frame data. After the integrity verification confirms that there are no errors, the reconstructed video frame is obtained.

[0073] The ground station receiver receives the disease identification data packets via Transmission Control Protocol (TCP) sockets, reads the data stream byte-by-byte to obtain the raw data stream, parses the protocol header in the raw data stream to extract the data type identifier and data length information, determines the boundaries of the complete data packet based on the data length, and separates the corresponding complete data segment from the raw data stream to obtain the target data segment.

[0074] The target data segment is deserialized to restore it to structured data containing disease type, quantity, location, confidence level, and associated image identifiers, thus obtaining initial disease identification data. Subsequently, the initial disease identification data is checked for field completeness and logical rationality, and invalid data with missing fields or logical contradictions is removed, ultimately yielding accurate disease identification results.

[0075] Step 105: Perform time-series synchronization calibration on the reconstructed video frames and the disease identification results to obtain collaborative detection data.

[0076] In this embodiment, to address the timing misalignment issue between video stream data packets and disease identification results during transmission due to differences in protocol characteristics and network latency, timing synchronization calibration is required. First, the generation and reception timestamps of the reconstructed video frames are collected, along with the generation and reception timestamps of the disease identification results, resulting in multiple sets of timestamp data.

[0077] The differences between multiple sets of timestamp data are calculated to analyze the transmission delay differences between the reconstructed video frames and the disease identification results, thus obtaining delay difference data. Based on the delay difference data, the associated timestamps of the disease identification results are calibrated and adjusted to ensure that the time bases of the two are consistent, resulting in calibrated timestamps.

[0078] The first matching process is performed between the associated image identifiers of the reconstructed video frames and the associated image identifiers of the disease identification results to filter out the initial matching data with consistent identifiers. Then, the timestamps of the reconstructed video frames in the initial matching data are matched and verified with the timestamps of the calibrated disease identification results to correct minor time-series deviations. Finally, a one-to-one correspondence between the reconstructed video frames and the disease identification results is established to obtain collaborative detection data.

[0079] Step 106: Overlay rendering processing is performed on the disease identification data in the collaborative detection data and the corresponding video frames to obtain an enhanced video stream.

[0080] In this embodiment of the application, the location information of the disease in the collaborative detection data is processed by coordinate transformation, and the location coordinates of the disease in the actual detection scene are mapped to the pixel coordinate system of the corresponding reconstructed video frame to obtain the pixel-level disease location, ensuring that the disease annotation can be accurately superimposed on the corresponding position of the video frame.

[0081] Rendering parameters are set for pixel-level lesion locations, lesion type text, and confidence scores. The style of the annotation box, the font, size, and color parameters of the text are determined to ensure that the annotation information is clearly identifiable and does not obscure key detection details in the video frame, thus obtaining the rendering configuration information.

[0082] Based on the rendering configuration information, highlighted bounding boxes are drawn at specified pixel positions in the corresponding video frames, and the defect type and confidence level are overlaid to obtain enhanced video frames. All enhanced video frames are stitched together in chronological order to form a continuous enhanced video stream. This enhanced video stream can intuitively display the location, type, and identification reliability of bridge defects.

[0083] The ground station will enhance the video stream for real-time decoding and display, allowing operators to intuitively observe the bridge's condition and defects. At the same time, it will classify and store the enhanced video stream, collaborative detection data, raw video data, and defect identification data, establishing an indexed detection database to form a traceable and analyzable detection data system.

[0084] This application's embodiments effectively address the core contradiction in bridge inspection data transmission under bandwidth-constrained environments through differentiated processing and collaborative optimization. On one hand, video data segmentation and encapsulation adapt to bandwidth characteristics, while the serialization and encapsulation of defect identification data ensure transmission reliability, achieving a dual-objective optimization of low-latency video stream transmission and high-completeness defect identification data transmission. On the other hand, through parsing and reconstruction, temporal synchronization calibration, and overlay rendering, a precise correlation between video frames and defect information is established, generating an intuitive enhanced video stream and constructing an end-to-end closed loop of acquisition-transmission-processing-display. This significantly improves the real-time performance, accuracy, and ease of operation of remote bridge inspection, adapting to bandwidth-constrained and unstable wireless transmission scenarios.

[0085] In the above Figure 1 Based on the above, this application also provides a schematic diagram of the implementation process for obtaining collaborative detection data, as shown below. Figure 2 As shown, steps 201 to 205 are included:

[0086] Step 201: Obtain the timestamps of the reconstructed video frames and the timestamps of the disease identification results to obtain multiple sets of timestamp data.

[0087] In this embodiment, the ground station receiver first collects multiple sets of timestamp data for the reconstructed video frames and the disease identification results. The timestamp for the reconstructed video frames includes two parts: the original generation timestamp generated when the UAV's high-definition camera captures the video frame, and the receiving timestamp recorded by the ground station receiver after completing the video frame reconstruction. Similarly, the timestamp for the disease identification results includes two parts: the original generation timestamp when the UAV's edge computing device completes disease identification and generates the result, and the receiving timestamp recorded by the ground station receiver after parsing the disease identification data packet. All timestamps are recorded using a unified time base to ensure data comparability, ultimately forming multiple sets of timestamp data.

[0088] Step 202: Perform difference calculation on multiple sets of timestamp data to obtain delay difference data.

[0089] In this embodiment, the difference between multiple sets of timestamp data is calculated to obtain delay difference data. The specific calculation logic is as follows: the transmission delay of the reconstructed video frame (video frame reception timestamp minus its generation timestamp) and the transmission delay of the disease identification result (disease identification result reception timestamp minus its generation timestamp) are calculated respectively, and then the transmission delays of the two types of data are differentially calculated to obtain delay difference data reflecting the difference in transmission delay between the two.

[0090] Step 203: The timestamp of the disease identification result is calibrated based on the delay difference data to obtain the calibrated timestamp.

[0091] In this embodiment, the original generation timestamp of the disease identification result is calibrated based on the calculated delay difference data. If the transmission delay of the disease identification result is greater than the transmission delay of the reconstructed video frame, the corresponding delay difference is added to the original generation timestamp of the disease identification result; if the transmission delay of the disease identification result is less than the transmission delay of the reconstructed video frame, the corresponding delay difference is subtracted from the original generation timestamp of the disease identification result, thus obtaining the calibrated timestamp. This calibration operation ensures that the time reference of the reconstructed video frame and the disease identification result is consistent.

[0092] Step 204: Perform a first-level matching process between the associated image identifiers in the reconstructed video frames and the associated image identifiers in the disease identification results to obtain initial matching data.

[0093] In this embodiment, the associated image identifiers in the reconstructed video frames are matched with the associated image identifiers in the disease identification results. The associated image identifier is a unique identifier assigned synchronously by the UAV when acquiring video frames and generating disease identification results, ensuring that video frames and corresponding disease identification results in the same detection scenario have the same associated image identifier. During the matching process, reconstructed video frames and disease identification results with completely identical associated image identifiers are selected to form initial matching data, preliminarily eliminating invalid data with no corresponding relationship.

[0094] Step 205: Perform a second matching and verification process on the timestamps of the reconstructed video frames in the initial matching data and the calibrated timestamps to obtain collaborative detection data.

[0095] In this embodiment, a second matching verification process is performed on the original generation timestamp of the reconstructed video frame and the calibrated timestamp of the disease identification result in the initial matching data. A reasonable time tolerance threshold is set. If the difference between the original generation timestamp of the reconstructed video frame and the calibrated timestamp of the disease identification result in the same set of initial matching data is within the preset tolerance threshold range, the match is considered valid. If the difference exceeds the tolerance threshold, it is determined that the timing deviation has not been fully corrected, and the set of data is discarded. Through this dual matching verification, collaborative detection data corresponding one-to-one between the reconstructed video frame and the disease identification result is finally obtained, ensuring that the identified disease information can be accurately associated with the corresponding video frame.

[0096] This application employs a dual matching and timestamp calibration mechanism to accurately resolve the timing misalignment issue between video streams and disease identification results caused by differences in transmission protocols and network latency. The calculation and calibration of differences between multiple sets of timestamp data ensures consistency in the time base of the two types of data. A first-level matching of associated image identifiers initially filters valid data, while a second-level verification of timestamps further corrects minor deviations, ultimately achieving a precise one-to-one correspondence between the reconstructed video frames and the disease identification results. This improves the accuracy and reliability of data association, avoids mismatches between disease information and video footage due to timing misalignment, provides high-quality collaborative detection data for subsequent overlay rendering, and ensures the accuracy of remote diagnosis.

[0097] In some embodiments, the disease identification data in the collaborative detection data is overlaid with the corresponding video frames to obtain an enhanced video stream, including: performing coordinate transformation processing on the disease location information in the collaborative detection data to obtain pixel-level disease locations.

[0098] Specifically, the disease location information in the collaborative detection data is based on the actual physical coordinates recorded when the drone captured the video. It needs to be converted into the pixel coordinates of the corresponding video frames to achieve accurate labeling. First, the shooting parameters of the drone's high-definition camera (including focal length, shooting distance, sensor size, etc.) and the resolution information of the video frames (such as pixel width and height) are obtained.

[0099] Based on the above parameters, a mapping relationship between physical coordinates and pixel coordinates is established: taking the upper left corner of the video frame as the origin of pixel coordinates, and combining the perspective projection principle of the shooting angle, the physical position coordinates of the lesion in the actual scene are converted into pixel coordinate values ​​within the video frame, ensuring that the position of the lesion at the pixel level completely corresponds to the actual position in the video screen, and finally obtaining the pixel-level lesion position.

[0100] Furthermore, the pixel-level location, type, and confidence level of the disease are rendered to obtain rendering configuration information.

[0101] Specifically, rendering parameters are configured for pixel-level defect locations, defect types, and confidence levels to form rendering configuration information. For different defect types (such as cracks, concrete spalling, and steel corrosion), a unique annotation box style is set for each defect, including different highlight colors (e.g., red for cracks, yellow for spalling) and line thickness (usually 2-3 pixels), making it easier for operators to quickly distinguish defect types.

[0102] Configure the disease type text, determining a uniform font, font size, and text color; standardize the confidence level format and configure a display style with the same color as the text to avoid obscuring key information in the video. All configuration parameters are then integrated to form complete rendering configuration information.

[0103] Furthermore, the corresponding video frames are drawn based on the rendering configuration information to obtain enhanced video frames.

[0104] Specifically, based on the rendering configuration information, drawing operations are performed on the corresponding reconstructed video frames. First, the pixel-level lesion location is located, and a highlighted annotation box is drawn according to the configured style, ensuring that the annotation box accurately surrounds the lesion area, neither exceeding the lesion range nor omitting key parts; then, lesion type text and confidence information are overlaid in the blank area near the annotation box, with the text position maintaining a reasonable distance from the annotation box to ensure visual continuity.

[0105] During the drawing process, layer overlay technology is used to draw the annotation boxes and text information on an independent visualization layer, which is then overlaid and merged with the original video frame layer without changing the integrity of the original video data, ultimately resulting in an enhanced video frame with clear disease annotations.

[0106] Furthermore, the enhanced video frames are spliced ​​and integrated to obtain the enhanced video stream.

[0107] Specifically, all enhanced video frames are stitched together in chronological order. The frames are sorted according to their capture timestamps to ensure the frame order matches the actual drone shooting order, avoiding stuttering or timing discrepancies. During the stitching process, the video stream's frame rate is kept stable to guarantee smooth playback.

[0108] After integration, an enhanced video stream is generated. This stream retains the continuous footage of the original video while overlaying key identification information such as disease type, location, and confidence level in real time. Simultaneously, the enhanced video stream is decoded and displayed in real time for intuitive observation by ground station operators. The enhanced video stream, along with corresponding collaborative detection data and original detection data, is then categorized and stored in the detection database, forming a traceable and analyzable detection record.

[0109] This application's embodiments transform abstract disease identification data into intuitive visual information through precise coordinate transformation and standardized rendering configuration. Pixel-level coordinate transformation ensures that the disease marking location perfectly matches the actual disease location in the video image; differentiated rendering configuration makes different types of diseases clearly distinguishable without obscuring key video information; enhanced video frame stitching and integration ensure the continuity and smoothness of the video stream. It allows for intuitive observation of the location, type, and reliability of bridge diseases, reducing detection difficulty while enhancing the video stream's traceability and analyzability.

[0110] In some embodiments, video data is segmented to obtain multiple fragments, including: probing the network's maximum transmission unit to obtain the maximum data packet size.

[0111] Specifically, to avoid transmission failures or additional overhead caused by exceeding network transmission limits after video data is fragmented, the maximum transmission unit (MTB) of the wireless communication link between the UAV and the ground station is first detected. A standard MTB detection mechanism is used, sending test data packets of different sizes to the ground station receiver and combining this with the data packet reception feedback status to determine the maximum data packet size that the link can transmit normally.

[0112] During the probing process, the test is gradually increased from a smaller packet size until packet loss or fragmentation is detected, ultimately determining the maximum packet size supported by the current link.

[0113] Furthermore, the video data is adapted based on the maximum data packet size to obtain the segmentation rules.

[0114] Specifically, based on the detected maximum data packet size and the characteristics of the single frame size of the video data, adaptation processing is performed to determine the segmentation rules. First, the average size and fluctuation range of the single frame of the video data collected by the drone are statistically analyzed. The optimal data payload size of a single segment is calculated based on the maximum data packet size. Typically, the segment data payload is set to a value that does not exceed the maximum data packet size minus the overhead of subsequent custom header information.

[0115] For single-frame video data, fixed-length segmentation is performed according to the calculated optimal payload size. If the size of a single-frame data cannot be divided evenly by the optimal payload size, the last segment is processed as a variable-length segment, containing only the remaining data without the need for additional padding of invalid data, in order to optimize bandwidth utilization. At the same time, the segmentation order identification rules are clearly defined to ensure that each segment can be accurately identified as belonging to its video frame and its position within the frame, providing a basis for subsequent header information addition and receiver reassembly.

[0116] Furthermore, the video data is segmented frame by frame based on the segmentation rules to obtain multiple basic video segments.

[0117] Specifically, the video data is segmented frame by frame according to the determined segmentation rules. Starting from the first frame of video data, data segments of a set size are extracted sequentially as a basic video segment. After the segmentation of all data in the current frame is completed, the same operation is performed on the next frame of video data until all video data is segmented.

[0118] During the segmentation process, the original order of the video data is strictly maintained, and the integrity of the data content is not changed, ensuring that each basic video segment is a continuous fragment of the original data of the corresponding video frame. The final result is multiple basic video segments arranged in frame order and segment order. Each segment contains only the original video data and no additional identification information has been added.

[0119] Furthermore, after adding header information to each basic video segment, each basic video segment is integrated to obtain multiple segments.

[0120] Specifically, to enable the receiving end to accurately reassemble video frames, header information needs to be added to each basic video segment and then integrated. The header information includes frame identifier, segment index, total number of segments, data size, and total frame size: the frame identifier is a unique identifier assigned to each video frame to distinguish different video frames; the segment index identifies the sequential number of the current segment within its respective video frame; the total number of segments is the total number of segments in the corresponding video frame; the data size is the actual data volume of the current segment; and the total frame size is the complete data volume of the corresponding original video frame.

[0121] The header information and corresponding basic video segments are concatenated and integrated in a fixed order, with the header information first and the basic video segment data last, forming a final segment with a complete structure and containing the key information required for reassembly. During the integration process, seamless connection between the header information and the segment data is ensured, without introducing any additional redundant data, ultimately resulting in multiple segments that meet the transmission requirements.

[0122] This application's embodiments lay the foundation for efficient video stream transmission through network characteristic adaptation and standardized processing. The maximum packet size detected by the network's maximum transmission unit ensures that secondary fragmentation at the IP layer is not triggered after fragmentation and encapsulation, reducing additional transmission overhead. Fragmentation rules based on the maximum packet size balance bandwidth utilization and transmission stability; the combination of fixed-length and variable-length segmentation methods avoids invalid data padding. The addition of header information provides a complete basis for reassembly at the receiving end. This enhances the adaptability of video data transmission in bandwidth-constrained environments, improves packet loss resistance, and ensures the efficiency and stability of video stream transmission.

[0123] In some embodiments, parsing, frame identifier grouping, and reassembly processing of the transmitted video stream data packets to obtain reassembled video frames includes: parsing the transmitted video stream data packets to obtain parsing results.

[0124] Specifically, the ground station receiver receives the transmitted video stream data packets via a User Datagram Protocol (UDP) socket and first initiates the parsing process. The parsing process focuses on the header information of the data packets, extracting key fields according to the pre-set encapsulation format on the UAV side, including frame identifier, fragment index, total number of fragments, data size, and total frame size. The frame identifier distinguishes different original video frames and is the core basis for subsequent grouping; the fragment index indicates the sequential position of the current data packet within its respective video frame; the total number of fragments clarifies the total number of fragments in the corresponding video frame, used to determine whether all data has been received. After parsing, a parsing result containing the above key information is generated, ensuring that the core attributes of each data packet can be accurately identified.

[0125] Furthermore, the video stream data packets are grouped based on the frame identifiers in the parsing results to obtain a set of frame packets.

[0126] Specifically, based on the frame identifiers in the parsing results, all received video stream data packets are grouped. Each original video frame is assigned a unique identifier. The receiving end maintains a temporary buffer structure to group data packets with the same frame identifier into the same set, forming a frame group set. Each frame group set corresponds to all fragments of an original video frame to be reassembled. The grouping process strictly adheres to the principle of frame identifier uniqueness to avoid confusion between data packets from different video frames and ensure that each data packet within a set corresponds to only a single video frame.

[0127] Furthermore, based on the fragment index in the parsing results, the video stream data packets are sequentially arranged to obtain an ordered data packet sequence.

[0128] Specifically, for each set of frame packets, the packets are sequentially arranged according to the fragment index in the parsing result. The fragment indexes are assigned in natural number order (e.g., 1, 2, 3…), clearly identifying the logical order of each data packet within its respective video frame. During arrangement, the data packets within the frame packet set are rearranged in ascending order based on the fragment index, forming an ordered sequence of data packets. This operation ensures that data packets that might have arrived out of order due to network transmission are restored to their original sending order.

[0129] Furthermore, timeout monitoring, discarding, and buffer release are performed on the ordered data packet sequence to obtain valid frame groups.

[0130] Specifically, to prevent incomplete frame fragment sets from occupying buffer resources for extended periods while ensuring the real-time performance of the video stream, a frame-level timeout monitoring mechanism is implemented for each ordered data packet sequence. The timeout period is flexibly set based on the actual wireless communication link conditions to ensure that fragment reception is completed within an acceptable network latency range, while avoiding video stuttering caused by timeouts.

[0131] During monitoring, the fragment reception status of each frame packet set is tracked in real time to determine whether all fragments corresponding to that frame have been received. If all fragments are received within the timeout period, it is determined to be a valid frame packet, and the ordered data packet sequence is retained; if all fragments are not received after the timeout, it is determined to be an invalid frame packet, and all data packets corresponding to that sequence are discarded, and the buffer resources they occupy are released to prevent memory leaks and ensure efficient utilization of the receiving end system resources.

[0132] Furthermore, the video stream data packets in the valid frame group are spliced ​​and integrated to obtain reconstructed video frames.

[0133] Specifically, for frame groups deemed valid, their corresponding ordered data packet sequences are extracted and then spliced ​​together. During splicing, the video data payload is extracted from each data packet in the order of the ordered data packet sequence, the header information is removed, and all video data fragments are seamlessly spliced ​​together in their original order to restore a complete single-frame video data. After splicing, the integrity of the complete video data is verified. Once it is confirmed that there are no missing or incorrect data, a reconstructed video frame that can be decoded and played normally is obtained. This reconstructed video frame retains the image quality and content integrity of the original video.

[0134] This application's embodiments utilize precise parsing, ordered grouping, and timeout control to efficiently reconstruct complete video frames in bandwidth-constrained scenarios where packet loss is possible. Data packet parsing extracts key information, providing a basis for grouping and sorting; frame-identifier-based grouping avoids data confusion between different video frames, and fragment index-based sorting restores the original data order; frame-level timeout monitoring and cache release mechanisms effectively clean up invalid data from incomplete fragments, preventing memory leaks and ensuring efficient utilization of system resources. This ensures that video frames with all fragments are completely reconstructed while maintaining the continuity and real-time nature of the video stream, improving system robustness and adapting to unstable transmission conditions under bandwidth constraints.

[0135] In some embodiments, parsing and boundary determination processing are performed on the transmitted disease identification data packets to obtain disease identification results, including: reading the transmitted disease identification data packets to obtain the original data stream.

[0136] Specifically, the ground station receiver receives the transmitted disease identification data packets via Transmission Control Protocol (TCP) sockets. Since TCP is a streaming protocol, multiple data packets may be transmitted in a fragmented manner. Therefore, the received data stream needs to be read continuously in byte order. During the reading process, the receiver maintains a data buffer, continuously storing the received bytes to ensure no transmitted bytes are missed, ultimately forming a continuous and complete raw data stream.

[0137] Furthermore, the original data stream is parsed to obtain the protocol header parsing result.

[0138] Specifically, for the raw data stream, the protocol header is specially parsed according to the pre-set encapsulation format of the UAV. The protocol header is a fixed-length field added during the encapsulation process of the UAV, containing two core pieces of information: first, a data type identifier, used to clearly identify that the current data stream is bridge defect identification data, avoiding confusion with other types of transmitted data; second, data length, used to identify the effective data volume of a single complete defect identification data packet.

[0139] During the parsing process, the receiving end extracts the protocol header fields according to the agreed byte order. First, it identifies the data type identifier and confirms that it is the target disease identification data. Then, it extracts the data length information to form a protocol header parsing result containing the data type and the effective data length.

[0140] Furthermore, the original data stream is separated based on the protocol header parsing results to obtain the target data segment.

[0141] Specifically, based on the data length information in the protocol header parsing result, the original data stream is separated to resolve the packet merging problem that may occur during streaming transmission of the Transmission Control Protocol. The receiving end extracts a byte segment from the original data stream, starting from the first byte after the protocol header, with a length equal to the data length. This byte segment is the serialized data corresponding to a single complete disease identification data point, which is the target data segment.

[0142] After separation, if there are still remaining bytes in the original data stream, the remaining bytes are used as a new original data stream, and the above protocol header parsing and separation process is repeated until all received data streams are processed, ensuring that each complete disease identification data packet can be accurately separated without data confusion or interception errors.

[0143] Furthermore, the target data segment is deserialized to obtain initial disease identification data.

[0144] Specifically, since the drone serializes the disease identification data, the receiving end needs to deserialize the target data segment to restore it to directly usable structured data. The receiving end calls the deserialization tool corresponding to the serialization format to parse the byte stream of the target data segment and restore it to structured data containing specific fields, i.e., the initial disease identification data.

[0145] The initial disease identification data is complete structured information, including at least the disease type (such as cracks, concrete spalling, etc.), the number of diseases, the specific location of the disease in the corresponding video frame, the identification confidence level, and the associated image identifiers. All fields are consistent with the disease identification data generated by the UAV to ensure the accuracy of data reconstruction.

[0146] Furthermore, the initial disease identification data is verified to obtain the disease identification results.

[0147] Specifically, to ensure the reliability of disease identification results, the initial disease identification data needs to undergo dual verification. First, a field integrity check is performed to verify whether the initial disease identification data contains all necessary fields (such as disease type, location, associated image identifiers, etc.). If any fields are missing, the data is deemed invalid. Second, a logical validity check is performed to verify the logical validity of each field's data. For example, it checks whether the confidence level is within a reasonable range of 0-100%, and whether the disease location coordinates are within the pixel coordinate range of the corresponding video frame. If logical contradictions exist, the data is deemed invalid.

[0148] After verification, all invalid data is removed, and the initial disease identification data with complete fields and reasonable logic is retained, so as to obtain accurate and effective disease identification results.

[0149] This application's embodiments ensure the integrity, accuracy, and validity of disease identification data through step-by-step parsing and dual verification. Reading the original data stream and parsing the protocol header precisely resolves the packet fragmentation problem in streaming transmission of the transmission control protocol, ensuring accurate separation of individual complete data packets. Deserialization restores the structured initial disease identification data, maintaining consistency with the original data from the UAV. Dual verification of field integrity and logical rationality eliminates invalid data, preventing misjudgments of disease information due to missing data or logical contradictions. This ensures that the transmitted disease identification data can be completely and accurately restored, providing reliable structured data for time-series synchronization calibration and overlay rendering, avoiding missed or erroneous disease information reports, and guaranteeing the accuracy of remote detection.

[0150] While this application provides method operation steps as shown in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the method can be executed sequentially according to this embodiment or the accompanying drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0151] like Figure 3 As shown in the illustration, this application also provides a video transmission device 300 for bridge inspection. The device includes:

[0152] The acquisition module 301 is used to acquire video data and disease identification data, wherein the disease identification data includes at least one of disease type, quantity, location, confidence level and associated image identifier.

[0153] The processing module 302 is used to segment the video data to obtain multiple segments. Each segment includes a frame identifier, a segment index, a total number of segments, a data size, and a total frame size.

[0154] The processing module 302 is also used to encapsulate multiple segments to obtain video stream data packets and to serialize and encapsulate the disease identification data to obtain disease identification data packets.

[0155] The transmission module 303 is used to transmit the video stream data packet and the disease identification data packet, parse the transmitted video stream data packet, perform frame identification grouping and reassembly processing to obtain the reassembled video frame, and parse and perform boundary determination processing on the transmitted disease identification data packet to obtain the disease identification result.

[0156] The processing module 302 is further used to perform time-series synchronization calibration processing on the reconstructed video frame and the disease identification result to obtain collaborative detection data.

[0157] The processing module 302 is further configured to overlay and render the disease identification data in the collaborative detection data with the corresponding video frames to obtain an enhanced video stream, wherein the enhanced video stream includes disease annotation boxes, type text, and confidence level indicators.

[0158] In some embodiments, the acquisition module 301 is further configured to acquire the timestamp of the reconstructed video frame and the timestamp of the disease identification result, thereby obtaining multiple sets of timestamp data.

[0159] Processing module 302 is also used to perform difference calculation processing on the multiple sets of timestamp data to obtain delay difference data.

[0160] The processing module 302 is further configured to calibrate the timestamp of the disease identification result based on the delay difference data to obtain a calibrated timestamp.

[0161] The processing module 302 is further configured to perform a first-level matching process on the associated image identifier in the recombined video frame and the associated image identifier in the disease identification result to obtain initial matching data.

[0162] The processing module 302 is further configured to perform a second matching verification process on the timestamps of the recombined video frames in the initial matching data and the calibrated timestamps to obtain collaborative detection data.

[0163] In some embodiments, the processing module 302 is further configured to perform coordinate transformation processing on the disease location information in the collaborative detection data to obtain pixel-level disease locations.

[0164] The processing module 302 is also used to perform rendering processing on the pixel-level disease location, disease type and confidence level to obtain rendering configuration information.

[0165] The processing module 302 is also used to perform drawing processing on the corresponding video frame according to the rendering configuration information to obtain an enhanced video frame.

[0166] The processing module 302 is also used to splice and integrate the enhanced video frames to obtain an enhanced video stream.

[0167] In some embodiments, the processing module 302 is further configured to perform detection processing on the network maximum transmission unit to obtain the maximum data packet size.

[0168] The processing module 302 is also used to perform adaptation processing on the video data based on the maximum data packet size to obtain the segmentation rules.

[0169] The processing module 302 is also used to perform frame-by-frame segmentation of the video data based on the segmentation rules to obtain multiple basic video segments.

[0170] The processing module 302 is also used to add header information to each video basic segment and then integrate each video basic segment to obtain multiple segments.

[0171] In some embodiments, the processing module 302 is further configured to parse the transmitted video stream data packets to obtain the parsing result.

[0172] The processing module 302 is further configured to group the video stream data packets based on the frame identifiers in the parsing results to obtain a set of frame packets.

[0173] Processing module 302 is further configured to perform sequential arrangement processing on the video stream data packets based on the fragment index in the parsing result to obtain an ordered data packet sequence.

[0174] The processing module 302 is also used to perform timeout monitoring, discarding, and buffer release processing on the ordered data packet sequence to obtain valid frame groups.

[0175] The processing module 302 is also used to splice and integrate the video stream data packets in the effective frame group to obtain reconstructed video frames.

[0176] In some embodiments, the processing module 302 is further configured to read and process the transmitted disease identification data packet to obtain the original data stream.

[0177] The processing module 302 is also used to parse the original data stream to obtain the protocol header parsing result.

[0178] The processing module 302 is further configured to perform separation processing on the original data stream based on the protocol header parsing result to obtain the target data segment.

[0179] The processing module 302 is also used to deserialize the target data segment to obtain initial disease identification data.

[0180] The processing module 302 is also used to perform verification processing on the initial disease identification data to obtain the disease identification result.

[0181] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0182] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0183] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0184] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.

[0185] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.

[0186] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.

[0187] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.

[0188] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0189] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0190] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A video transmission method for bridge inspection, characterized in that, include: Acquire video data and disease identification data, wherein the disease identification data includes at least one of disease type, quantity, location, confidence level, and associated image identifiers; The video data is segmented to obtain multiple segments, each segment including a frame identifier, segment index, total number of segments, data size, and total frame size; Multiple segments are encapsulated to obtain video stream data packets, and the disease identification data is serialized and encapsulated to obtain disease identification data packets. The video stream data packets and the disease identification data packets are transmitted. The transmitted video stream data packets are parsed, frame identifiers are grouped and reassembled to obtain reassembled video frames. The transmitted disease identification data packets are parsed and boundary determination is performed to obtain disease identification results. The reconstructed video frames and the disease identification results are subjected to time-series synchronization calibration to obtain collaborative detection data; The disease identification data in the collaborative detection data is overlaid and rendered with the corresponding video frames to obtain an enhanced video stream, which includes disease annotation boxes, type text and confidence labels. The step of performing time-series synchronization calibration on the reconstructed video frames and the disease identification results to obtain collaborative detection data includes: The timestamps of the reconstructed video frames and the disease identification results are obtained to get multiple sets of timestamp data. The multiple sets of timestamp data are processed by difference calculation to obtain delay difference data; The timestamp of the disease identification result is calibrated based on the delay difference data to obtain the calibrated timestamp. The associated image identifiers in the reconstructed video frames are matched with the associated image identifiers in the disease identification results to obtain initial matching data. A second matching and verification process is performed on the timestamps of the reconstructed video frames in the initial matching data and the calibrated timestamps to obtain collaborative detection data. The process of overlaying and rendering the disease identification data in the collaborative detection data with the corresponding video frames to obtain an enhanced video stream includes: The disease location information in the collaborative detection data is subjected to coordinate transformation to obtain pixel-level disease locations; The pixel-level disease location, disease type, and confidence level are rendered to obtain rendering configuration information; The corresponding video frames are drawn according to the rendering configuration information to obtain enhanced video frames; The enhanced video frames are spliced ​​and integrated to obtain an enhanced video stream.

2. The method according to claim 1, characterized in that, The video data is segmented to obtain multiple segments, including: The maximum transmission unit of the network is detected to obtain the maximum data packet size; Based on the maximum data packet size, the video data is adapted to obtain the segmentation rules; Based on the segmentation rules, the video data is segmented frame by frame to obtain multiple basic video segments. After adding header information to each basic video segment, the basic video segments are integrated to obtain multiple segments.

3. The method according to claim 1, characterized in that, The process of parsing, grouping, and reassembling the transmitted video stream data packets to obtain reassembled video frames includes: The transmitted video stream data packets are parsed to obtain the parsing results; Based on the frame identifiers in the parsing results, the video stream data packets are grouped to obtain a set of frame packets. Based on the fragmentation index in the parsing result, the video stream data packets are sequentially arranged to obtain an ordered data packet sequence; The ordered data packet sequence is subjected to timeout monitoring, discarding, and buffer release processing to obtain valid frame packets; The video stream data packets in the effective frame group are spliced ​​and integrated to obtain reconstructed video frames.

4. The method according to claim 1, characterized in that, The process of parsing and boundary determination of the transmitted disease identification data packet to obtain the disease identification result includes: The transmitted disease identification data packet is read and processed to obtain the original data stream; The original data stream is parsed to obtain the protocol header parsing result; Based on the protocol header parsing results, the original data stream is separated to obtain the target data segment; The target data segment is deserialized to obtain initial disease identification data; The initial disease identification data is verified to obtain the disease identification results.

5. The method according to claim 2, characterized in that, The header information includes at least one of the following: identifier, fragment index, total number of fragments, data size, and total frame size.

6. A video transmission device for bridge inspection, characterized in that, include: The acquisition module is used to acquire video data and disease identification data, wherein the disease identification data includes at least one of disease type, quantity, location, confidence level, and associated image identifiers; The processing module is used to segment the video data to obtain multiple segments, each segment including a frame identifier, segment index, total number of segments, data size, and total frame size; The processing module is also used to encapsulate multiple segments to obtain video stream data packets and to serialize and encapsulate the disease identification data to obtain disease identification data packets. The transmission module is used to transmit the video stream data packets and the disease identification data packets, and to parse, frame identify group and reassemble the transmitted video stream data packets to obtain reassembled video frames, and to parse and boundary determine the transmitted disease identification data packets to obtain disease identification results. The processing module is also used to perform time-series synchronization calibration processing on the reconstructed video frames and the disease identification results to obtain collaborative detection data; The processing module is also used to overlay and render the disease identification data in the collaborative detection data with the corresponding video frames to obtain an enhanced video stream, which includes disease annotation boxes, type text and confidence indicators. The processing module is further configured to perform time-series synchronization calibration processing on the reconstructed video frames and the disease identification results to obtain collaborative detection data, wherein: The timestamps of the reconstructed video frames and the disease identification results are obtained to get multiple sets of timestamp data. The multiple sets of timestamp data are processed by difference calculation to obtain delay difference data; The timestamp of the disease identification result is calibrated based on the delay difference data to obtain the calibrated timestamp. The associated image identifiers in the reconstructed video frames are matched with the associated image identifiers in the disease identification results to obtain initial matching data. A second matching and verification process is performed on the timestamps of the reconstructed video frames in the initial matching data and the calibrated timestamps to obtain collaborative detection data. The processing module is further configured to overlay and render the disease identification data in the collaborative detection data with the corresponding video frames to obtain an enhanced video stream, wherein: The disease location information in the collaborative detection data is subjected to coordinate transformation to obtain pixel-level disease locations; The pixel-level disease location, disease type, and confidence level are rendered to obtain rendering configuration information; The corresponding video frames are drawn according to the rendering configuration information to obtain enhanced video frames; The enhanced video frames are spliced ​​and integrated to obtain an enhanced video stream.

7. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-channel video data processing method and device, electronic equipment and medium

    CN113259715A

  • Road disease display method, system, device and equipment and readable storage medium

    CN117041685A

  • Video information processing method and system

    CN120583200A

  • Intelligent event identification method and system based on high-speed camera

    CN121121021A