A time delay test method and system for a remote driving system

By inserting disturbance frames into the remote driving system and performing signal processing analysis, a sparse feature space and delay inversion model are constructed, which solves the problem of large delay measurement errors in the existing technology, realizes accurate identification and correction of link delay, and improves measurement accuracy and availability.

CN120567348BActive Publication Date: 2026-01-27WUHU SIMBA NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510672487.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2026-01-27
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In existing remote driving systems, the time-stamp-based delay measurement mechanism cannot accurately reflect the real time of data processing or transmission, resulting in a large invisible error in delay measurement. Especially under non-explicit processing behaviors such as asynchronous scheduling and data caching and queuing, it is impossible to achieve microsecond-level accurate identification and correction.

Method used

By inserting preset disturbance frames into image frames, processing and transmission timestamps of each node are collected, and signal processing analysis is performed at the cabin end to construct a sparse feature space and delay inversion optimization model, extract delay features of each stage, and correct the link delay calculated based on timestamps.

Benefits of technology

It enables effective identification and correction of ghost delays in the link, improves the accuracy and engineering availability of delay measurement, and can identify implicit scheduling delays and anomalies in the link, achieving microsecond-level measurement accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567348B_ABST
    Figure CN120567348B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of delay test, especially to a kind of remote control driving system's delay test method and system, the present application proposes the following scheme, by inserting the preset disturbance frame in image frame, the processing and transmission timestamp of image frame in each node is collected, and the image response signal of disturbance frame is analyzed in cabin end signal processing, the delay characteristic of each stage is extracted, for correcting the link delay calculated based on timestamp. By constructing sparse feature space and delay inversion optimization model, the accurate identification and microsecond-level correction of multi-stage processing delay are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of latency testing technology, and in particular to a latency testing method and system for a remote-controlled driving system. Background Technology

[0002] In existing remote-controlled driving systems, to achieve real-time closed-loop control of vehicle-side image data transmission and cabin-side control commands, a timestamp-based latency measurement mechanism is commonly used. This involves marking key nodes such as image acquisition, encoding, transmission, decoding, and display, and calculating link latency based on the time difference between each node. However, in actual deployment environments, due to non-explicit processing behaviors such as asynchronous scheduling, data caching and queuing, and protocol layer packet reassembly, timestamps cannot accurately reflect the physical moment when data is "actually processed or sent," resulting in significant invisible errors.

[0003] This application presents a time delay testing method and system for a remote-controlled driving system. Summary of the Invention

[0004] The technical problem this invention aims to solve is to address the shortcomings of existing technologies by providing a latency testing method and system for remote-controlled driving systems. This method involves inserting preset disturbance frames into image frames, acquiring the processing and transmission timestamps of the image frames at each node, and performing signal processing analysis on the image response signals of the disturbance frames at the cabin end to extract latency features at each stage. These features are then used to correct the link latency calculated based on timestamps. By constructing a sparse feature space and a latency inversion optimization model, accurate identification and microsecond-level correction of multi-stage processing latency are achieved.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A latency testing method for a remote driving system is provided, applied to testing a remote driving system including a vehicle-side terminal, a cloud-side terminal, and a cabin-side terminal. The latency testing method includes:

[0007] The timestamps of the collected image frames are acquired and transmitted and processed at the vehicle end, cloud end, and cabin end, wherein preset disturbance frames are inserted into the image frames;

[0008] The latency of each sending and receiving stage and processing stage is determined based on each timestamp;

[0009] The image response signal of the disturbance frame at the cabin end is processed and analyzed to extract delay features, and the delay calculated based on the timestamp is corrected according to the delay features.

[0010] Before obtaining each timestamp, the latency testing method further includes:

[0011] A time synchronization module based on the PTP protocol is configured for the vehicle end and the cabin end respectively, and a time synchronization module based on the NTP protocol is configured for the cloud.

[0012] A clock deviation prediction model is constructed based on historical deviation sequences. The dynamic offset between cloud NTP time and other terminal PTP time is predicted by the clock deviation prediction model, and the cloud time reference is corrected according to the dynamic offset.

[0013] Before acquiring the timestamps of image frames being transmitted, received, and processed at the vehicle end, the latency testing method further includes:

[0014] In the vehicle-side image acquisition and encoding stage, image frames are acquired by the vehicle-side image sensor and input to the motion estimation stage of the encoder. Before being input to the encoder, a perturbation frame with preset characteristics is inserted into the image frame.

[0015] Based on the system time of PTP time synchronization, a timestamp T1 is injected into the image frame through the encoder's software development interface;

[0016] In the image processing flow, after the controller at the vehicle end receives the image frame, it marks the image reception timestamp T2, and after the processing is completed, it marks the processing completion timestamp T3.

[0017] The processed image frames are encapsulated into TCP packets for streaming, and the streaming timestamp T4 is marked.

[0018] Before acquiring the timestamps of image frames being sent, received, and processed in the cloud, the latency testing method further includes:

[0019] The cloud receives TCP packets from the vehicle and performs unpacking processing;

[0020] The received timestamp is dynamically corrected according to the clock deviation prediction model to generate a corresponding cloud-corrected timestamp T5.

[0021] The cloud-corrected timestamp T5 is written into the RTP protocol extension header and the message data is forwarded to the cabin via the RTSP protocol.

[0022] Before acquiring the timestamps of image frames being transmitted, received, and processed at the cabin end, the latency testing method further includes:

[0023] The cabin receives the message data forwarded by the cloud through the communication equipment, and unpacks and extracts the timestamps T1 to T5 from the message;

[0024] The marker module receives the current timestamp T6 of the image frame;

[0025] The image frame is rendered and output through a display device, and the image display timestamp T7 is marked.

[0026] The step of inserting a perturbation frame with preset characteristics into the image frame before it is input to the encoder includes:

[0027] Based on the feature deconstruction requirements of the transmission and reception and processing stages of delay detection, the perturbation frame is encoded into multiple non-overlapping delay coding regions.

[0028] Structural perturbation features are added to the time-delay coding region according to the stage characteristics, wherein the structural perturbation features are generated by frequency distribution coding, phase offset coding and block matching coding.

[0029] The disturbance frames are periodically inserted into the image frames and transmitted alternately with the image frames at a preset ratio.

[0030] Signal processing analysis is performed on the image response signal of the disturbance frame at the cabin end to extract delay features, including:

[0031] The response frame sequence of the disturbance frame in the cabin-end display device is collected, and an image response observation matrix is ​​constructed based on the response frame sequence;

[0032] The image response observation matrix is ​​projected onto the sparse feature space through the sparse transform domain to obtain a sparse representation of the perturbation signal in the multi-scale feature domain.

[0033] Construct a joint optimization model for delay inversion with sparse representation as the observed variable and propagation delay at each stage as the sparse coefficient, wherein the joint optimization model for delay inversion includes sparse regularization terms and structural consistency constraint terms;

[0034] According to the aforementioned delay inversion joint optimization model, the sparse representations of each stage are jointly inverted and solved to obtain the sparse coefficient distribution, which is then used as the delay feature.

[0035] The step of projecting the image response observation matrix into the sparse feature space through the sparse transform domain includes:

[0036] The image response observation matrix is ​​transformed by an orthogonal sparse transformation operator to construct a response sparse representation tensor in the mapping domain. The orthogonal sparse transformation operator is a preset segmented modulation basis, and its construction maintains structural consistency with the coding structure corresponding to each stage in the perturbed frame.

[0037] The principal components of the response sparse representation tensor are concentrated in the sparse subspace corresponding to each stage by controlling the dimensionality suppression function, so as to generate a sparse feature space.

[0038] The joint inversion solution of the sparse representations at each stage includes:

[0039] The sparse representations corresponding to each time-delay coding region in the perturbation frame are solved in parallel to obtain the sparse coefficient vector.

[0040] The sparse coefficient vectors of each region are converged to the optimal solution domain through iterative optimization.

[0041] After iterative optimization, the non-zero components in the sparse coefficient vector are calibrated in stages, and the response time represented by their corresponding positions is used as the sparse coefficient distribution of each stage.

[0042] A latency testing system for a remote-controlled driving system, the system comprising:

[0043] The acquisition module is used to acquire image frames and insert preset disturbance frames into the image frames. The image frames are transmitted and processed sequentially via the vehicle end, cloud end, and cabin end.

[0044] The timestamp recording module is used to record the timestamp information of image frames at each transmission and reception stage and processing stage at the vehicle end, cloud end, and cabin end, respectively.

[0045] The latency calculation module is used to determine the latency of each transmission and reception stage and processing stage based on the timestamps of each stage.

[0046] The response acquisition module is used to acquire the image response signal of the disturbance frame in the cabin-end display device;

[0047] The signal analysis module is used to perform signal processing analysis on the image response signal and extract delay features that reflect the propagation delay at each stage;

[0048] The delay correction module is used to correct the delay calculated based on the timestamp according to the delay characteristics.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] This invention constructs a delay inversion model that can be mapped to each link processing stage by inserting disturbance frames with staged structural features into image frames and combining image response signal acquisition and sparse feature analysis at the cabin end, thereby realizing the effective identification and correction of ghost delays in the link. Attached Figure Description

[0051] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0052] Figure 1 This is a schematic diagram illustrating an exemplary application scenario of an embodiment of the present invention;

[0053] Figure 2 This is a flowchart illustrating a delay testing method for a remote-controlled driving system according to an embodiment of the present invention.

[0054] Figure 3 This is a flowchart illustrating a timestamp marking method according to an embodiment of the present invention;

[0055] Figure 4 This is a schematic diagram illustrating the principle of the perturbation frame in an embodiment of the present invention;

[0056] Figure 5 This is a schematic diagram of the signal processing and analysis method according to an embodiment of the present invention. Detailed Implementation

[0057] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0058] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0059] Please see Figure 1 This figure is a schematic diagram of an exemplary application scenario provided by an embodiment of this application.

[0060] like Figure 1 As shown, the application scenario is a remote driving architecture consisting of the vehicle, cloud, and cabin. These three components are interconnected via a transmission network and each accesses a time synchronization mechanism to achieve time synchronization across multiple nodes.

[0061] Figure 1 The PTP (Precision Time Protocol) timing mechanism for vehicle-side access is shown. Specifically, it includes a GMSL camera for acquiring raw image frames and periodically inserting perturbation frames into the image stream, and a vehicle-side controller for receiving, processing, encoding and streaming the acquired image frames, while recording timestamps T1 to T4 at different processing stages.

[0062] Figure 1 The cloud access NTP (Network Time Protocol) timing mechanism is shown, which includes an unpacking module for receiving TCP packets from the vehicle and extracting image frames, and an error compensation module for correcting the time base difference between the cloud time and PTP time based on a clock deviation prediction model and generating a corrected timestamp T5.

[0063] Figure 1 The diagram shows that the cabin also uses the PTP timing mechanism, which includes an unpacking module for receiving and parsing image data streams from the cloud and recording the display time timestamp T6; a display output and rendering module for rendering the received image frames and recording the display time timestamp T7; a disturbance frame decoding module for decoding and extracting features from the disturbance structures in the image frames; and a timestamp delay correction module for correcting the delay calculated based on the timestamp based on the delay characteristics of the disturbance frames in the image response.

[0064] In one example Figure 1 The transmission network shown can be configured using a 5G network to reduce transmission latency and achieve real-time control of autonomous vehicles.

[0065] In one example, a cloud server is used as a repeater to establish a data connection between the remote-controlled vehicle and the remote-controlled cabin. A PTP time synchronization system unifies the time base for both devices, and the cloud uses a clock skew prediction model to compensate for timing errors. The remote-controlled vehicle, cloud server, and remote-controlled cabin each inject timestamps through a software SDK to calculate the transmission latency. The remote-controlled vehicle includes an onboard controller for sending and receiving data, the remote-controlled cabin connects to the cabin display via an HDMI cable, and the cloud server has RTSP transceiver capabilities.

[0066] In one example, T2-T1 represents the camera GMSL transmission delay;

[0067] T3-T2 represents the vehicle-side image processing latency;

[0068] T4-T3 represent the vehicle-side encoding push latency;

[0069] T5-T4 represent the transmission latency from the vehicle to the cloud;

[0070] T6-T5 represent the transmission latency from the cloud to the cabin.

[0071] T7-T6 represent cabin display delay;

[0072] T7-T1 represents the end-to-end video stream latency of the remote driving system.

[0073] Next, with reference to the accompanying drawings, a delay testing method for a remote-controlled driving system provided in this application will be described.

[0074] Please see Figure 2 The figure is a flowchart illustrating a delay testing method for a remote driving system provided in an embodiment of this application. Figure 2 The method shown can be applied to testing remote driving systems, which include vehicle-side, cloud-side, and cabin-side components. Figure 2The method shown includes the following steps S1-S4, and the specific steps are as follows:

[0075] S1: Obtain the timestamps of the acquired image frames during transmission, reception, and processing at the vehicle, cloud, and cabin terminals;

[0076] In this embodiment, the vehicle-side acquires image frames via a GMSL camera. Before the image enters the encoding process, a perturbation frame is inserted to embed a reference signal for subsequent processing delay analysis. Subsequently, the image frames are marked with timestamps T1, T2, T3, and T4 at the image sensor encoding stage, the controller receiving stage, and the image streaming stage, respectively. After receiving the image frame message, the cloud unpacks it and calculates the time offset from NTP to PTP based on the historical time deviation model. The corrected cloud-received timestamp is T5. The cabin-side records timestamps T6 and T7 sequentially during the reception, rendering, and display of the image frames. By constructing the timestamp sequence T1 to T7, a complete model of the processing and transmission trajectory of the image frame in the entire remote control link can be created, facilitating subsequent delay decomposition and localization. This solves the technical defect of traditional delay measurement that only focuses on the start and end points and ignores link details.

[0077] S2: Determine the latency of each sending and receiving stage and processing stage based on each timestamp;

[0078] In this embodiment, timestamp sequences T1 to T7 are used to calculate the latency of each stage of image frame acquisition and processing at the vehicle end (T1-T3), transmission from the vehicle end to the cloud (T4-T5), transmission from the cloud to the cabin end (T5-T6), and cabin end rendering and display (T6-T7), thereby achieving a refined latency distribution estimation for each processing sub-stage in the entire remote video link. This enables the detection of potential bottlenecks or abnormal waiting in a certain stage of the system, thereby improving the latency visualization capability of the remote control system.

[0079] S3: Perform signal processing and analysis on the image response signal of the disturbance frame at the cabin end to extract delay features;

[0080] In this embodiment, the cabin end captures the image sequence of disturbance frames output by the display screen and constructs an image response observation matrix. Based on the design structure of each delay coding region in the disturbance frame, an orthogonal sparse transformation operator is used to map the response observation matrix to a sparse feature space. Subsequently, based on the correspondence between the sparse representation and the predefined disturbance structure, a delay inversion joint optimization model is constructed, and the sparse coefficient distribution of each coding region is extracted through an iterative sparse reconstruction method. The position of the coefficients has a one-to-one correspondence with the delay of each processing stage. Through the above-mentioned disturbance response signal modeling and joint solution mechanism, the proactive identification of "non-causal delays" such as implicit scheduling delays and queuing waiting that cannot be detected by traditional timestamps can be achieved, breaking through the technical blind spots of existing timestamp-based measurements.

[0081] S4: Correct the latency calculated based on the timestamp according to the latency characteristics;

[0082] In this embodiment, the obtained sparse delay features are used to correct the delay results of the timestamp calculation. The timestamp difference at each stage is combined with the physical response delay extracted from the sparse coefficients for error compensation and boundary smoothing, thereby constructing a high-precision delay estimate that more closely resembles the actual physical link. This correction process can effectively eliminate error offsets caused by multi-level caching, operating system scheduling, and asynchronous transmission, improve the engineering availability and microsecond-level accuracy of the measurement values, and solve the measurement bias problem caused by unobservable delays in existing pure timestamp mechanisms.

[0083] While timestamp-based end-to-end latency measurement is a relatively mature solution in existing remote driving links, its core assumption remains that timestamps accurately represent the actual moment of data processing or transmission. However, this assumption does not hold true in real-world environments with multi-level asynchronous transmission, buffer queuing, and uncontrollable module scheduling. Taking vehicle-side video streaming as an example, even if a timestamp is added when the image frame is encoded, the actual message transmission time may be offset by tens of milliseconds due to network queuing and system scheduling jitter. This ghost latency cannot be directly captured by timestamps, resulting in imperceptible errors in the calculated link latency even when time points T1 to T7 are measured. Furthermore, existing technologies have attempted to optimize timestamp strategies through high-frequency sampling, clock synchronization, or a combination of hardware and software, but these methods essentially remain at the level of optimizing the timestamp itself and cannot escape the logical bottleneck of relying on the system to describe itself. The end result is that, in remote control scenarios with high real-time requirements, existing timestamp measurement methods are unable to support microsecond-level latency assessment and accurate segmentation identification, especially when there is an opaque intermediate layer scheduling mechanism in the link, the measured values ​​deviate significantly from the physical reality.

[0084] For example, in the video image processing flow at the vehicle end, after the image frame is encoded, it is pushed and transmitted by the controller through the network protocol stack. Typically, to calculate the processing and transmission latency of this frame, a timestamp is added before encoding is completed or before the data is submitted to the network interface. However, in practical applications, this timestamp cannot accurately reflect the actual moment the image frame is sent. This is because, at the engineering level, after data is submitted from the application layer to the network interface, it usually enters the operating system's kernel buffer and is then asynchronously scheduled by the network protocol stack. This scheduling process involves multiple non-deterministic factors such as queuing, buffer merging, priority sorting, and system interrupt response. The final data transmission often lags behind the timestamp recording time by tens of milliseconds, and this process lacks explicit callbacks or event hooks that can be invoked by user space. Therefore, even if an attempt is made to timestamp when the transmission occurs, only the data submission action can be marked, not the actual departure of the data, resulting in an unavoidable error between the timestamp record and the actual physical behavior.

[0085] Furthermore, without modifying the driver layer code or introducing external hardware for data acquisition, existing operating systems and network protocol stacks do not provide any standard mechanism for accurately recording data transmission time. Even if forced intervention is achieved through driver layer instrumentation or packet capture mechanisms, it faces engineering problems such as high system stability risks, poor cross-platform adaptability, and high resource consumption.

[0086] To address this, this embodiment introduces a perturbation frame as an active physical reference signal, constructing a visualized, responsive, and resolvable embedded probe mechanism within the image frame transmission link, thereby overcoming the limitations of the timestamp observation system. In its implementation, the perturbation frame employs a structured design, embedding perturbation features with different encoding modes into multiple regions of the image frame. These perturbation regions remain unchanged during each processing stage (e.g., acquisition, processing, encoding, and transmission), but exhibit different temporal response characteristics due to system scheduling and queuing behavior. After the output is displayed at the cabin end, the perturbation frame undergoes time-series sampling and image observation. By combining sparse transform mapping and joint optimization algorithms, the response time of the actual processing action at each stage can be deduced from the physical response, thus obtaining a true time delay characteristic based on signal behavior rather than system claims. Through this physical signal entity, the perturbation frame transforms the originally unobservable and unmeasurable temporal delay in the link into a computable delay feature.

[0087] Furthermore, by fusing and correcting the delay features extracted from perturbation frame signal processing with the timestamp calculation value, not only can the delay calculation results of each processing stage be dynamically adjusted, but also hidden bottlenecks and abnormal delay segments in the link can be effectively identified, thus completing and enhancing the traditional timestamp scheme.

[0088] Please see Figure 3The figure is a flowchart illustrating a timestamp marking method provided in an embodiment of this application. Figure 3 The method shown can be applied as a preliminary step in a delay testing method for a remote driving system according to this application. Figure 3 The method shown includes the following A1-A4, and the specific steps are as follows:

[0089] A1: Configure a time synchronization module based on the PTP protocol for the vehicle end and the cabin end respectively, and configure a time synchronization module based on the NTP protocol for the cloud.

[0090] Specifically, in order to ensure that the time base of each node in the remote driving link is consistent and comparable, it is necessary to establish time synchronization mechanisms for the vehicle, cloud and cabin ends to support the synchronization of subsequent timestamps.

[0091] In this embodiment, the vehicle-side and cabin-side terminals, acting as terminal execution nodes, are typically deployed in edge hardware with industrial-grade or near-real-time control capabilities. Therefore, the IEEE 1588 Precision Time Protocol (PTP) is chosen as the timing scheme, achieving nanosecond-level synchronization through a master clock device configured in the local area network. The PTP master clock can be driven by a GPS module or a high-precision temperature-compensated oscillator, completing the timing process within the network layer through message exchange procedures such as Sync, Follow_Up, Delay_Req, and Delay_Resp. The cabin-side terminal also uses the same PTP client module to ensure that its local time is in the same time domain as the vehicle-side terminal.

[0092] Furthermore, since the cloud is deployed in a virtual server or data center environment, it is limited by network isolation and cross-area transmission bandwidth. To reduce deployment complexity and ensure cross-regional accessibility, this embodiment selects a network time synchronization service based on the standard NTP (Network Time Protocol) protocol. The cloud periodically synchronizes its time with a trusted NTP time source and collects offset samples based on its own operating status to establish a long-term time reference. Through the above configuration, the vehicle end and cabin end in the remote control link construct a high-precision time synchronization domain based on PTP, while the cloud establishes a relatively low-precision but predictable time reference through NTP, providing a foundation for subsequent multi-terminal timestamp fusion and deviation correction.

[0093] A2: After acquiring image frames, timestamps T1-T4 are sequentially marked on the vehicle end;

[0094] Specifically, in this embodiment, the timestamp marking process for image frames is embedded in a key node of the vehicle-side image processing link, recording the actual processing time of image frames at different stages. First, the image sensor module acquires image frames at a fixed frame rate and transmits them to the encoding chip via the image bus. When the acquisition is completed and the image enters the image buffer, a timestamp T1 is injected through the SDK interface provided by the encoder. T1 represents the time when the image content is initially captured.

[0095] Furthermore, after the image frame enters the encoding module, the vehicle-side controller acquires the image frame through interrupt callbacks or DMA event listening. Upon receiving a buffered copy of the complete image data, the controller immediately records a timestamp T2, reflecting the start time of image content processing. After image processing is complete, including image format conversion, preprocessing, and encoding, the controller records a timestamp T3, indicating the processing endpoint when the image data is ready for streaming. Finally, when the TCP streaming module encapsulates the image frame into a network packet and submits it to the kernel network buffer, it records a timestamp T4, reflecting the final node behavior before the vehicle-side completes streaming. All four timestamps are read from the system time via the PTP time synchronization module and written into the image frame metadata using a high-resolution clock source (accuracy down to the microsecond level), serving as raw data for subsequent link latency assessment.

[0096] A3: Forward TCP packets to the cloud. The cloud dynamically corrects the received timestamp based on the clock skew prediction model and generates the corresponding cloud-corrected timestamp T5.

[0097] Specifically, after the vehicle-mounted terminal completes the TCP packet encapsulation of the image frame, it pushes it to the cloud server node via a remote communication link. Since the cloud uses the NTP protocol for time synchronization, there is a systematic deviation between the time base and the vehicle-mounted terminal's PTP time domain. Without correction, the difference between the cloud's received time T5 and T4 will be incomparable. In this embodiment, to solve this problem, the cloud introduces a deviation prediction model based on a historical clock offset sequence. This model periodically collects the time offset value between the vehicle-mounted terminal and the cloud, which can be achieved through synchronization messages or standard protocol interactions, forming an offset sequence buffer. An LSTM network is used to model the deviation trend, predicting the estimated offset between the current NTP time and PTP time in real time. When the TCP packet arrives at the cloud, the system obtains the received time T5 and calculates the corresponding corrected timestamp based on the estimated offset to compensate for the systematic differences between different time domains. Finally, the generated timestamp T5 is written into the image frame metadata, maintaining a unified time domain with T1 to T4.

[0098] A4: The cabin receives message data from the cloud and marks it as T6 and T7 in sequence;

[0099] Specifically, in this embodiment, the cabin receives image frame data streams forwarded from the cloud via a standard RTSP or TCP client interface. Upon receiving a complete image frame and writing it to the local receive buffer, the system marks the receive timestamp T6 by calling the local PTP clock. T6 represents the time when the cabin system first acquires the complete frame content. Subsequently, the cabin passes the image frame to the rendering module for image display. At the moment the image frame is decoded, synthesized, and sent to the display device's frame buffer, the system rendering engine listens for display trigger events and records the display timestamp T7. T7 represents the final output time when the image content is visually presented.

[0100] Furthermore, to improve the accuracy of timestamp recording, the cabin employs a hardware rendering pipeline combined with V-Sync synchronization technology to ensure that the time point marked by T7 accurately reflects the actual frame display trigger cycle. In addition, to support the acquisition of perturbation frame responses at the display layer, the image frame buffering mechanism in this embodiment fully preserves the content of perturbation frames and activates an external camera or photoelectric detection device after T7 to acquire the display response, which is then used to construct the subsequent image response observation matrix.

[0101] Taking the insertion of perturbation frames as an example, please refer to... Figure 4 To understand, Figure 4 The schematic diagram of the perturbation frame principle provided in the embodiment of this application firstly encodes the perturbation frame into multiple non-overlapping delay coding regions based on the feature deconstruction requirements of the transmission and reception stage and the processing stage of delay detection.

[0102] Then, structural perturbation features are added to the time-delay coding region according to the stage characteristics, wherein the structural perturbation features are generated by frequency distribution coding, phase offset coding and block matching coding.

[0103] Finally, the perturbation frames are periodically inserted into the image frames and transmitted alternately with the image frames at a preset ratio.

[0104] Specifically, the purpose of inserting a perturbation frame is to address the problem that traditional timestamps cannot accurately capture the unobservable delays caused by asynchronous queuing and non-explicit scheduling in remote-controlled driving systems. Because data processing may experience uncontrollable transmission interruptions and buffering waits, while system timestamps possess synchronization, they lack verifiable physical responses, resulting in a lack of realism and resolvability in overall latency assessments. Therefore, it is necessary to construct a perturbation structure that can be embedded into the image link, possessing response consistency and signal reconstruction characteristics, as explicit reference information. This allows the cabin end to deduce the path delay experienced by the perturbation frame based on the image's physical response itself, thus compensating for the dynamic behavior within the link that the timestamp mechanism cannot observe.

[0105] In this embodiment, the disturbance frame is divided into three non-overlapping time-delay coding regions. Each region corresponds to a typical processing stage in the remote driving link, such as image acquisition, image encoding, data packetization and transmission, etc. The disturbance frame is encoded using a significant structural interference signal, and different types of disturbance features are introduced into different regions, as detailed below:

[0106] The first region corresponds to the image acquisition and preprocessing stage (from the vehicle-side image sensor output to the encoder input), and employs a frequency distribution coding structure. This structure consists of high-frequency / low-frequency modulation stripes, and its spectral characteristics are fixed after the sensor output and remain almost unchanged throughout the entire image processing chain, stably reflecting the signal morphology at the image acquisition time point. Therefore, the response displacement observed in this region at the cabin end can accurately calibrate the time span from T1 to T2.

[0107] The second region corresponds to the image encoding and network packetization stage (the vehicle-side controller encodes and packets the image and pushes it to the cloud), employing a phase-shift encoding structure. This structure is constructed using a combination of stripes or interference patterns with preset phase shifts, preserving the phase relationship even after compression and network transmission. By calculating the phase shift recovery amount in the image, the vehicle-side controller can extract the temporal changes that may occur between T2 and T5 due to buffer accumulation, asynchronous transmission, or retransmission, thus reflecting the actual propagation time difference between the image leaving the transmitter and arriving at the cloud.

[0108] The third region corresponds to the cloud forwarding and cabin-end receiving and display stages, employing a block-matching coding structure. This structure embeds a set of high-contrast structural tiles into the perturbed frame. Its characteristic is that after network transmission and cabin-end rendering, it is affected by data reassembly, decoding order, and output queuing, resulting in slight image position shifts or display temporal response misalignments. The cabin end extracts the response time characteristics of this region using inter-frame differential and block-matching methods, capturing frame drift caused by display queuing and processing delays during T5 to T7. A preset ratio ensures link probe density and system load balance.

[0109] Furthermore, the aforementioned disturbance frame design is not a static structure, but rather allows for dynamic parameter adjustment in conjunction with the link model. The layout, frequency selection, and phase control parameters of the disturbance regions can all be set in real-time through the controller parameter table to match the link response characteristics under different test conditions, thereby improving signal reconstruction accuracy. In actual cabin-side response observation, a camera or synchronous acquisition module can accurately record the response time sequence of the disturbance frames on the screen. By constructing an image response observation matrix and performing sparse transformation and joint inversion, the actual response time of each structural disturbance can be identified, and the physical delay of each stage in the link can be deduced. It should be noted that the specific number of regions can be set according to actual needs.

[0110] Please see Figure 5 The figure is a schematic flowchart of the signal processing and analysis method provided in the embodiment of this application. Figure 5 The method shown can be applied to step S3 of the aforementioned method, and the specific steps are as follows:

[0111] S3.1: Collect the response frame sequence of the disturbance frame in the cabin-end display device, and construct an image response observation matrix based on the response frame sequence;

[0112] Specifically, to achieve precise perception of the physical delays at each stage of the remote control driving link, this embodiment collects response image sequences of disturbance frames in the cabin-end display device and constructs an image response observation matrix for delay inversion. Since the disturbance frame traverses multiple processing modules in the transmission link, it carries the processing and transmission delay information of the entire link when it is displayed and output in the cabin. Therefore, the response characteristics of each stage within the link can be reconstructed by analyzing the image changes during its actual display process.

[0113] In this embodiment, the cabin-end display device outputs image content at a fixed refresh rate (e.g., 60Hz). To ensure acquisition accuracy, an image acquisition device synchronized with the display device is set up, such as an industrial camera with a frame rate of not less than 240fps, fixedly installed in front of the display screen, and synchronized with the display controller via hardware triggering to ensure that each frame of display content can be completely acquired. The acquisition device records the actual image output by the display screen in the form of a frame sequence, which includes normal image frames and periodically inserted disturbance frames. For easy identification, the disturbance frames are pre-coded with a unique structure when inserted, and they have stable and detectable spatial characteristics in the cabin-end display image. Whenever a disturbance frame is decoded and sent to the display buffer of the display device, the camera will continuously capture images within a time window before and after it, recording the entire process of the disturbance frame from its first display to the complete disappearance of the structural graphic. The acquired data is a continuous high frame rate image sequence in three-channel RGB format, with each frame having the same resolution as the display device, and organized in chronological order. Each round of perturbation frame acquisition will generate a set of response frame sequences, which have dynamic changes in perturbation response in the time dimension and maintain a stable spatial structure.

[0114] In one example, when constructing the observation matrix, the position coordinates of each perturbation coding region (e.g., frequency perturbation region A, phase perturbation region B, block perturbation region C) in each group of response frame sequences are extracted based on the original structure template of the perturbation frame. The pixel content of these regions in consecutive frames is then truncated and stacked. Each perturbation region is used as a unit to form a three-dimensional data block composed of several frames of image content, with a size of W×H×T, where W and H represent the spatial dimensions of the region, and T represents the number of acquisition frames.

[0115] Furthermore, assuming the spatial coordinate range of the first perturbation region is (x1, y1) - (x2, y2), the pixel content of this region is extracted in each frame of the image and stacked along the time dimension to form a region response tensor of size (x2–x1) × (y2–y1) × T. Repeating the above operation for all perturbation regions allows the construction of a complete image response observation matrix G, where each channel corresponds to multi-frame response data for one perturbation region.

[0116] Preferably, to improve signal extraction stability, this embodiment also performs preprocessing operations such as grayscale normalization, inter-frame difference enhancement, and regional smoothing filtering on the observation matrix data to eliminate background noise and system acquisition errors, and enhance the response edges of the perturbation structure in the time domain. The constructed image response observation matrix G will serve as the core input of the delay inversion algorithm, which has the ability to physically map the response of the perturbation frame and can support the stage-by-stage and structure-by-structure reconstruction modeling of link delay.

[0117] S3.2: Project the image response observation matrix into the sparse feature space through the sparse transform domain to obtain the sparse representation of the perturbation signal in the multi-scale feature domain;

[0118] Specifically, since each delay coding region of the perturbation frame only produces a significant response at a very few time points in the time domain, the image response observation matrix is ​​essentially a high-dimensional, sparsely distributed, and noise-superimposed data structure. Directly using the observation matrix for delay solving not only results in a huge amount of data and excessively high dimensionality, but also in the sparse response state of the perturbation signal in the observation matrix. Therefore, it is necessary to use the sparse feature transformation method to reduce its dimensionality, compress it, and reconstruct its information.

[0119] In this embodiment, the image response observation matrix is ​​obtained by continuously acquiring perturbation frames output by the display device at a high frame rate, and has a structure with a fixed spatial size and temporal series length. To improve data processing efficiency and feature separability, the system divides the observation matrix according to preset structural regions in the perturbation frames. The image response data of each perturbation coding region is extracted separately to form a set of three-dimensional image response unit data blocks. Each data block represents the continuous response of a specific perturbation structure during the cabin display process, including the visual changes of the structure in multiple frames.

[0120] Furthermore, this embodiment introduces sparse transform technology to map the aforementioned response data blocks to a set of specially designed sparse basis structures. Specifically, for each perturbation structure region, the system selects a sparse transform method that matches the perturbation type: for frequency perturbation regions, an orthogonal transform method based on frequency decomposition is used to extract the principal components of the periodic response; for phase perturbation regions, a multi-level wavelet transform method is used to capture the changing trend caused by local phase shift; for block matching perturbation regions, a sparse dictionary matching method based on texture direction extraction is used to extract image texture drift features.

[0121] Furthermore, during the sparse transformation, the system first reassembles the multi-frame response images of each perturbed structure region along the time axis and projects them into a sparse transformation domain specifically designed for that region. After projection, the original response image is represented as a set of sparse structure characteristic coefficients. The system retains the set of sparse principal components with the highest energy according to a preset threshold and discards sparse response terms with smaller amplitudes and insignificant contributions, ultimately forming the response representation of each perturbed structure in the sparse space.

[0122] Furthermore, to ensure the independence of delay characteristics across different stages of the link, this embodiment introduces a dimensionality suppression strategy in the sparse representation stage. This strategy sets activation windows for the sparse response data of different perturbation structures based on the stage division during perturbation frame structure design, allowing only sparse responses within specific time periods to participate in subsequent processing. For example, the system sets that frequency perturbation structures should only respond in the initial stage of perturbation frame display, phase perturbations in the middle stage, and block matching structures in the final stage; response components exceeding this time range are automatically set to zero. Through this suppression mechanism, the sparse representation of the perturbation frame achieves temporal matching with the link stages, effectively reducing cross-stage interference and improving the accuracy and physical interpretability of subsequent inversion solutions.

[0123] In one example, the specific steps for sparse representation of the perturbation signal are as follows:

[0124] S3.2.1: The image response observation matrix is ​​transformed by an orthogonal sparse transformation operator to construct a response sparse representation tensor in the mapping domain. The orthogonal sparse transformation operator is a preset segmented modulation basis, and its construction maintains structural consistency with the coding structure corresponding to each stage in the perturbation frame.

[0125] S3.2.2: The principal components of the response sparse representation tensor are concentrated in the sparse subspace corresponding to each stage by using a dimensionality suppression function to generate a sparse feature space;

[0126] S3.3: Construct a joint optimization model for delay inversion with sparse representation as the observed variable and propagation delay at each stage as the sparse coefficient, wherein the joint optimization model for delay inversion includes sparse regularization terms and structural consistency constraint terms;

[0127] Specifically, although sparse transformation can compress the image response observation matrix into the temporal activation features of multiple perturbation regions, this sparse representation only indicates where the response occurred, but cannot express which processing stage the response belongs to, or the specific corresponding latency. The fundamental purpose of this application is to deduce from the image response the actual propagation time of the image frame at each stage of the entire link. Therefore, in this embodiment, a mathematically solvable inversion model needs to be constructed to establish a clear causal mapping between the image response features and the processing link stages.

[0128] In this embodiment, during the construction process, the model input is first determined, namely the sparse representation features extracted from the image response observation matrix. These features represent the activation response intensity of multiple perturbation regions on the time axis. Each region originates from a predefined structural coding block in the perturbation frame, and each structural coding corresponds one-to-one with a specific link processing stage. Then, based on this structural design logic, each stage in the link (such as acquisition, encoding, streaming, transmission, reception, and display) is treated as an independent parameter to be estimated, i.e., the possible propagation delay of each stage.

[0129] Furthermore, to establish a solvable relationship, in this embodiment, the aforementioned link stages are considered as a set of independent but time-delay variables with a limited total duration. By analyzing the sparse response of each perturbation region, a preliminary time window range can be obtained, representing when the structural perturbation may have changed. However, since interference (such as illumination, reflection, and display refresh interference) is unavoidable in the image response, the maximum activation point cannot be directly taken as the result. Instead, the entire time window needs to be treated as a fitting object, and a set of processing stage delay variables needs to be found to match the theoretical response pattern generated by it with the actual observed sparse response.

[0130] Furthermore, two constraints are added to the model: First, a sparse regularization constraint is added to control the distribution of the solved propagation delay as concentrated as possible on the time axis, allowing only a few activation points to exist, thus preventing delay value drift or duplicate coverage; second, a structural consistency constraint is introduced, which forces each perturbation structural region to only respond to its own link stage. This structural correspondence is determined during the design of the perturbation frame, and therefore can be loaded into the model as a hard constraint to prevent activation overlap between stages during the optimization process.

[0131] As an example, to ensure the model is solvable in actual deployment, this embodiment uses an iterative modeling method: first, based on the activation trend of the perturbation region in the sparse representation, an initial estimate of the delay range for each stage is given; then, combined with the sparse activation values, these estimates are gradually narrowed and adjusted, with the stage position updated in each iteration until the overall error is reduced to below a set threshold. Throughout this iteration, the one-to-one correspondence between the perturbation structure and the link stages remains unchanged, ensuring that the final output delay of each element originates from its designated structural region.

[0132] Furthermore, this model construction method can deconstruct the originally mixed image response features into a set of delay variables corresponding to the link stages, and eliminate problems such as signal crossover and inter-frame interference by gradually fitting and deconstraining.

[0133] S3.4: Based on the aforementioned delay inversion joint optimization model, perform joint inversion solution on the sparse representation of each stage to obtain the sparse coefficient distribution, and use the sparse coefficient distribution as the delay feature;

[0134] Specifically, after the model is built, in order to obtain the time delay results of each stage of the link, it is necessary to jointly solve the sparse feature tensor and output the optimal estimate of the sparse coefficients. In this embodiment, an accelerated sparse solver based on FISTA is selected to perform parallel inversion of sparse coefficient vectors in multiple regions. A multi-threaded decoupling method is used to construct gradient descent paths for each perturbed structural region and update the sparse delay components of the corresponding stage in real time.

[0135] During the iteration process, the system dynamically evaluates the error convergence rate, automatically adjusts the sparsity threshold to improve the stability and robustness of the solution, and outputs the final sparse coefficient distribution map when the convergence condition is met. Each non-zero position in this distribution corresponds to a disturbance response activation time point, that is, the physical response time of a specific stage in the link. By matching the difference with the system timestamps T1 to T7, the offset of the actual propagation delay of the link can be derived, achieving high-precision correction of the timestamp delay.

[0136] Furthermore, the sparse coefficient distribution obtained through joint solution not only provides a quantitative estimate of the multi-stage delay of the link, but also enables the construction of a delay distribution map at the level of the response observation matrix. This map can be used to identify performance issues such as bottlenecks, asymmetric loads, or dynamic jitter in the link, thereby giving traditional remote control links physical measurability and response reversibility.

[0137] In one example, the specific steps for solving the joint inversion are as follows:

[0138] S3.4.1: Perform parallel solving on the sparse representations corresponding to each time-delay coded region in the perturbed frame to obtain the sparse coefficient vector;

[0139] Specifically, each delay-coded region in the perturbation frame essentially represents an independent processing or transmission stage in the link. After sparse transformation, each region acquires sparse activation features over a time dimension. To determine the response time location of the corresponding stage, i.e., the propagation delay, these sparse activation features must be parameterized. However, since the stages are essentially logically connected, have independent responses, and non-overlapping delay distributions, a parallel partitioning strategy should be adopted during the solution process. Solution paths should be constructed separately for each perturbation region to improve model computational efficiency and avoid cross-interference of solution errors.

[0140] In this embodiment, the perturbation frame is divided into multiple time-delay coding regions. Each region, after sparse transformation, generates a sparse activation feature, representing the activation level of that region within the time window. When constructing the inversion model, each region is treated as an independent solution unit, and its required propagation delay value is solved in parallel within a unified optimization framework. Specifically, the inversion model employs a sparse fitting strategy. Given the sparse response of a region, the propagation delay estimate most likely explaining the response behavior is derived. Formally, this is equivalent to deriving the triggering position that generated the mode from the sparse activation mode, i.e., the time point at which the perturbation structure experiences the physical response.

[0141] S3.4.2: Iterative optimization converges the sparse coefficient vector of each region to the optimal solution domain;

[0142] Specifically, the sparse solution process is constrained by non-ideal factors in the perturbation response, such as blurring of the perturbation signal caused by image display refresh, activation errors caused by external noise, and misjudgments caused by inter-frame interference, which may cause the initial solution to deviate from the true response time. Therefore, it is necessary to introduce an iterative optimization mechanism to perform multiple rounds of correction and convergence on the sparse coefficient vector of the initial solution in order to obtain a solution that approximates the real physical behavior.

[0143] In this embodiment, a two-stage optimization mechanism based on iterative threshold updates and response alignment reconstruction is used. In the initial stage, a predicted response pattern is constructed based on the activation positions of the sparse coefficients, and similarity calculations are performed with the actual observed sparse vectors to generate error feedback. Subsequently, a contraction stage is entered, where the parameter space is gradually compressed and approximated by adjusting non-zero positions in the sparse vectors, adding or deleting activation elements, and updating the response window position. Each iteration includes a response error minimization objective and a structural consistency constraint recalculation mechanism to ensure that the solution process maintains a good fit to the true signal without deviating from the semantic boundaries of the perturbed structure.

[0144] S3.4.3: After iterative optimization, the non-zero components in the sparse coefficient vector are stage-calibrated, and the response time represented by their corresponding positions is taken as the sparse coefficient distribution of each stage.

[0145] Specifically, the non-zero components retained in the sparse coefficient vector represent stable and significant response activations in the perturbation region at a certain point in time. According to the design principles of perturbation frames, each perturbation structure corresponds to only one processing stage in the link. Therefore, the location of the non-zero component is equivalent to the moment the physical response occurs in that stage, i.e., the actual propagation delay of that processing stage. By assigning both structural affixation and temporal location, a complete link delay distribution can be constructed.

[0146] In this embodiment, based on a pre-established one-to-one mapping relationship between the perturbation structure and the link stage, the sparse coefficient vector of each perturbation region is labeled as the response source of a certain fixed processing stage. After iterative optimization, the system extracts the positions of the non-zero components in the vector and converts the frame order of this position in the response observation window into time units as the propagation delay of that stage. Since all sparse responses come from the physical image acquisition at the cabin end, reflecting the actual time point at which the perturbation becomes visible at the visual output level, it is more physically measurable and stable than the system's internal timestamps.

[0147] Furthermore, this sparse coefficient distribution set can not only be directly used to correct the timestamp delay difference in system calculations, but also for engineering applications such as constructing dynamic response characteristic maps of links, identifying stage bottlenecks, and assisting in scheduling priority optimization.

[0148] As a preferred embodiment, this preferred embodiment proposes a link load response method based on periodic feedback stress test frame triggering. Its core idea is to not rely on disturbance frames, but to inject feedback stress test frames with trigger signals into the image data stream, and to use the system end-to-end feedback mechanism to monitor its physical backhaul delay in real time, thereby establishing a stress test response spectrum of whether the link processing at each stage is blocked.

[0149] Specifically, the system controller inserts feedback stress test frames with specific encoding characteristics into the image stream at fixed intervals (e.g., every 20 frames). These frames do not structurally alter the content; instead, they embed stress test identifiers in their image metadata and message headers. Once the frame arrives at the cabin end and is processed, the cabin-end system automatically generates a feedback acknowledgment message, which is returned to the vehicle end via the reverse link. The controller calculates the physical response time of each hop in the link based on the difference between this feedback return delay and the current normal image frame processing rhythm. This time includes: the time the frame is queued, the time delayed by protocol buffering and packaging, and the time delayed by operating system scheduling—that is, the ghost latency that attempts to probe but the timestamp cannot cover. The triggering and return of the stress test frames are based on real links and real processing flows, not analog signals, thus their feedback delay has high physical realism. Furthermore, the stress test frame insertion frequency is adjustable, and its triggering behavior can be embedded into existing network protocol frameworks without modifying image encoding methods or business content, resulting in high system compatibility. More importantly, the feedback stress testing mechanism essentially provides a means for external entities to proactively inquire about the link status, avoiding the structural defect of the system relying solely on internal data points to provide its own information.

[0150] In another preferred scheme, in order to accurately capture the link-level processing latency in the remote driving system, a multi-clock trajectory mapping mechanism based on cross-layer frame boundary alignment is proposed. By detecting the image frame boundary at each key node of the link and synchronously recording the local system clock state when the boundary occurs, multi-stage latency inference is finally achieved through trajectory alignment.

[0151] Specifically, at any stage of image frame acquisition, encoding, packaging, decoding, rendering, and display, there will be captureable frame boundary events (such as frame synchronization signals, header writing, decoding start, frame rendering submission, etc.). These events are naturally the trigger points for the start or end of the processing behavior. Therefore, as long as the occurrence time of these boundaries is accurately recorded at each stage, the propagation trajectory of the entire link can be reconstructed.

[0152] In this embodiment, the vehicle-side encoder records the first frame boundary marker and timestamp before the image frame enters motion estimation; the controller marks the second frame boundary event at the packet start position; the cloud-based unpacking module records the third frame boundary before the decoder starts; and the cabin-side records the final boundary before the image is submitted for display. By aligning the clocks of these key frame boundary points and combining them with the timing systems of each node, a multi-segment trajectory of the complete image frame propagation link is constructed. To compensate for the problem of cross-node clock inconsistency, a trajectory mapping function is introduced in the scheme. The average of the stable propagation time between historical frames is used to align the trajectories of each stage, so that even when the clocks are not strictly synchronized, the actual dwell and transfer process of the frame in the system can be restored. The system can finally output a set of trajectory time periods for each stage, thereby separating the processing delay of each segment. Without introducing any new frames, changing the image content, or relying on external detection signals, the delay modeling is performed only by extracting the clocks of the original frame events in the system, which has a completely non-intrusive system friendliness. At the same time, by using trajectory alignment rather than strict timestamp comparison, it has a stronger cross-clock domain adaptability, which is a low-coupling compensation mechanism for scenarios where the accuracy of traditional timing mechanisms is insufficient.

[0153] As an example, this embodiment provides a latency testing system for a remote-controlled driving system, the system comprising:

[0154] The acquisition module is used to acquire image frames and insert preset disturbance frames into the image frames. The image frames are transmitted and processed sequentially via the vehicle end, cloud end, and cabin end.

[0155] The timestamp recording module is used to record the timestamp information of image frames at each transmission and reception stage and processing stage at the vehicle end, cloud end, and cabin end, respectively.

[0156] The latency calculation module is used to determine the latency of each transmission and reception stage and processing stage based on the timestamps of each stage.

[0157] The response acquisition module is used to acquire the image response signal of the disturbance frame in the cabin-end display device;

[0158] The signal analysis module is used to perform signal processing analysis on the image response signal and extract delay features that reflect the propagation delay at each stage;

[0159] The delay correction module is used to correct the delay calculated based on the timestamp according to the delay characteristics.

[0160] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A latency testing method for a remote driving system, applied to test a remote driving system, wherein the remote driving system includes a vehicle end, a cloud end, and a cabin end, characterized in that, The latency testing method includes: The timestamps of the collected image frames are acquired and transmitted and processed at the vehicle end, cloud end, and cabin end, wherein preset disturbance frames are inserted into the image frames; The latency of each sending and receiving stage and processing stage is determined based on each timestamp; The image response signal of the disturbance frame at the cabin end is processed and analyzed to extract delay features, and the delay calculated based on the timestamp is corrected according to the delay features; Before obtaining each timestamp, the latency testing method further includes: A time synchronization module based on the PTP protocol is configured for the vehicle end and the cabin end respectively, and a time synchronization module based on the NTP protocol is configured for the cloud. A clock deviation prediction model is constructed based on historical deviation sequences. The dynamic offset between cloud NTP time and other terminal PTP time is predicted by the clock deviation prediction model, and the cloud time reference is corrected according to the dynamic offset. Before acquiring the timestamps of image frames being transmitted, received, and processed at the vehicle end, the latency testing method further includes: In the vehicle-side image acquisition and encoding stage, image frames are acquired by the vehicle-side image sensor and input to the motion estimation stage of the encoder. Before being input to the encoder, a perturbation frame with preset characteristics is inserted into the image frame. Based on the system time of PTP time synchronization, a timestamp T1 is injected into the image frame through the encoder's software development interface; In the image processing flow, after the controller at the vehicle end receives the image frame, it marks the image reception timestamp T2, and after the processing is completed, it marks the processing completion timestamp T3. The processed image frames are encapsulated into TCP packets for streaming, and the streaming timestamp T4 is marked at the same time. The step of inserting a perturbation frame with preset characteristics into the image frame before it is input to the encoder includes: Based on the feature deconstruction requirements of the transmission and reception and processing stages of delay detection, the perturbation frame is encoded into multiple non-overlapping delay coding regions. Structural perturbation features are added to the time-delay coding region according to the stage characteristics, wherein the structural perturbation features are generated by frequency distribution coding, phase offset coding and block matching coding. The disturbance frames are periodically inserted into the image frames and transmitted alternately with the image frames at a preset ratio.

2. The delay testing method for a remote-controlled driving system according to claim 1, characterized in that, Before acquiring the timestamps of image frames being sent, received, and processed in the cloud, the latency testing method further includes: The cloud receives TCP packets from the vehicle and performs unpacking processing; The received timestamp is dynamically corrected according to the clock deviation prediction model to generate a corresponding cloud-corrected timestamp T5. The cloud-corrected timestamp T5 is written into the RTP protocol extension header and the message data is forwarded to the cabin via the RTSP protocol.

3. The delay testing method for a remote-controlled driving system according to claim 2, characterized in that, Before acquiring the timestamps of image frames being transmitted, received, and processed at the cabin end, the latency testing method further includes: The cabin receives the message data forwarded by the cloud through the communication equipment, and unpacks and extracts the timestamps T1 to T5 from the message; The marker module receives the current timestamp T6 of the image frame; The image frame is rendered and output through a display device, and the image display timestamp T7 is marked.

4. The delay testing method for a remote-controlled driving system according to claim 1, characterized in that, Signal processing analysis is performed on the image response signal of the disturbance frame at the cabin end to extract delay features, including: The response frame sequence of the disturbance frame in the cabin-end display device is collected, and an image response observation matrix is ​​constructed based on the response frame sequence; The image response observation matrix is ​​projected onto the sparse feature space through the sparse transform domain to obtain a sparse representation of the perturbation signal in the multi-scale feature domain. Construct a joint optimization model for delay inversion with sparse representation as the observed variable and propagation delay at each stage as the sparse coefficient, wherein the joint optimization model for delay inversion includes sparse regularization terms and structural consistency constraint terms; According to the aforementioned delay inversion joint optimization model, the sparse representations of each stage are jointly inverted and solved to obtain the sparse coefficient distribution, which is then used as the delay feature.

5. The delay testing method for a remote-controlled driving system according to claim 4, characterized in that, The step of projecting the image response observation matrix into the sparse feature space through the sparse transform domain includes: The image response observation matrix is ​​transformed by an orthogonal sparse transformation operator to construct a response sparse representation tensor in the mapping domain. The orthogonal sparse transformation operator is a preset segmented modulation basis, and its construction maintains structural consistency with the coding structure corresponding to each stage in the perturbed frame. The principal components of the response sparse representation tensor are concentrated in the sparse subspace corresponding to each stage by controlling the dimensionality suppression function, so as to generate a sparse feature space.

6. The delay testing method for a remote-controlled driving system according to claim 4, characterized in that, The joint inversion solution of the sparse representations at each stage includes: The sparse representations corresponding to each time-delay coding region in the perturbation frame are solved in parallel to obtain the sparse coefficient vector. The sparse coefficient vectors of each region are converged to the optimal solution domain through iterative optimization. After iterative optimization, the non-zero components in the sparse coefficient vector are calibrated in stages, and the response time represented by their corresponding positions is used as the sparse coefficient distribution of each stage.

7. A latency testing system for a remote-controlled driving system, used to implement the latency testing method for a remote-controlled driving system as described in any one of claims 1-6, characterized in that, The system includes: The acquisition module is used to acquire image frames and insert preset disturbance frames into the image frames. The image frames are transmitted and processed sequentially via the vehicle end, cloud end, and cabin end. The timestamp recording module is used to record the timestamp information of image frames at each transmission and reception stage and processing stage at the vehicle end, cloud end, and cabin end, respectively. The latency calculation module is used to determine the latency of each transmission and reception stage and processing stage based on the timestamps of each stage. The response acquisition module is used to acquire the image response signal of the disturbance frame in the cabin-end display device; The signal analysis module is used to perform signal processing analysis on the image response signal and extract delay features that reflect the propagation delay at each stage; The delay correction module is used to correct the delay calculated based on the timestamp according to the delay characteristics.

Citation Information

Patent Citations

  • Video processing method and device, electronic equipment and readable storage medium

    CN113286194A

  • Remote control anti-delay video transmission method based on future scene generation network

    CN119254998A