Time delay test method and system of remote control driving system
By inserting disturbed frames into the remote control driving system and performing signal processing and analysis, combining sparse feature space and delay inversion optimization model, the problem of large delay measurement error in the prior art is solved, and accurate identification and correction of link delay is achieved, and measurement accuracy and availability are improved.
Patent Information
- Application Number
- CN202510672487.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the existing remote driving system, the time stamp-based delay measurement mechanism cannot accurately reflect the real time of data processing or transmission, resulting in large invisible errors in delay measurement. Especially under non-explicit processing behaviors such as asynchronous scheduling and data cache queueing, it is impossible to achieve microsecond accurate identification and correction.
By inserting preset perturbation frames into the image frame, the processing and transmission timestamps of each node are collected, and signal processing and analysis is performed at the cabin. Combining the sparse feature space and delay inversion optimization model, the delay characteristics of each stage are extracted, and the link delay calculated based on timestamps is corrected.
It realizes effective identification and correction of ghost delays in the link, improves the accuracy and engineering availability of delay measurement, and can identify non-causal delays such as implicit scheduling delays and queue waiting, achieving a microsecond delay correction effect.
Smart Images

Figure CN120567348A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time delay testing, and in particular to a time delay testing method and system for a remote control driving system. Background Art
[0002] In existing remote-controlled driving systems, a timestamp-based latency measurement mechanism is commonly used to achieve real-time closed-loop control of vehicle-side image data transmission and cabin-side control commands. This mechanism uses time stamps at key nodes such as image acquisition, encoding, transmission, decoding, and display, and calculates link latency based on the time difference between each node. However, in actual deployments, due to non-explicit processing behaviors such as asynchronous scheduling, data caching, and protocol-layer packet reassembly, timestamps cannot accurately reflect the physical moment when data is actually processed or sent, resulting in significant invisible errors.
[0003] This application designs a time delay testing method and system for a remote control driving system. Summary of the Invention
[0004] The present invention addresses the shortcomings of existing technologies and provides a method and system for measuring latency in remote-controlled pilot systems. This method inserts a pre-set perturbation frame into an image frame, collects the processing and transmission timestamps of the image frame at each node, and performs signal processing and analysis on the image response signal of the perturbation frame at the cabin end. This extracts delay features at each stage, which are used to correct link latency calculated based on timestamps. By constructing a sparse feature space and a delay inversion optimization model, this method achieves precise identification and microsecond-level correction of multi-stage processing delays.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A time delay test method for a remote control driving system is applied to test a remote control driving system including a vehicle side, a cloud side, and a cabin side. The time delay test method includes:
[0007] Obtaining respective timestamps of transmission, reception, and processing of the captured image frames at the vehicle side, the cloud side, and the cabin side, wherein a preset disturbance frame is inserted into the image frames;
[0008] Determine the latency of each sending, receiving, and processing stage based on each timestamp;
[0009] Signal processing and analysis are performed on the image response signal of the disturbance frame at the cabin end to extract delay features, and the time delay calculated based on the timestamp is corrected according to the delay features.
[0010] Before obtaining each timestamp, the delay testing method further includes:
[0011] The vehicle side and the cabin side are respectively configured with a timing module based on the PTP protocol, and the cloud side is configured with a timing module based on the NTP protocol;
[0012] A clock deviation prediction model is constructed based on the historical deviation sequence. The dynamic offset between the cloud NTP time and the other end PTP time is predicted by the clock deviation prediction model, and the cloud time base is corrected according to the dynamic offset.
[0013] Before obtaining the timestamps of each image frame being sent, received, and processed on the vehicle side, the latency testing method further includes:
[0014] In the vehicle-side image acquisition and encoding stage, image frames are acquired by the vehicle-side image sensor and input into the motion estimation stage of the encoder, wherein a disturbance frame with preset characteristics is inserted into the image frame before input into the encoder;
[0015] Based on the PTP time-serving system time, inject the timestamp T1 into the image frame through the software development interface of the encoder;
[0016] In the image processing process, the controller on the vehicle side marks the image receiving timestamp T2 after receiving the image frame, and marks the processing completion timestamp T3 after the processing is completed;
[0017] The processed image frames are encapsulated into TCP packets for streaming, and the streaming timestamp T4 is marked.
[0018] Before obtaining the timestamps of each image frame being sent, received, and processed in the cloud, the latency testing method further includes:
[0019] The cloud receives the TCP message from the vehicle and unpacks it;
[0020] Dynamically correct the received timestamp according to the clock deviation prediction model to generate the corresponding cloud-corrected timestamp T5;
[0021] The cloud-side correction timestamp T5 is written into the RTP protocol extension header, and the message data is forwarded to the cabin end via the RTSP protocol.
[0022] Before obtaining the timestamps of each image frame being sent, received, and processed at the cabin end, the latency testing method further includes:
[0023] The cabin receives the message data forwarded by the cloud through the communication equipment, and unpacks and extracts the timestamps T1 to T5 in the message;
[0024] Marking the current timestamp T6 of the image frame received by the cabin end;
[0025] The image frame is rendered and outputted through a display device, and an image display timestamp T7 is marked.
[0026] The inserting a disturbance frame having preset characteristics into the image frame before inputting into the encoder comprises:
[0027] Based on the feature deconstruction requirements of the transmission and reception phase and the processing phase of delay detection, the disturbance frame is encoded into multiple non-overlapping delay coding regions;
[0028] adding a structural disturbance feature in the time delay coding region according to the phase characteristics, wherein the structural disturbance feature is generated by frequency distribution coding, phase offset coding and block matching coding;
[0029] The disturbance frame is periodically inserted into the image frame and transmitted alternately with the image frame at a preset ratio.
[0030] Performing signal processing and analysis on the image response signal of the disturbance frame at the cabin end to extract delay features includes:
[0031] Acquiring a response frame sequence of the disturbance frame in a cabin-side display device, and constructing an image response observation matrix according to the response frame sequence;
[0032] Projecting the image response measurement matrix into a sparse feature space through a sparse transform domain to obtain a sparse representation of the disturbance signal in a multi-scale feature domain;
[0033] Constructing a delay-inversion joint optimization model with sparse representation as observation variable and propagation delay of each stage as sparse coefficient, wherein the delay-inversion joint optimization model includes a sparse regularization term and a structural consistency constraint term;
[0034] According to the delay inversion joint optimization model, the sparse representation of each stage is jointly inverted and solved to obtain a sparse coefficient distribution, which is used as a delay feature.
[0035] The projecting the image response measurement matrix to a sparse feature space through a sparse transform domain includes:
[0036] The image response measurement matrix is transformed by an orthogonal sparse transformation operator to construct a response sparse expression tensor in the mapping domain, wherein the orthogonal sparse transformation operator is a preset piecewise modulation basis, and its construction maintains structural consistency with the coding structure corresponding to each stage in the perturbation frame;
[0037] The principal components of the response sparse expression tensor are controlled to be concentrated in the sparse subspace corresponding to each stage by a dimension suppression function to generate a sparse feature space.
[0038] The joint inversion solution of the sparse representation of each stage includes:
[0039] The sparse representation corresponding to each delay coding region in the disturbance frame is solved in parallel to obtain a sparse coefficient vector;
[0040] Through iterative optimization, the sparse coefficient vectors of each region are converged to the optimal solution domain;
[0041] After iterative optimization, the non-zero components in the sparse coefficient vector are calibrated in stages, and the response moments represented by the corresponding positions are used as the sparse coefficient distribution of each stage.
[0042] A time delay test system for a remote control driving system, the system comprising:
[0043] An acquisition module is used to acquire image frames and insert preset disturbance frames into the image frames. The image frames are sequentially transmitted and processed via the vehicle side, the cloud side, and the cabin side.
[0044] A timestamp recording module is used to record the timestamp information of the image frame at each sending and receiving stage and processing stage on the vehicle side, cloud side and cabin side respectively;
[0045] The delay calculation module is used to determine the delay of each receiving and sending stage and processing stage based on the timestamps of each stage;
[0046] A response acquisition module, configured to acquire an image response signal of the disturbance frame in a cabin-side display device;
[0047] a signal analysis module, configured to perform signal processing and analysis on the image response signal and extract delay features reflecting propagation delays at each stage;
[0048] The delay correction module is used to correct the delay calculated based on the timestamp according to the delay characteristics.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] The present invention inserts disturbance frames with stage-specific structural characteristics into image frames, and combines image response signal acquisition and sparse feature analysis at the cabin end to construct a delay inversion model that can be mapped to each link processing stage, thereby achieving effective identification and correction of ghost delays in the link. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0052] Figure 1 This is a schematic diagram of an exemplary application scenario of an embodiment of the present invention;
[0053] Figure 2 This is a flow chart of a method for testing the time delay of a remote control driving system according to an embodiment of the present invention;
[0054] Figure 3 A schematic diagram of a timestamp marking method according to an embodiment of the present invention;
[0055] Figure 4 This is a schematic diagram of the disturbance frame principle according to an embodiment of the present invention;
[0056] Figure 5 Schematic diagram of the signal processing and analysis method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0058] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It will be understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0059] See also Figure 1 , which is a schematic diagram of an exemplary application scenario provided in an embodiment of the present application.
[0060] like Figure 1 As shown in the figure, the application scenario is a remote driving architecture consisting of the vehicle side, cloud side, and cabin side. The three are interconnected through a transmission network and are connected to the timing mechanism to achieve multi-node time synchronization.
[0061] Figure 1 The figure shows the vehicle-side access to the PTP (Precision Time Protocol) timing mechanism, which specifically includes a GMSL camera for acquiring raw image frames and periodically inserting disturbance frames into the image stream, and a vehicle-side controller for receiving, processing, encoding, and streaming the acquired image frames, while recording timestamps T1 to T4 at different processing stages.
[0062] Figure 1 The figure shows the cloud access NTP (Network Time Protocol) timing mechanism, which specifically includes an unpacking module for receiving TCP messages from the vehicle and extracting image frames, and an error compensation module for correcting the time base difference between the cloud time and PTP time based on the clock deviation prediction model, and generating a corrected timestamp T5.
[0063] Figure 1 It shows that the cabin side is also connected to the PTP timing mechanism, which specifically includes an unpacking module for receiving and parsing the image data stream from the cloud, and recording the display time timestamp T6, a display output and rendering module for rendering and outputting the received image frame, and recording the display time timestamp T7, a disturbance frame decoding module for decoding and extracting features of the disturbance structure in the image frame, and a timestamp delay correction module for correcting the delay calculated based on the timestamp based on the delay characteristics of the disturbance frame in the image response.
[0064] In one example, Figure 1 The transmission network shown can be set up through a 5G network to reduce transmission latency and achieve real-time control of unmanned vehicles.
[0065] In one example, a cloud server acts as a relay to establish a data connection between the remote-control vehicle and the remote-control cabin. A PTP timing synchronization system aligns the time base between the two terminals, and a clock deviation prediction model is used in the cloud to compensate for timing errors. The remote-control vehicle, cloud server, and remote-control cabin each inject timestamps via a software SDK, ultimately calculating transmission delay. The remote-control vehicle includes an onboard controller for sending and receiving data, while the remote-control cabin connects to the cabin display via an HDMI cable. The cloud server also provides RTSP transceiver functionality.
[0066] In one example, T2-T1 represents the camera GMSL transmission delay;
[0067] T3-T2 represents the vehicle-side image processing delay;
[0068] T4-T3 represents the encoding and streaming delay at the vehicle end;
[0069] T5-T4 represents the transmission delay from vehicle to cloud;
[0070] T6-T5 represents the transmission delay from the cloud to the cabin;
[0071] T7-T6 represents the cabin display delay;
[0072] T7-T1 represents the full-link video stream latency of the remote control driving system.
[0073] Next, in conjunction with the accompanying drawings, a time delay testing method for a remote control driving system provided in an embodiment of the present application is introduced.
[0074] See also Figure 2 , which is a flow chart of a time delay testing method for a remote control driving system provided in an embodiment of the present application. Figure 2 The method shown can be applied to test a remote driving system, which includes a vehicle side, a cloud side, and a cabin side. Figure 2The method shown includes the following steps S1-S4, and the specific steps are as follows:
[0075] S1: Obtaining the timestamps of each time when the captured image frame is sent, received, and processed on the vehicle side, cloud side, and cabin side;
[0076] In this embodiment, the vehicle captures image frames via a GMSL camera and inserts a disturbance frame before the image enters the encoding process to embed a reference signal for subsequent processing delay analysis. Subsequently, the image frames are timestamped with T1, T2, T3, and T4 during the image sensor encoding phase, controller reception phase, and image streaming phase, respectively. Upon receiving the image frame message, the cloud unpacks it and calculates the NTP to PTP time offset based on a historical time deviation model. The corrected cloud reception timestamp is T5. The cabin side records timestamps T6 and T7 in sequence during the process of receiving, rendering, and displaying the image frame. By constructing the timestamp sequence T1 to T7, the processing and transmission trajectory of the image frame throughout the remote control link can be fully modeled, facilitating subsequent delay decomposition and location, thereby resolving the technical flaw of traditional delay measurement, which focuses only on the start and end points while ignoring link details.
[0077] S2: Determine the latency of each receiving and sending stage and processing stage based on each timestamp;
[0078] In this embodiment, the timestamp sequence T1 to T7 is used to calculate the latency of each stage of the image frame, from vehicle-side acquisition to processing (T1-T3), vehicle-side to cloud-side transmission (T4-T5), cloud-side to cabin-side transmission (T5-T6), and cabin-side rendering and display (T6-T7). This allows for a refined estimation of the latency distribution of each processing sub-stage in the entire remote video link. This allows for the identification of potential bottlenecks or abnormal waits at any stage in the system, thereby improving the latency visualization capabilities of the remote control system.
[0079] S3: Perform signal processing and analysis on the image response signal of the disturbance frame at the cabin end to extract delay features;
[0080] In this embodiment, the cabin captures the disturbance frame image sequence output by the display screen and constructs an image response observation matrix. Based on the design structure of each delay coding region in the disturbance frame, an orthogonal sparse transform operator is used to map the response observation matrix to a sparse feature space. Subsequently, based on the correspondence between the sparse representation and the predefined disturbance structure, a delay inversion joint optimization model is constructed, and the sparse coefficient distribution of each coding region is extracted through an iterative sparse reconstruction method. The coefficient position shows a one-to-one correspondence with the delay of each processing stage. Through the above-mentioned disturbance response signal modeling and joint solution mechanism, it is possible to actively identify "non-causal delays" such as implicit scheduling delays and queue waiting that cannot be detected by traditional timestamps, breaking through the technical blind spots of existing timestamp-based measurement.
[0081] S4: Modify the delay calculated based on the timestamp according to the delay characteristics;
[0082] In this embodiment, the obtained sparse delay features are used to correct the delay results of the timestamp calculation. The timestamp difference at each stage is combined with the physical response delay extracted from the sparse coefficients to perform error compensation and boundary smoothing, thereby constructing a high-precision delay estimate that is closer to the actual physical link. This correction process effectively eliminates error offsets caused by multi-level caches, operating system scheduling, and asynchronous transmission, improving the engineering usability and microsecond-level accuracy of the measurement value, and addressing the measurement bias caused by unobservable delays in existing pure timestamp mechanisms.
[0083] While timestamp-based end-to-end latency measurement is a relatively mature solution for remote driving links, its core assumption remains that timestamps accurately represent the actual moment of data processing or transmission. However, this assumption fails in real-world environments with multi-level asynchronous transmission, buffer queues, and uncontrollable module scheduling. For example, in vehicle-side video streaming, even if an image frame is timestamped upon encoding, the actual message transmission time may be offset by tens of milliseconds due to network queuing and system scheduling jitter. This ghost latency cannot be directly captured using timestamps, resulting in imperceptible errors in the calculated link latency even when time points T1 to T7 are measured. Existing technologies have also attempted to optimize timestamp strategies through high-frequency sampling, clock synchronization, or a combination of hardware and software. However, these approaches essentially remain at the optimization level of the timestamp itself and fail to overcome the logical bottleneck of relying on the system to describe itself. The end result is that in remote control scenarios with high real-time requirements, the existing timestamp measurement method is unable to support microsecond-level delay assessment and precise segment identification, especially when there is an opaque intermediate layer scheduling mechanism in the link, the measured value deviates seriously from the physical reality.
[0084] For example, in the vehicle-side video image processing process, after encoding, the image frame is streamed by the controller through the network protocol stack. Typically, to calculate the processing and transmission latency of the frame, a timestamp is added after encoding is completed or before the data is submitted to the network interface. However, in practice, this timestamp does not accurately reflect the actual time the image frame was sent. This is because, at an engineering level, after data is submitted from the application layer to the network interface, it typically enters the operating system's kernel buffer and is then asynchronously scheduled by the network protocol stack. This scheduling process involves multiple non-deterministic factors such as queuing, cache merging, prioritization, and system interrupt response. The final data outbound transmission often lags by tens of milliseconds from the time the timestamp is recorded. Furthermore, this process lacks clear callbacks or event hooks that can be invoked by user space. Therefore, even if attempts are made to timestamp the sending behavior, this only marks the data submission action, not the actual data transmission behavior, resulting in inevitable discrepancies between the timestamp record and the actual physical behavior.
[0085] Furthermore, existing operating systems and network protocol stacks don't provide any standard mechanisms for accurately recording data transmission times without modifying driver-level code or introducing external hardware for data collection. Even forcible intervention through driver-level instrumentation or packet capture mechanisms still presents engineering challenges such as high system stability risks, poor cross-platform adaptability, and high resource consumption.
[0086] To this end, in this embodiment, a disturbance frame is introduced as an active physical reference signal to construct a visual, responsive, and analyzable embedded probe mechanism in the image frame transmission link, thereby breaking through the limitations of the timestamp observation system. In specific implementation, the disturbance frame adopts a structured design method to embed disturbance features with different coding modes into multiple areas of the image frame. These disturbance areas will not be tampered with in each processing stage (such as acquisition, processing, encoding, and transmission), but will produce different time domain response characteristics due to system scheduling and queuing behavior. After the display output at the cabin end, the disturbance frame is subjected to time-series sampling and image observation. Combined with sparse transformation mapping and joint optimization solution algorithm, the response time of the actual processing action in each stage can be inferred from the physical response, thereby obtaining a real delay feature based on signal behavior rather than system declaration. Through the physical signal entity of the disturbance frame, the originally unobservable and unmeasurable time domain delay in the link is converted into a computable delay feature.
[0087] Furthermore, by fusing and correcting the delay features extracted based on disturbance frame signal processing with the timestamp calculation value, not only can the delay calculation results of each processing stage be dynamically adjusted, but also the hidden bottlenecks and abnormal delay segments in the link can be effectively identified, thereby complementing and enhancing the traditional timestamp solution.
[0088] See also Figure 3, which is a flow chart of a timestamp marking method provided in an embodiment of the present application. Figure 3 The method shown can be applied to the pre-step of a time delay test method of a remote control driving system of the present application. Figure 3 The method shown includes the following A1-A4, and the specific steps are as follows:
[0089] A1: A timing module based on the PTP protocol is configured on the vehicle side and the cabin side respectively, and a timing module based on the NTP protocol is configured on the cloud side;
[0090] Specifically, in order to ensure that the time base of each node in the remote driving link is unified and comparable, it is necessary to establish timing mechanisms for the vehicle side, cloud side and cabin side respectively to support the synchronization of subsequent timestamps.
[0091] In this embodiment, the vehicle-side and cabin-side serve as terminal execution nodes, typically deployed in edge hardware with industrial-grade or quasi-real-time control capabilities. Therefore, the IEEE1588 Precision Time Protocol (PTP) is selected as the timing solution, achieving nanosecond-level synchronization through a master clock device configured in the local area network. The PTP master clock can use a GPS module or a high-precision temperature-compensated oscillator to drive the clock source, and the timing process is completed within the network layer through message exchange processes such as Sync, Follow_Up, Delay_Req, and Delay_Resp. The cabin-side also uses the same PTP client module to ensure that its local time is in the same time domain as the vehicle-side.
[0092] Furthermore, since the cloud is deployed in a virtual server or data center environment, it is limited by network isolation and cross-wide area transmission bandwidth. In order to reduce deployment complexity and ensure cross-regional accessibility, this embodiment uses a network timing service based on the standard NTP (Network Time Protocol) protocol. The cloud performs periodic time synchronization with a trusted NTP time source and collects offset samples based on its own operating status to establish a long-term time reference. Through the above configuration, the vehicle end and the cabin end in the remote control link build a high-precision timing domain based on PTP, and the cloud establishes a relatively low-precision but predictable time reference through NTP, providing a basis for subsequent multi-terminal timestamp fusion and deviation correction.
[0093] A2: After capturing the image frame, timestamps T1-T4 are marked in sequence on the vehicle side;
[0094] Specifically, in this embodiment, the image frame timestamp process is embedded in key nodes of the vehicle-side image processing chain, recording the actual processing moments of the image frames at different stages. First, the image sensor module captures image frames at a fixed frame rate and transmits them to the encoder chip via the image bus. When the image frame is captured and entered into the image buffer, a timestamp T1 is injected through the SDK interface provided by the encoder. T1 represents the time when the image content was initially captured.
[0095] Furthermore, after the image frame enters the encoding module, the vehicle-side controller will obtain the image frame through an interrupt callback or DMA completion event monitoring. After receiving the cached copy of the complete image data, the controller immediately records the timestamp T2, which is used to reflect the starting moment of the controller processing the image content. After the image processing is completed, including image format conversion, preprocessing and encoding actions, the controller records the timestamp T3, which indicates the processing end point when the image data is available for streaming. Finally, when the TCP streaming module encapsulates the image frame as a network message and submits it to the kernel network buffer, the timestamp T4 is recorded. This timestamp is used to reflect the last node behavior before the vehicle-side completes the streaming. The four timestamps read the system time through the PTP timing module, and all use a high-resolution clock source (accuracy can be down to microseconds) to write into the image frame metadata as the original data for subsequent link delay evaluation.
[0096] A3: Forwards the TCP packet to the cloud. The cloud dynamically corrects the received timestamp based on the clock deviation prediction model and generates the corresponding cloud-corrected timestamp T5.
[0097] Specifically, after the vehicle completes the TCP message encapsulation of the image frame, it pushes it to the cloud server node through the remote communication link. Since the cloud uses the NTP protocol for time synchronization, there is a systematic deviation between the time base and the vehicle-side PTP time domain. If no correction is made, the difference between the cloud reception time T5 and T4 will be incomparable. In this embodiment, to solve this problem, the cloud introduces a deviation prediction model based on a historical clock offset sequence. The model periodically collects the time offset value between the vehicle and the cloud, which can be implemented through synchronization messages or standard protocol interactions, and forms an offset sequence cache. The deviation trend is modeled through the LSTM network, and the estimated offset between the current NTP time and the PTP time is predicted in real time. When the TCP message arrives at the cloud, the system obtains the receiving time T5 and calculates the corresponding corrected timestamp based on the estimated offset to compensate for the systematic differences between different time synchronization domains. The final generated timestamp T5 is written into the image frame metadata, maintaining the same time domain as T1 to T4.
[0098] A4: The cabin receives the message data from the cloud and marks it with T6 and T7 in sequence;
[0099] Specifically, in this embodiment, the cabin receives the image frame data stream forwarded from the cloud via a standard RTSP or TCP client interface. When the system receives a complete image frame and writes it to the local receive buffer, it uses the local PTP clock to mark the reception timestamp T6. T6 represents the time when the cabin system first acquired the complete frame content. Subsequently, the cabin passes the image frame to the rendering module for image display operations. When the image frame is decoded, synthesized, and sent to the display device's frame buffer, the system rendering engine monitors the display trigger event and records the display timestamp T7, which represents the final output moment of the visual presentation of the image content.
[0100] Furthermore, to improve timestamp recording accuracy, the cabin uses a hardware rendering pipeline combined with V-Sync synchronization technology to ensure that the time point marked by T7 accurately reflects the actual frame display trigger period. Furthermore, to support the acquisition of perturbation frame responses at the display layer, the image frame buffer mechanism in this embodiment fully preserves the perturbation frame content and, after T7, activates an external camera or photoelectric detection device to collect the display response for use in constructing the subsequent image response observation matrix.
[0101] Taking the insertion of disturbance frame as an example, please refer to Figure 4 To understand, Figure 4 The disturbance frame principle diagram provided in the embodiment of the present application first encodes the disturbance frame into multiple non-overlapping delay coding regions based on the feature deconstruction requirements of the transmission and reception phase and the processing phase of delay detection;
[0102] Then, a structural disturbance feature is added to the time delay coding region according to the phase characteristics, wherein the structural disturbance feature is generated by frequency distribution coding, phase offset coding and block matching coding;
[0103] Finally, the disturbance frame is periodically inserted into the image frame and transmitted alternately with the image frame at a preset ratio.
[0104] Specifically, the purpose of inserting the disturbance frame is to solve the problem that traditional timestamps cannot accurately capture the unobservable delay caused by asynchronous queuing and non-explicit scheduling in remote control driving systems. Since the data may have uncontrollable transmission interruptions, cache waiting and other behaviors in the processing link, although the system timestamp has synchronization, it does not have physical response verifiability, resulting in the lack of authenticity and resolvability of the overall delay evaluation. To this end, it is necessary to construct a disturbance structure that can be embedded in the image link and has response consistency and signal restoration characteristics as external reference information, so that the path delay experienced by the disturbance frame can be deduced based on the physical response of the image itself at the cabin end, thereby supplementing the dynamic behavior within the link that the timestamp mechanism cannot observe.
[0105] In this embodiment, the perturbation frame is divided into three non-overlapping delay coding regions, each corresponding to a typical processing stage in the remote control driving link, such as image acquisition, image coding, data packaging and transmission. The perturbation frame is coded using a significant structural interference signal and introduces different types of perturbation features in different regions, which are specifically divided as follows:
[0106] The first region corresponds to the image acquisition and preprocessing stage (from vehicle-side image sensor output to encoder input) and employs a frequency-distributed encoding structure. This structure consists of high- and low-frequency modulated stripes. Its spectral characteristics are fixed immediately after sensor output and remain virtually unchanged throughout the entire image processing chain, providing a stable representation of the signal morphology at the time of image acquisition. Therefore, the response displacement observed in this region at the cabin end accurately calibrates the time span from T1 to T2.
[0107] The second region corresponds to the image encoding and network packetization stage (the vehicle-side controller encodes and packages the image and pushes it to the cloud), using a phase-shift encoding structure. This structure is composed of strip combinations or interference patterns with preset phase slips, which can retain their phase relationship after compression and network transmission. By calculating the phase offset recovery in the image, the cabin side can extract the time domain changes that may occur between T2 and T5 due to buffer accumulation, transmission asynchrony, or reordering and retransmission, thereby reflecting the actual propagation time difference between the image leaving the transmitter and arriving at the cloud.
[0108] The third region, corresponding to the cloud-side forwarding and on-board reception and display stages, utilizes a block-matching encoding structure. This structure embeds a set of high-texture-contrast structural tiles within the perturbed frame. Its characteristic is that after network transmission and on-board rendering, slight image position shifts or misaligned display temporal responses may occur due to data reorganization, decoding order, and output queuing. The on-board side extracts the response time characteristics of this region through inter-frame differencing and block matching, capturing frame time drift caused by display queuing and processing delays from T5 to T7. A preset ratio is used to ensure link detection density and system load balance.
[0109] Furthermore, the above-mentioned disturbance frame design is not a static structure, but can be combined with the link model for dynamic parameter control. The layout position, frequency selection, phase control parameters, etc. of the disturbance area can be set in real time through the controller parameter table to match the link response characteristics under different test conditions, thereby improving the signal restoration accuracy. In the actual cabin-end response observation, the response time series of the disturbance frame on the screen can be accurately recorded using a camera or a synchronous acquisition module. By constructing an image response observation matrix and performing sparse transformation and joint inversion solution, the actual response time corresponding to each structural disturbance can be identified, and the physical delay of each stage in the link can be inferred. It should be noted that the specific number of areas can be set according to actual needs.
[0110] See also Figure 5 , which is a flow chart of the signal processing and analysis method provided in an embodiment of the present application. Figure 5 The method shown can be applied to step S3 of the aforementioned method, and the specific steps are as follows:
[0111] S3.1: Acquire a response frame sequence of the disturbance frame in a cabin-side display device, and construct an image response observation matrix based on the response frame sequence;
[0112] Specifically, to achieve detailed perception of the physical delays at each stage of the remote control link, this embodiment constructs an image response observation matrix for delay inversion by capturing a sequence of images of the perturbation frame's response on the cabin-side display device. Because the perturbation frame traverses multiple processing modules in the transmission link, it carries information about the processing and transmission delays of the entire link when it is displayed on the cabin-side. Therefore, the response characteristics of each stage within the link can be restored by analyzing the image changes during the actual display process.
[0113] In this embodiment, the cabin-side display device outputs image content at a fixed refresh rate (e.g., 60 Hz). To ensure acquisition accuracy, an image acquisition device synchronized with the display device is provided, such as an industrial camera with a frame rate of no less than 240 fps. This device is fixedly mounted directly in front of the display screen and synchronized with the display controller via hardware triggering to ensure that each frame of display content is fully captured. The acquisition device records the actual image output by the display screen in the form of a frame sequence, which includes normal image frames and periodically inserted disturbance frames. To facilitate identification, the disturbance frames are pre-set with a unique coding structure when inserted, and they have stable and detectable spatial features in the cabin-side display image. Whenever the disturbance frame is decoded and sent to the display buffer of the display device, the camera will continuously capture images within a time domain window before and after it, recording the entire process from the first display of the disturbance frame to the complete disappearance of the structural pattern. The acquired data is a continuous high-frame-rate image sequence in the format of a three-channel RGB image. The resolution of each frame is consistent with that of the display device and is organized in chronological order. Each round of perturbation frame acquisition will generate a set of response frame sequences, which have dynamic changes of perturbation response in the time dimension and the spatial structure remains stable.
[0114] In one example, when constructing the observation matrix, the position coordinates of each perturbed coding region (e.g., frequency perturbation region A, phase perturbation region B, and block perturbation region C) in each response frame sequence are extracted based on the original perturbation frame structure template. The pixel content of these regions in consecutive frames is then intercepted and stacked. For each perturbation region, a three-dimensional data block consisting of the content of several frames is formed, with dimensions of W×H×T, where W and H represent the spatial dimensions of the region and T represents the number of acquired frames.
[0115] Furthermore, assuming that the spatial coordinate range of the first perturbed region is (x1, y1)-(x2, y2), the pixel content of this region is extracted from each frame of the image and stacked along the time dimension to form a regional response tensor of size (x2–x1)×(y2–y1)×T. Repeating the above operation for all perturbed regions can construct a complete image response measurement matrix G, where each channel corresponds to multiple frames of response data for a perturbed region.
[0116] Preferably, to improve signal extraction stability, this embodiment also performs preprocessing operations on the measurement matrix data, such as grayscale normalization, inter-frame difference enhancement, and regional smoothing filtering, to eliminate background noise and system acquisition errors and enhance the response edge of the perturbation structure in the time domain. The constructed image response measurement matrix G will serve as the core input of the delay inversion algorithm. It has the ability to physically map the perturbation frame response and can support the stage-by-stage and structure-by-structure restoration modeling of the link delay.
[0117] S3.2: Projecting the image response measurement matrix into a sparse feature space via a sparse transform domain to obtain a sparse representation of the disturbance signal in a multi-scale feature domain;
[0118] Specifically, since each delayed coding area of the disturbance frame only produces a significant response at a very few time points in the time domain, the image response observation matrix is essentially a high-dimensional, sparsely distributed, noise-superimposed data structure. Directly using the observation matrix for delay solution not only results in a huge amount of data and too high a dimension, but also the disturbance signal is sparsely responded in the observation matrix. Therefore, it is necessary to use a sparse feature transformation method to reduce its dimensionality and reconstruct its information.
[0119] In this embodiment, the image response observation matrix is obtained by continuously capturing the perturbation frames output by the display device at a high frame rate. It has a structure with fixed spatial dimensions and time series length. To improve data processing efficiency and feature separability, the system divides the observation matrix according to pre-set structural regions within the perturbation frames. The image response data for each perturbation encoding region is extracted separately, forming a set of three-dimensional image response unit data blocks. Each data block represents the continuous response of a specific perturbation structure during cabin-side display, including the visual changes of that structure across multiple frames.
[0120] Furthermore, this embodiment introduces sparse transformation technology to map the above-mentioned response data blocks into a set of specially designed sparse basis structures. Specifically, for each disturbance structure region, the system selects a sparse transformation method that matches the disturbance type: for the frequency disturbance region, an orthogonal transformation method based on frequency decomposition is used to extract the main components of the periodic response; for the phase disturbance region, a multi-level wavelet transform method is used to capture the variation trend caused by local phase offset; for the block matching disturbance region, a sparse dictionary matching method based on texture direction extraction is used to extract the image texture drift characteristics.
[0121] Furthermore, when performing a sparse transformation, the system first reorganizes the multi-frame response images of each perturbed structural region along the time axis and projects them into a sparse transformation domain designed specifically for that region. After projection, the original response image is represented as a set of sparse structural feature coefficients. Based on a preset threshold, the system retains the set of sparse principal components with the highest energy and discards sparse response terms with smaller amplitudes and less significant contributions, ultimately forming a response representation of each perturbed structure in a sparse space.
[0122] Furthermore, in order to achieve the independence of the delay characteristics of different stages of the link, this embodiment also introduces a dimensionality suppression strategy in the sparse representation stage. This strategy sets an activation window for the sparse response data of different disturbance structures according to the stage division when the disturbance frame structure is designed, and only allows sparse responses within a specific time period to participate in subsequent processing. For example, the system sets that the frequency disturbance structure should only respond in the initial stage of the disturbance frame display, the phase disturbance should respond in the middle stage, and the block matching structure should respond in the final stage of the display. The response components that exceed this time range will be automatically reset to zero. Through this suppression mechanism, the sparse expression of the disturbance frame completes the matching with the link stage in the time dimension, effectively reduces cross-stage interference, and improves the accuracy and physical interpretability of the subsequent inversion solution.
[0123] In an example, the specific steps of sparsely representing the disturbance signal are as follows:
[0124] S3.2.1: Transform the image response measurement matrix using an orthogonal sparse transform operator to construct a sparse representation tensor of the response in the mapping domain, wherein the orthogonal sparse transform operator is a preset piecewise modulation basis whose construction maintains structural consistency with the coding structure corresponding to each stage in the perturbation frame;
[0125] S3.2.2: Controlling the principal components of the response sparse representation tensor to be concentrated in the sparse subspace corresponding to each stage through a dimensionality suppression function to generate a sparse feature space;
[0126] S3.3: Construct a delay-inversion joint optimization model using sparse representation as observation variables and propagation delays at each stage as sparse coefficients, wherein the delay-inversion joint optimization model includes a sparse regularization term and a structural consistency constraint term;
[0127] Specifically, although the sparse transformation can compress the image response observation matrix into the activation features of multiple disturbance regions in time, this sparse expression only indicates where the response occurred, but cannot express which processing stage the response belongs to, and the specific corresponding delay. The fundamental purpose of this application is to derive the time it takes for the image frame to actually propagate in each stage of the entire link from the image response. Therefore, in this embodiment, it is necessary to construct a mathematically solvable inversion model to establish a clear causal mapping between the image response features and the processing link stages.
[0128] In this embodiment, during the construction process, the model input is first determined, namely the sparse expression features extracted from the image response observation matrix. This feature is represented as the activation response intensity of multiple perturbed regions on the time axis, each region coming from a predefined structural coding block in the perturbed frame, and each structural coding corresponds one-to-one to a certain link processing stage. Then, based on this structural design logic, each stage in the link (such as acquisition, encoding, streaming, transmission, reception, display, etc.) is regarded as an independent parameter to be estimated, that is, the possible propagation delay of each stage.
[0129] Furthermore, in order to establish a solvable relationship, in this embodiment, the above-mentioned link stages are regarded as a set of independent but limited delay variables. By analyzing the sparse response of each disturbance area, a time window range can be preliminarily obtained, representing when the structural disturbance may have changed. However, due to the inevitable interference in the image response (such as illumination, reflection, and display refresh interference), the maximum activation point cannot be directly taken as the result. Instead, the entire time window needs to be treated as an object to be fitted, and a set of processing stage delay variables is found so that the theoretical response pattern generated by it matches the sparse response of the actual observation.
[0130] Furthermore, two constraints are added to the model: First, a sparse regularization constraint is added to control the distribution of the solved propagation delay to be as concentrated as possible on the time axis, allowing only a small number of activation points to exist to prevent delay value drift or overlap. Second, a structural consistency constraint is introduced, which forces each perturbation structure region to respond only to the link phase to which it belongs. This structural correspondence is determined during the perturbation frame design and can be loaded into the model as a hard constraint to prevent activation overlap between phases during the optimization process.
[0131] As an example, to ensure the model is solvable in actual deployment, this embodiment uses an iterative modeling approach: Initially, based on the activation trends of the perturbed region in the sparse representation, a preliminary estimate of the delay range for each stage is generated. These estimates are then gradually shrunk and adjusted based on the sparse activation values, with each iteration updating the stage position until the overall error falls below a set threshold. During this iterative process, the one-to-one correspondence between the perturbation structure and the link stage is maintained, ensuring that each output delay ultimately originates from its designated structural region.
[0132] Furthermore, through this model construction method, the originally mixed image response features can be deconstructed into a set of delay variables corresponding to the link stage, and through step-by-step fitting and constraint stripping, problems such as signal crossing and inter-frame interference can be eliminated.
[0133] S3.4: performing a joint inversion on the sparse representations of each stage according to the delay-inversion joint optimization model to obtain a sparse coefficient distribution, and using the sparse coefficient distribution as a delay feature;
[0134] Specifically, after the model is built, to obtain the delay results for each link stage, the sparse feature tensors need to be jointly solved to output the optimal estimate of the sparse coefficients. In this embodiment, a FISTA-based accelerated sparse solver is used to perform parallel inversion of multi-region sparse coefficient vectors. A multi-threaded decoupling approach is used to construct a gradient descent path for each perturbation structure region, and the sparse delay components of the corresponding stage are updated in real time.
[0135] During the iteration process, the system dynamically evaluates the error convergence rate, automatically adjusts the sparsity threshold to improve the stability and robustness of the solution, and outputs the final sparse coefficient distribution map when the convergence conditions are met. Each non-zero position in this distribution corresponds to a disturbance response activation time point, that is, the physical response moment at a specific stage in the link. By matching the difference between the system timestamps T1 to T7, the offset of the actual link propagation delay can be derived, achieving high-precision correction of timestamp delay.
[0136] Furthermore, the sparse coefficient distribution obtained through joint solution not only provides a quantitative estimate of the multi-stage delay of the link, but also can construct a delay distribution map at the response observation matrix level, which is used to identify performance issues such as bottlenecks, asymmetric loads or dynamic jitter in the link, thereby giving traditional remote control links physical measurability and response reversibility.
[0137] In an example, the specific steps of the joint inversion solution are as follows:
[0138] S3.4.1: Parallel solve the sparse representation corresponding to each delayed coding region in the perturbation frame to obtain a sparse coefficient vector;
[0139] Specifically, each delay-coded region in the perturbation frame essentially represents an independent processing or transmission stage in the link. After sparse transformation, each region acquires sparse activation features over a time dimension. To determine the response time position of the stage corresponding to that region, the so-called propagation delay, parameter estimation of these sparse activation features is necessary. However, since the stages are essentially logically connected in series, their responses are independent, and their delay distributions are non-overlapping, a parallel partitioning strategy should be adopted during the solution process to construct a solution path for each perturbation region separately, thereby improving model computational efficiency and avoiding cross-interference of solution errors.
[0140] In this embodiment, the disturbance frame is divided into multiple time-delay coding regions, and each region generates a sparse activation feature after sparse transformation, which represents the activation degree of the region in the time window. When constructing the inversion model, each region is treated as an independent solution unit, and the required propagation delay value is solved in parallel under a unified optimization framework. Specifically, a sparse fitting strategy is adopted in the solution of the inversion model. Under the premise of the sparse response of the known region, the propagation delay estimate that is most likely to explain the response behavior is inferred. In form, it is equivalent to inferring the trigger position that generates the pattern from the sparse activation pattern, that is, the time point when the disturbance structure experiences a physical response.
[0141] S3.4.2: Converge the sparse coefficient vectors in each region to the optimal solution domain through iterative optimization;
[0142] Specifically, the sparse solution process is limited by non-ideal factors in the disturbance response, such as the blurring of the disturbance signal caused by image display refresh, activation errors caused by external noise, and misjudgments caused by inter-frame interference. These issues can cause the initial solution to deviate from the actual response time. Therefore, an iterative optimization mechanism must be introduced to perform multiple rounds of correction and convergence on the sparse coefficient vector of the initial solution to obtain a solution that approximates the actual physical behavior.
[0143] In this embodiment, a two-stage optimization mechanism based on iterative threshold update and response alignment reconstruction is used. In the initial stage, a predicted response pattern is constructed based on the activation position of the sparse coefficient, and the similarity is calculated with the actual observed sparse vector and error feedback is generated; then the contraction stage is entered, and the parameter space is gradually compressed and approximated by adjusting the non-zero position in the sparse vector, adding and deleting activation elements, and updating the response window position. Each round of iteration includes the response error minimization objective and the structural consistency constraint recalculation mechanism to ensure that the solution process not only maintains the fit to the real signal, but also does not deviate from the semantic boundaries of the perturbation structure.
[0144] S3.4.3: After iterative optimization, perform stage calibration on the non-zero components in the sparse coefficient vector, and use the response time represented by the corresponding position as the sparse coefficient distribution of each stage;
[0145] Specifically, the non-zero components retained in the sparse coefficient vector represent stable and significant response activation in the perturbed region at a specific point in time. Due to the design principles of the perturbation frame, each perturbation structure corresponds to only one processing stage in the link. Therefore, the position of the non-zero component is equivalent to the moment when the physical response occurred in that stage, that is, the actual propagation delay of that processing stage. By dually calibrating it with structural attribution and time position, a complete link delay distribution can be constructed.
[0146] In this embodiment, based on a pre-established one-to-one mapping between perturbation structures and link stages, the sparse coefficient vector of each perturbation region is calibrated as the response source for a fixed processing stage. After iterative optimization is complete, the system extracts the position of the non-zero component in the vector and converts the frame sequence of that position within the response observation window into time units, which serve as the propagation delay for that stage. Because all sparse responses are derived from physical image acquisition at the cabin end, they reflect the time point at which the perturbation visibility actually occurs at the visual output level, and are therefore more physically measurable and stable than the system's internal timestamps.
[0147] Furthermore, this sparse coefficient distribution set can not only be used directly to correct the timestamp delay difference calculated by the system, but can also be used for engineering purposes such as constructing link dynamic response characteristic diagrams, identifying stage bottlenecks, and assisting in scheduling priority optimization.
[0148] As a preferred embodiment, a link load response method based on periodic feedback stress test frame triggering is proposed in this preferred embodiment. The core idea is not to rely on disturbance frames, but to inject feedback stress test frames with trigger signals into the image data stream and use the system end-to-end feedback mechanism to monitor its physical return delay in real time, thereby establishing a stress test response map of whether the link processing is blocked at each stage.
[0149] Specifically, the system controller inserts feedback stress test frames with specific coding characteristics into the image stream at a fixed interval (for example, every 20 frames). These frames maintain no structural perturbations, but instead embed stress test identifiers in the image metadata and message headers. Once the frame arrives and is processed at the cabin, the cabin system automatically generates a feedback confirmation message, which is sent back to the vehicle via the reverse link. Based on the difference between this feedback return latency and the current normal image frame processing rhythm, the controller calculates the physical response time for each hop of the link. This latency includes factors such as the time the frame was queued, protocol buffering delays, and operating system scheduling delays—these are ghost delays that are attempted to be detected but cannot be covered by the timestamp. The triggering and return of stress test frames are based on the behavior of real links and real processing flows, rather than simulated signals. Therefore, the feedback latency is highly realistic. Furthermore, the frequency of stress test frame insertion is adjustable, and its triggering behavior can be embedded into existing network protocol frameworks without modifying image encoding methods or service content, ensuring high system compatibility. More importantly, the feedback stress testing mechanism essentially provides a means for external parties to proactively query link status, avoiding the structural defect of the system relying solely on internal management and speaking to itself.
[0150] In another preferred solution, in order to accurately capture the link-level processing delay in the remote-controlled driving system, a multi-clock trajectory mapping mechanism based on cross-layer frame boundary alignment is proposed. By detecting the image frame boundary at each key node of the link and synchronously recording the local system clock state when the boundary occurs, multi-stage delay inference is finally achieved through trajectory alignment.
[0151] Specifically, at any stage of image frame acquisition, encoding, packaging, decoding, rendering, and display, there will be captureable frame boundary events (such as frame synchronization signals, packet header writing, decoding start, frame rendering submission, etc.). These events are naturally the trigger points for the start or end of processing behavior. Therefore, by accurately recording the occurrence of these boundaries at each stage, the propagation trajectory of the entire link can be reconstructed.
[0152] In this embodiment, the vehicle-side encoder records the first frame boundary marker and timestamps the image frame before it enters motion estimation. The controller marks the second frame boundary event at the start of the packet. The cloud-side unpacking module records the third frame boundary before the decoder starts. The cabin-side system records the final boundary before the image is submitted to the display buffer. By aligning the clocks at these key frame boundaries and integrating them with the timing systems of each node, a multi-segment trajectory of the complete link propagation of the image frame is constructed. To compensate for clock inconsistencies across nodes, the solution introduces a trajectory mapping function that aligns the trajectories of each stage using the mean of the stable propagation time between historical frames. This allows the actual frame's stay and transition within the system to be reconstructed even when clocks are not strictly synchronized. The system ultimately outputs a set of trajectory time segments for each stage, thereby isolating each processing delay. This approach does not introduce any new frames, does not alter image content, and does not rely on external detection signals. Delay modeling is performed solely through clock extraction from the system's existing frame events, making it completely non-intrusive and system-friendly. Furthermore, by aligning the trajectories rather than strictly comparing timestamps, the solution offers enhanced cross-clock domain adaptability, providing a low-coupling compensation mechanism for scenarios where traditional timing mechanisms lack sufficient accuracy.
[0153] As an example, this embodiment provides a time delay test system for a remote control driving system, the system comprising:
[0154] An acquisition module is used to acquire image frames and insert preset disturbance frames into the image frames. The image frames are sequentially transmitted and processed via the vehicle side, the cloud side, and the cabin side.
[0155] A timestamp recording module is used to record the timestamp information of the image frame at each sending and receiving stage and processing stage on the vehicle side, cloud side and cabin side respectively;
[0156] The delay calculation module is used to determine the delay of each receiving and sending stage and processing stage based on the timestamps of each stage;
[0157] A response acquisition module, configured to acquire an image response signal of the disturbance frame in a cabin-side display device;
[0158] a signal analysis module, configured to perform signal processing and analysis on the image response signal and extract delay features reflecting propagation delays at each stage;
[0159] The delay correction module is used to correct the delay calculated based on the timestamp according to the delay characteristics.
[0160] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A time delay test method for a remote control driving system, applied to testing a remote driving system, wherein the remote driving system includes a vehicle side, a cloud side, and a cabin side, and is characterized in that: The delay testing method includes: Obtaining respective timestamps of transmission, reception, and processing of the captured image frames at the vehicle side, the cloud side, and the cabin side, wherein a preset disturbance frame is inserted into the image frames; Determine the latency of each sending, receiving, and processing stage based on each timestamp; Signal processing and analysis are performed on the image response signal of the disturbance frame at the cabin end to extract delay features, and the time delay calculated based on the timestamp is corrected according to the delay features.
2. The time delay testing method of a remote control driving system according to claim 1, characterized in that: Before obtaining each timestamp, the delay testing method further includes: The vehicle side and the cabin side are respectively configured with a timing module based on the PTP protocol, and the cloud side is configured with a timing module based on the NTP protocol; A clock deviation prediction model is constructed based on the historical deviation sequence. The dynamic offset between the cloud NTP time and the other end PTP time is predicted by the clock deviation prediction model, and the cloud time base is corrected according to the dynamic offset.
3. The time delay testing method of a remote control driving system according to claim 2, characterized in that: Before obtaining the timestamps of each image frame being sent, received, and processed on the vehicle side, the latency testing method further includes: In the vehicle-side image acquisition and encoding stage, image frames are acquired by the vehicle-side image sensor and input into the motion estimation stage of the encoder, wherein a disturbance frame with preset characteristics is inserted into the image frame before input into the encoder; Based on the PTP time-serving system time, inject the timestamp T1 into the image frame through the software development interface of the encoder; In the image processing process, the controller on the vehicle side marks the image receiving timestamp T2 after receiving the image frame, and marks the processing completion timestamp T3 after the processing is completed; The processed image frames are encapsulated into TCP packets for streaming, and the streaming timestamp T4 is marked.
4. The time delay testing method of a remote control driving system according to claim 3, characterized in that: Before obtaining the timestamps of each image frame being sent, received, and processed in the cloud, the latency testing method further includes: The cloud receives the TCP message from the vehicle and unpacks it; Dynamically correct the received timestamp according to the clock deviation prediction model to generate the corresponding cloud-corrected timestamp T5; The cloud-side correction timestamp T5 is written into the RTP protocol extension header, and the message data is forwarded to the cabin end via the RTSP protocol.
5. The time delay testing method of a remote control driving system according to claim 4, characterized in that: Before obtaining the timestamps of each image frame being sent, received, and processed at the cabin end, the latency testing method further includes: The cabin receives the message data forwarded by the cloud through the communication equipment, and unpacks and extracts the timestamps T1 to T5 in the message; Marking the current timestamp T6 of the image frame received by the cabin end; The image frame is rendered and outputted through a display device, and an image display timestamp T7 is marked.
6. The time delay testing method of a remote control driving system according to claim 3, characterized in that: The inserting a disturbance frame having preset characteristics into the image frame before inputting into the encoder comprises: Based on the feature deconstruction requirements of the transmission and reception phase and the processing phase of delay detection, the disturbance frame is encoded into multiple non-overlapping delay coding regions; adding a structural disturbance feature in the time delay coding region according to the phase characteristics, wherein the structural disturbance feature is generated by frequency distribution coding, phase offset coding and block matching coding; The disturbance frame is periodically inserted into the image frame and transmitted alternately with the image frame at a preset ratio.
7. The time delay testing method of a remote control driving system according to claim 6, characterized in that: Performing signal processing and analysis on the image response signal of the disturbance frame at the cabin end to extract delay features includes: Acquiring a response frame sequence of the disturbance frame in a cabin-side display device, and constructing an image response observation matrix according to the response frame sequence; Projecting the image response measurement matrix into a sparse feature space through a sparse transform domain to obtain a sparse representation of the disturbance signal in a multi-scale feature domain; Constructing a delay-inversion joint optimization model with sparse representation as observation variable and propagation delay of each stage as sparse coefficient, wherein the delay-inversion joint optimization model includes a sparse regularization term and a structural consistency constraint term; According to the delay inversion joint optimization model, the sparse representation of each stage is jointly inverted and solved to obtain a sparse coefficient distribution, which is used as a delay feature.
8. The time delay testing method of a remote control driving system according to claim 7, characterized in that: The projecting the image response measurement matrix to a sparse feature space through a sparse transform domain includes: The image response measurement matrix is transformed by an orthogonal sparse transformation operator to construct a response sparse expression tensor in the mapping domain, wherein the orthogonal sparse transformation operator is a preset piecewise modulation basis, and its construction maintains structural consistency with the coding structure corresponding to each stage in the perturbation frame; The principal components of the response sparse expression tensor are controlled to be concentrated in the sparse subspace corresponding to each stage by a dimension suppression function to generate a sparse feature space.
9. The time delay testing method of a remote control driving system according to claim 7, characterized in that: The joint inversion solution of the sparse representation of each stage includes: The sparse representation corresponding to each delay coding region in the disturbance frame is solved in parallel to obtain a sparse coefficient vector; Through iterative optimization, the sparse coefficient vectors of each region are converged to the optimal solution domain; After iterative optimization, the non-zero components in the sparse coefficient vector are calibrated in stages, and the response moments represented by the corresponding positions are used as the sparse coefficient distribution of each stage.
10. A time delay test system for a remote control driving system, used to implement a time delay test method for a remote control driving system according to any one of claims 1 to 9, characterized in that: The system comprises: An acquisition module is used to acquire image frames and insert preset disturbance frames into the image frames. The image frames are sequentially transmitted and processed via the vehicle side, the cloud side, and the cabin side. A timestamp recording module is used to record the timestamp information of the image frame at each sending and receiving stage and processing stage on the vehicle side, cloud side and cabin side respectively; The delay calculation module is used to determine the delay of each receiving and sending stage and processing stage based on the timestamps of each stage; A response acquisition module, configured to acquire an image response signal of the disturbance frame in a cabin-side display device; a signal analysis module, configured to perform signal processing and analysis on the image response signal and extract delay features reflecting propagation delays at each stage; The delay correction module is used to correct the delay calculated based on the timestamp according to the delay characteristics.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and readable storage medium
CN113286194A
Data transmission delay test method and system and electronic equipment
CN118368225A
Forward collision early warning method and system based on remote control driving system
CN119126632A
Full-link image transmission delay test method and system of remote control driving system
CN119232920A
Remote control anti-delay video transmission method based on future scene generation network
CN119254998A