Remote data transmission method, sending end, receiving end and transmission system
By extracting the semantic information of video data frames in the remote driving system and generating future frames on the receiving end, the problem of video data transmission delay in remote driving is solved, and lower transmission delay and higher data transmission quality are achieved.
Patent Information
- Application Number
- CN202510176280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
In the existing remote driving technology, the real-time transmission delay of video data is high, which cannot meet the needs of remote driving, especially in complex network environments.
By extracting semantic information of video data frames at the sending end and generating future frames using a deep neural network on the receiving end, the data that has not been received is predicted, thereby reducing transmission delay.
It minimizes the amount of data, reduces the demand for transmission bandwidth, alleviates network congestion, reduces data transmission delay, and ensures the continuity and real-timeness of the data transmission process.
Smart Images

Figure CN120111241A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote control technology, and in particular to a remote data transmission method, a sending end, a receiving end and a transmission system. Background Art
[0002] Currently, there are still obvious technical barriers between AI-driven self-driving cars and the vision of achieving fully autonomous driving, and remote-controlled cars are considered to be a unique solution to fill this gap. The salient feature of remote driving is that the cockpit and motion actuators are separated in space, and information about the vehicle's surroundings and the vehicle's status are transmitted to the cockpit via the network. The driver makes decisions in the remote cockpit with the help of the returned information, and transmits the control commands back to the vehicle through the network to achieve remote control. Remote driving is considered to be an important safety redundancy for autonomous driving systems, opening up a new path for the full implementation and popularization of autonomous driving technology; at the same time, it can significantly improve the working environment in park scenarios such as ports and mines, and provide safer and more comfortable working conditions.
[0003] In the field of remote driving, real-time transmission of video data is crucial to ensure driving safety and improve operational efficiency. The low latency and high stability of 5G networks provide a technical basis for real-time video transmission, but in actual applications, the volatility of cellular network signal strength has a negative impact on network connection quality; at the same time, the high-speed movement of vehicles in remote driving scenarios will amplify the impact of delays on driving safety. Existing solutions encode and compress video data based on video coding standards such as H.264 and transmit them over the network through the RTP / RTCP protocol. In actual deployment, the real-time delay is usually more than 120ms, which cannot meet the needs of remote driving. Therefore, how to reduce the end-to-end delay of vehicle video data in a complex network environment has become a problem that technicians in this field need to solve urgently. Summary of the invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a remote data transmission method, a sending end, a receiving end and a transmission system to solve the problems of high transmission delay in the prior art.
[0005] The technical solution adopted by the present invention to solve the above problems is:
[0006] A remote data transmission method, wherein a sending end extracts semantic information of a data frame and transmits the extracted semantic information to a receiving end, the receiving end receives the semantic information and generates a future frame prediction based on the semantic information for data that has not yet been actually received by the receiving end; wherein the data is video data or image data.
[0007] As a preferred technical solution, it includes: the sending end uses a deep neural network to perform intra-frame encoding on the key frames of the data frame, thereby performing semantic extraction on the I frame to obtain the key frame semantic information; then, referring to the key frame semantic information, the predicted frame is inter-encoded to obtain the motion vector semantics and residual semantics of the predicted frame; and then the key frame semantic information, the motion vector semantics and the residual semantics of the predicted frame are transmitted to the receiving end.
[0008] As a preferred technical solution, it includes: the receiving end receives the semantic information transmitted by the sending end, uses a deep neural network to refer to the key frame semantic information to restore the complete predicted frame semantic information, and then generates future frames based on the complete predicted frame semantic information and the key frame semantic information.
[0009] As a preferred technical solution, the receiving end performs the following steps:
[0010] J1, predicted frame semantic recovery: restore the complete semantics of the predicted frame based on motion vector semantics, residual semantics and reference frame semantics;
[0011] J2, variational inference: infer the posterior distribution of the latent variables of the semantic time series through variational inference and sample the latent variables;
[0012] J3, multi-scale semantic feature fusion: extract multi-scale semantic features based on CNN network and perform feature fusion;
[0013] J4, Future frame sequence prediction: Generate future frames based on latent variables and fused multi-scale semantic features to predict data that has not yet been actually received by the receiver.
[0014] As a preferred technical solution, it also includes: the receiving end uses a deep neural network to compensate for data that has not yet been actually received by the receiving end.
[0015] As a preferred technical solution, the method further comprises the following steps:
[0016] J5, data compensation: Calculate the loss value of the real frame and the future frame, and train the deep neural network based on the loss value to compensate for the data that has not been actually received by the receiver. The method is:
[0017] First, the receiver calculates the gradient information of the loss function relative to the channel output and updates the parameters of the deep neural network at the receiver, and transmits the gradient information back to the transmitter.
[0018] Then, the sender uses the returned gradient information and its own forward propagation information to continue calculating the gradients of the parameters of the sender's deep neural network for semantic extraction and updates the parameters of the semantic extraction.
[0019] As a preferred technical solution, the optimization goal of the deep neural network is:
[0020]
[0021] In the formula,
[0022]
[0023] Among them, arg represents the parameter value that makes the subsequent expression reach the minimum value, θ represents the parameter of the sending end neural network, and ψ represents the parameter of the receiving end neural network. represents the optimization of θ and ψ to find the parameter value that minimizes the subsequent expression, α d The hyperparameter representing the difference between the future frame and the real frame, T represents the number of frames predicted in the future, Represents the difference between the future frame and the real frame, represents the future frame set, X T represents the real frame set relative to the future frame time, α KL represents the KL divergence hyperparameter, KL(·||·) represents the KL divergence of two probability distributions, Represents probability distribution and the probability distribution p(z i ), p(·) represents the probability distribution of random variables, z i Representing semantic information The latent variables, | represents the conditional probability, i represents the number of video frames included in an image group, r 1 The dimension representing semantic information, express R 1 Dimensional semantic information, represents the i-th frame in an image group, Indicates z i The variational distribution of || represents the connection symbol of KL divergence, p(z i ) represents z i The prior probability of , λ represents the weight parameter, t represents the number of the current frame, represents the future frame numbered t, x t represents the real frame numbered t, ln(·) represents the natural logarithm, represents the probability distribution of the future frame numbered t.
[0024] A remote data transmitter is used to: use a deep neural network to perform intra-frame encoding on key frames of data frames, thereby performing semantic extraction on I frames to obtain key frame semantic information; then, refer to the key frame semantic information to perform inter-frame encoding on predicted frames to obtain motion vector semantics and residual semantics of the predicted frames; and then transmit the key frame semantic information, motion vector semantics and residual semantics of the predicted frames to a receiving end.
[0025] A remote data receiving end is used to: receive semantic information transmitted by a sending end, use a deep neural network to refer to key frame semantic information to restore complete predicted frame semantic information, and then generate future frames based on the complete predicted frame semantic information and key frame semantic information.
[0026] A remote data transmission system comprises a remote data sending end and a remote data receiving end which are communicatively connected to each other.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The present invention reduces the amount of data to the maximum extent by extracting semantic information of video data, thereby reducing the demand for transmission bandwidth; this measure alleviates the network congestion problem that occurs in a weak network environment, thereby reducing data transmission delay;
[0029] (2) The present invention introduces a prediction model at the receiving end, which predicts and compensates for future frames that have not yet been actually received by the receiving end. The present invention estimates future frames through the prediction model, thereby ensuring that the continuity and real-time performance of the data transmission process are maintained;
[0030] (3) The present invention adopts a video compression method based on semantic information. The sending end uses a deep neural network to extract video semantic information. The content of the semantic information is defined by a prediction model. The semantic information extracted by the sending end makes the future frames predicted by the prediction model at the receiving end as close to the actual situation as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a schematic diagram of the remote driving system;
[0032] Figure 2 for Figure 1 One of the partial enlarged pictures;
[0033] Figure 3 for Figure 1 The second partial enlarged picture;
[0034] Figure 4 This is a schematic diagram of a remote data transmission framework of the present invention;
[0035] Figure 5 for Figure 4 One of the partial enlarged pictures;
[0036] Figure 6 for Figure 4 The second partial enlarged image. DETAILED DESCRIPTION
[0037] The present invention will be further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0038] It is worth noting that: although the following embodiments are only described using video data for vehicle remote driving scenarios as an example, the application scenarios of the present invention are not limited to remote driving, and the data is not limited to video data, but can also be image data, etc. Such application scenarios and data types are also in line with the inventive concept of the present invention, and can also solve the technical problems to be solved through the technical solution of the present invention to achieve unexpected technical effects, and are therefore also within the inventive concept of the present invention.
[0039] Example 1
[0040] like Figures 1 to 6 As shown, in response to the above problems, the present application proposes a low-latency video transmission solution based on semantic communication and video prediction to reduce the end-to-end delay of video data of vehicles in complex network environments.
[0041] The transmitting end and receiving end involved in the present invention are both implemented through neural network modules, specifically including: intra-frame coding module and inter-frame coding module at the transmitting end, predicted frame semantic recovery module, multi-scale feature extraction network, variational inference module, and time series prediction network at the receiving end. Each neural network module is combined into a remote data transmission system according to the idea proposed by the present invention to realize low-delay transmission of video data.
[0042] The specific method is as follows:
[0043] The present invention proposes a task-oriented transmission scheme based on semantic information, which is expected to significantly reduce the end-to-end (sender to receiver, in this embodiment, for video data transmission, the sender is the vehicle end, and the receiver is the remote driving end) delay of real-time transmission of video data when applied to the field of remote driving.
[0044] To achieve this goal, the present invention adopts the following two technical means:
[0045] First, the present invention reduces the amount of data to the maximum extent by extracting semantic information of video data, thereby reducing the demand for transmission bandwidth. This measure is intended to alleviate the network congestion problem that occurs in a weak network environment, thereby reducing data transmission delay.
[0046] Secondly, the present invention introduces a prediction model at the receiving end (this embodiment uses a deep neural network), which will predict and compensate for future frames that have not yet been actually received by the receiving end. In view of the inevitable delay in the data transmission process, the present invention estimates future frames through the prediction model, thereby ensuring that the continuity and real-time performance of the data transmission process are maintained. Regarding compensation: for example, the receiving end successfully predicts the future frames from t+1 to t+s based on the data of frames tn to t. The data of the s-1 frame is actually observed by the sending end but has not yet been received by the receiving end. Displaying the data of the s-1 frame is delay compensation.
[0047] The present invention adopts a video compression method based on semantic information. The sending end uses a deep neural network to extract video semantic information. The content of the semantic information is defined by a prediction model, that is, the semantic information extracted by the sending end needs to make the future frames predicted by the prediction model of the receiving end as close to the actual situation as possible.
[0048] Consider a GOP (Group of Picture) in video data, which contains a key frame I frame and several predicted frames P frames. Video data has a lot of spatial redundancy and temporal redundancy. Removing redundancy as much as possible without losing too much semantic information is the key to video data compression encoding.
[0049] At the sending end, the key frame I frame is used as the reference frame for encoding the predicted frame P frame. The semantic information of the I frame needs to be restored as clearly and completely as possible at the receiving end so that the semantic information of the predicted frame P frame can be reconstructed with high quality. Therefore, only the I frame is intra-coded, that is, a deep neural network is used to perform semantic extraction on the I frame alone under a certain compression rate; for the predicted frame P frame, the sending end performs inter-frame coding with reference to the semantic information of the I frame and the current predicted frame, and uses a deep neural network to extract the motion vector semantics and residual semantics.
[0050] At the receiving end, the semantic information of the I frame is directly used as the input of the subsequent tasks and the reference semantics of the P frame; at the receiving end, the semantics of the motion vector and residual information of the predicted frame are combined with the semantics of the reference I frame, and the deep neural network is used to restore the complete semantic information of the P frame (the residual semantics and motion vector semantics are similar to the result of subtracting the predicted frame from the reference frame, so as to avoid sending the complete predicted frame, and only need to send the difference with the reference frame. At the receiving end, the semantics of the predicted frame before the subtraction is restored based on the reference frame, residual, and motion vector). Then, the complete semantic information of the P frame and the reference frame together constitute the semantic information time series as the input of the subsequent prediction model. The variational distribution parameters of the latent variable are obtained by variational inference of the semantic information time series. This distribution is an approximation of the posterior distribution of the latent variable with respect to the data. The latent variable is randomly sampled in this distribution. The variational distribution learns the distribution characteristics of the data set during the training process. The latent variable contains the characteristics of the video data obtained from the perspective of probability distribution and all reasonable possibilities in the future; at the same time, the semantic information further extracts multi-scale semantic features and fuses them, understands the semantic information of the data from different receptive fields, and comprehensively considers the local and global information of the data. The fused multi-scale semantic features and latent variables are then sent to the time series prediction network as features of a time step for inference and prediction to generate future frames. The multi-scale semantic features here can be understood as feature pyramids in machine vision. The semantic information extracted by the neural network is essentially a feature. In the feature generation process, the model extracts local information from high-resolution features and global information from low-resolution features, and fuses information of different resolutions (scales) to obtain local and global information of a frame of image. Finally, the loss values of the corresponding time step real frame and future frame are calculated, and the deep neural network is trained based on the loss values to compensate for the data that has not yet been actually received by the receiver.
[0051] It should be noted that: because the P frame of the aforementioned video encoding is customarily called a prediction frame, which conflicts with the name of the result produced by the prediction behavior of the receiving end, in order to avoid confusion, the prediction data generated by the deep neural network is named "future frame".
[0052] The training of the end-to-end low-latency transmission model proposed in the present invention involves joint training of the transmitter and the receiver. The transmitter performs intra-frame coding on the key frame, and at the same time, the prediction frame refers to the key frame semantic information for inter-frame coding to obtain the motion vector semantics and residual semantics of the prediction frame. The extracted semantic information (three types of key frame semantics, motion vector semantics, and residual semantics) is encoded and sent to the channel to the receiver. Due to the influence of channel noise, the semantic information received by the receiver is polluted to a certain extent, and the prediction model needs to predict future frames under noise interference. The motion vector semantics and residual semantics of the predicted frame at the receiver refer to the key frame semantics to restore the complete predicted frame semantics and constitute the semantic information time series together with the reference frame semantics. The variational distribution of the latent variables of the semantic information time series is obtained by variational inference. The sampling method cannot use backpropagation to update parameters during training, so the re-parametrization trick is used to sample the latent variables in the posterior probability. At the same time, multi-scale feature extraction is performed on the semantic information time series. By fusing information at different levels, a feature map that integrates information at different scales is obtained. The fused multi-scale semantic features and latent variables together constitute an input vector of a time step and are sent to the time series prediction network for inference and prediction to generate future frames. Finally, the loss values of the corresponding time step real frame and future frame are calculated. More specifically: the receiving network calculates the gradient of the loss function relative to the channel output and updates the parameters. Since the sending and receiving ends are separated by the channel, direct back propagation cannot be performed. The receiving end needs to transmit the gradient information back to the sending end. The sending end uses the returned gradient information and its own forward propagation information to continue to calculate the gradients of the relevant parameters of the sending end's deep neural network for semantic extraction, and realizes joint back propagation and parameter update of the sending and receiving ends, thereby realizing end-to-end model optimization, and the model converges to the global optimal point (that is, the end-to-end optimization goal, under the action of these parameters, the difference between the future frame and the real frame at the corresponding moment is minimized).
[0053] Assumptions is a video data containing C GOPs, where It means that the nth GOP contains i video frames and the first frame in each GOP is is the key frame, and the rest of the frames are prediction frames. Among them, R represents the vector space, m represents the dimension of the video frame, represents the m-dimensional keyframe, represents the i-th video frame of m dimension, W represents the width of the video frame resolution, and H represents the height of the video frame resolution.
[0054] Using intra-frame semantic encoder coding Get key frame semantic information Using inter-frame semantic encoder Encode the predicted frame in GOP to obtain motion vector semantics and residual semantic information Will and The key frame semantic information, motion vector semantic information, and residual semantic information received by the receiver are respectively and (Motion vector semantics and residual semantics are used uniformly to express), then combined Will Restore the complete semantic information of the predicted frame So far, the complete semantic information of GOP has been successfully reconstructed at the receiving end. represents r of the i-th frame 1 dimensional semantic information. Among them, Represents key frame semantic information, Represents motion vector semantics and residual semantics.
[0055] Inferring the network through variational Infer the latent variable z i The variational distribution of As much as possible The prior probability p(z i ), usually p(z i ) is assumed to be a standard normal distribution, and the KL divergence is used to measure the closeness of the two distributions (one of the KL divergences involved in the present invention is used to maintain the structure and continuity of the potential distribution space. The KL divergence is one of the inherent requirements of the variational inference technique, and its purpose is to make This distribution is closer to the standard Gaussian distribution (i.e., the prior distribution of z), maintaining the structure and continuity of the potential distribution space. ), which can show the probability space modeling of the latent variables and ensure the structure and continuity of the latent variables, namely:
[0056]
[0057] from Sampling to obtain latent variable z i . Among them, min represents the minimum value function.
[0058] Using multi-scale semantic extraction network Extract features of semantic information under different receptive fields and fuse multi-scale information to obtain multi-scale features Finally, the latent variable z i and multi-scale features The same as the feature c of a time step i, using the time series prediction network to generate future frames in T represents the number of frames predicted in the future. The value of T is selected based on the model performance and latency. For example, if the overall system latency is about 200ms and the frame rate is 30FPS, T=6 can compensate for the 200ms latency. The goal of the model is to predict future frames that are consistent with the actual situation as much as possible, that is:
[0059]
[0060] The present invention uses mean square error (MSE) and KL divergence (the second KL divergence involved in the present invention is used to measure the difference between the predicted frame and the real frame at the corresponding moment, which is the originality of the present invention) to measure the difference from the pixel level and probability distribution perspective. and X T The difference, where λ is the weight parameter, is:
[0061]
[0062] Formula (3) is the expansion of (2). The first half of the right side of the equal sign in formula (3) is is the mean square error (MSE), the second half of the right side of the equal sign in formula (3) is the KL divergence.
[0063] Therefore, the end-to-end optimization goal of the model is:
[0064]
[0065] Among them, arg represents the parameter value that makes the subsequent expression reach the minimum value, θ represents the parameter of the sending end neural network, and ψ represents the parameter of the receiving end neural network. represents the optimization of θ and ψ to find the parameter value that minimizes the subsequent expression, α d The hyperparameter representing the difference between the future frame and the real frame, T represents the number of frames predicted in the future, Represents the difference between the future frame and the real frame, represents the future frame set, X T represents the real frame set relative to the future frame time, α KL represents the KL divergence hyperparameter, KL(·||·) represents the KL divergence of two probability distributions, Represents probability distribution and the probability distribution p(z i ), p(·) represents the probability distribution of random variables, z i Representing semantic information The latent variables, | represents the conditional probability, i represents the number of video frames included in an image group, r 1The dimension representing semantic information, express R 1 Dimensional semantic information, represents the i-th frame in an image group, Indicates z i The variational distribution of || represents the connection symbol of KL divergence, p(z i ) represents z i The prior probability of , λ represents the weight parameter, t represents the number of the current frame, represents the future frame numbered t ( The element numbered t in t represents the real frame numbered t (X T t in ), ln(·) represents the natural logarithm, represents the probability distribution of the future frame numbered t. d and α KL Used to balance the weights of the two parts of the loss function.
[0066] Example 2
[0067] like Figures 1 to 6 As shown, based on Example 1, this example provides a more detailed implementation method.
[0068] This implementation combines Figure 1 The remote driving framework shown illustrates the present invention, wherein a vehicle with reasoning and image acquisition capabilities is defined as the remote driving vehicle end of the present implementation, and a remote cockpit with reasoning, communication, display, and driving kits such as a steering wheel, brake pedal, etc. The configuration of the vehicle and the remote cockpit will be described below.
[0069] The vehicle includes a control unit, an information collection unit and a communication unit, and each unit is described below. The control unit is responsible for receiving relevant control commands and parameters (such as speed, angle, etc.) sent by the upper control system, and controlling the lateral and longitudinal movement of the vehicle accordingly. The vehicle is connected to each electronic control unit (ECU) through the CAN bus. The ECU can receive digital signals from sensors and signals from the upper system. The control command from the remote driving end is encoded and sent to the CAN bus according to the CAN protocol and is received and analyzed by the ECU and sends control commands to actuators such as motors to control the movement of the vehicle. The information collection unit is responsible for collecting information around the vehicle. Remote driving mainly relies on the camera to collect video data as the basis for judging the driving environment. The camera installation position is required to meet the human eye's forward field of view of about 200-220°, while taking into account the observation of the environment on both sides and the rear during driving behavior. The communication unit is responsible for communication between the vehicle and the outside world. Remote driving mainly communicates with the remote driving end. The vehicle is equipped with 5G CPE (Customer Premise Equipment, CPE, customer premise equipment). The 5G CPE is connected to the industrial computer via Ethernet and has a traffic SIM card inserted (either a public network or a private network SIM card).
[0070] The cockpit needs to be equipped with a high-performance industrial computer for reasoning and as a center to connect the display, steering wheel, communication unit, etc.; it also needs to be equipped with a steering wheel and brake pedal to output control commands. For example, if a Logitech G923 model kit is used, the kit is connected to the industrial computer via USB. After the driver is installed, the steering wheel angle and pedal depth can be converted into digital signals to simulate driving behavior. The cockpit is connected to the public network or private network via a network cable to ensure accessibility to the network and the vehicle.
[0071] This implementation will combine Figure 4 The application process of the present invention in remote driving is described.
[0072] Step 1: Model building. The semantic extraction and compression of the video transmission method based on semantic communication and video prediction proposed in the present invention, as well as the prediction and reconstruction of the receiving end, are all based on deep neural networks. Before being applied to actual projects, model building and training are required in advance. Considering the intra-frame semantic encoder in actual projects, Inter-frame semantic encoder Located on the vehicle side, identifying the network Multi-scale semantic extraction network and the time series prediction network g ψ (c 1 ,…,c i ) is located in the remote cockpit and needs to be modeled separately according to the application location. represents the multi-scale semantic extraction network, c 1 ,…,ci They respectively represent the fused feature sequence input into the time series prediction network (i.e., the features after the fusion of multi-scale semantic information and latent variable z).
[0073] One possible implementation of the semantic encoder at the transmitter of the present invention is to use Swin Transformer as the backbone network. Swin Transformer is a deep learning model designed specifically for computer vision tasks. Through a layered architecture and a sliding window self-attention mechanism, it achieves deep extraction and understanding of image features while maintaining high efficiency. The output of the model is determined by the compression rate C (which can be achieved by existing technologies and the specific process will not be repeated here):
[0074]
[0075] Where N sem and N img They represent the number of bits of the output semantics of the semantic extraction network with Swin Transformer as the backbone network and the number of bits of the original image (the original image has an 8-bit color depth). The compression rate can control the amount of data of the semantic features output by the model. The richer the semantic features, the richer the prediction information at the receiving end; conversely, the greater the compression rate, the sparser the semantic features, and the less prediction information the sending end obtains. Inter-frame semantic coding obtains motion vector semantics through optical flow estimation, and obtains predicted frames through motion compensation. The residual is calculated based on the real frame semantics and the predicted frame, and the residual and motion vector semantics are sent to the channel. The receiving end realizes semantic reconstruction of the predicted frame through the corresponding network architecture.
[0076] The recognition network at the receiving end converts the features into the mean and logarithmic variance of the Gaussian distribution by setting learnable parameters, and uses the reparameterization technique (which can be achieved by existing technology, and the specific process will not be repeated here) to achieve gradient descent sampling, that is:
[0077]
[0078] where z i is a latent variable, u is the expected value of the variational distribution of semantic information time series data, σ is the standard deviation of the variational distribution of semantic information time series data, ε is an auxiliary variable sampled from a standard Gaussian distribution, and ⊙ is an element-level multiplication.
[0079] The multi-scale semantic extraction network refers to the idea of residual connection, uses CNN network to perform deeper semantic extraction of semantic features, and fuses features of different scales together to obtain global and local detailed information. The process of extracting multi-scale features through convolutional networks involves using convolution kernels and pooling layers of different sizes to capture semantic information at multiple scales. These operations can produce feature maps with different resolutions, each of which contains important information at the corresponding scale. Feature maps of different scales are first adjusted to the same spatial dimension by upsampling or downsampling, and then spliced along the channel dimension. The spliced feature map contains rich information from different scales, thus providing a more comprehensive feature description for subsequent processing.
[0080] Time series prediction network ψ (c 1 ,…,c i ) can be implemented using a long short-term memory network (LSTM). LSTM can effectively capture and utilize the long-term dependencies of time series by selectively forgetting and retaining the information of time series. In order to output clearer and more realistic future frames, a CNN network is added at the end to generate future frames.
[0081] Step 2: The application scenario of this implementation is remote driving, and the transmission content mostly involves road traffic scenes. In order to achieve good results, the data set used when training the model should contain a large number of relevant scenes. The model training process proposed in this invention involves joint training of the transmitter and the receiver. The gradient information of the receiver needs to be transmitted back to the transmitter for gradient descent. The receiver calculates the loss of the real frame and the future frame at the corresponding time and updates the parameters, that is:
[0082]
[0083] Among them, ψ tt+1 represents the receiving end parameters after the tt+1th parameter update, ψ tt represents the receiving end parameters after the ttth parameter update, ψ is the receiving end parameter, η is the learning rate, Indicates the gradient of parameter ψ, L loss is the loss function of the model, It means to find the gradient of the corresponding parameter of the loss function.
[0084] To achieve end-to-end gradient update, the sender needs to obtain L loss The gradient information of the transmitter parameter θ, but the receiver cannot obtain the transmitter parameter, and the chain rule is used to obtain the loss function at the receiver about the channel output and Loss and will as well as The sender can calculate L by using the returned information. loss The gradient information about the sender parameter θ is:
[0085]
[0086] in, represents the gradient of the loss function with respect to the parameters of the neural network at the sender, represents the gradient of the key frame semantics, the predicted frame motion vector semantics, and the residual semantics with respect to the parameters of the neural network at the sender, represents multiplication, Represents the loss function with respect to and Gradient
[0087] Thus, the sender realizes the end-to-end optimization process, namely:
[0088]
[0089] Among them, θ tt+1 represents the sender parameters after the tt+1th parameter update, θ tt Indicates the sender parameters after the ttth parameter update, Represents the gradient of the loss function with respect to the parameters of the neural network at the sender.
[0090] Step 3: In an embodiment of the present invention, the fully trained transmitter and receiver models are integrated and deployed in the remote driving vehicle and the remote cockpit, respectively. When the vehicle switches to the remote driving mode, the remote cockpit will take over the full control of the vehicle. In this mode, the on-board camera is responsible for collecting image information around the vehicle in real time and inputting the image information into the transmitter model. The transmitter model performs deep semantic extraction on the collected images and encodes them into bit streams based on the algorithm proposed in the present invention. These bit streams are then transmitted to the receiver model in the remote cockpit via the 5G network. After receiving the semantic features from the transmitter, the receiver model will perform the following steps:
[0091] Prediction frame semantics recovery: restore the complete semantics of the prediction frame based on motion vector semantics, residual semantics and reference frame semantics;
[0092] Variational Inference: Infer the posterior distribution of the latent variables of the semantic time series through variational inference and obtain the latent variables through reparameterization techniques;
[0093] Multi-scale semantic feature fusion: extract and fuse multi-scale features based on CNN network;
[0094] Prediction of future frame sequences: Based on latent variables and multi-scale semantic features, images that are still being transmitted or collected are predicted and displayed, so as to realize low-latency display of the surrounding images of the vehicle in the cockpit and provide timely and effective driving information for the remote driver.
[0095] Among them, the two steps of variational inference and multi-scale semantic feature fusion can be executed simultaneously.
[0096] As described above, the present invention can be preferably implemented.
[0097] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.
[0098] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.
Claims
1. A remote data transmission method, characterized in that: The sending end extracts semantic information of the data frame and transmits the extracted semantic information to the receiving end, and the receiving end receives the semantic information and generates future frame prediction data based on the semantic information that has not yet been actually received by the receiving end; wherein the data is video data or image data.
2. A remote data transmission method according to claim 1, characterized in that: include: The sender uses a deep neural network to perform intra-frame encoding on the key frames of the data frame, thereby performing semantic extraction on the I frame to obtain the key frame semantic information; Then, the predicted frame is inter-coded with reference to the key frame semantic information to obtain the motion vector semantics and residual semantics of the predicted frame; and then the key frame semantic information, the motion vector semantics and residual semantics of the predicted frame are transmitted to the receiving end.
3. A remote data transmission method according to claim 2, characterized in that: include: The receiving end receives the semantic information transmitted by the sending end, uses a deep neural network to refer to the key frame semantic information to restore the complete predicted frame semantic information, and then generates future frames based on the complete predicted frame semantic information and key frame semantic information.
4. A remote data transmission method according to claim 3, characterized in that: The receiving end performs the following steps: J1, predicted frame semantic recovery: restore the complete semantics of the predicted frame based on motion vector semantics, residual semantics and reference frame semantics; J2, variational inference: infer the posterior distribution of the latent variables of the semantic time series through variational inference and sample the latent variables; J3, multi-scale semantic feature fusion: extract multi-scale semantic features based on CNN network and perform feature fusion; J4, Future frame sequence prediction: Generate future frames based on latent variables and fused multi-scale semantic features to predict data that has not yet been actually received by the receiver.
5. A remote data transmission method according to claim 4, characterized in that: Also includes: The receiver uses a deep neural network to compensate for the data that has not yet been actually received by the receiver.
6. A remote data transmission method according to claim 5, characterized in that: The following steps are also included: J5, data compensation: Calculate the loss value of the real frame and the future frame, and train the deep neural network based on the loss value to compensate for the data that has not been actually received by the receiver. The method is: First, the receiver calculates the gradient information of the loss function relative to the channel output and updates the parameters of the deep neural network at the receiver, and transmits the gradient information back to the transmitter. Then, the sender uses the returned gradient information and its own forward propagation information to continue calculating the gradients of the parameters of the sender's deep neural network for semantic extraction and updates the parameters of the semantic extraction.
7. A remote data transmission method according to any one of claims 1 to 6, characterized in that: The optimization goal of a deep neural network is: In the formula, Among them, arg represents the parameter value that makes the subsequent expression reach the minimum value, θ represents the parameter of the sending end neural network, and ψ represents the parameter of the receiving end neural network. represents the optimization of θ and ψ to find the parameter value that minimizes the subsequent expression, α d The hyperparameter representing the difference between the future frame and the real frame, T represents the number of frames predicted in the future, Represents the difference between the future frame and the real frame, represents the future frame set, X T represents the real frame set relative to the future frame time, α KL represents the KL divergence hyperparameter, KL(·||·) represents the KL divergence of two probability distributions, Represents probability distribution and the probability distribution p(z i ), p(·) represents the probability distribution of random variables, z i Representing semantic information The latent variable, | represents the conditional probability, i represents the number of video frames included in an image group, r1 represents the dimension of semantic information, express The r1-dimensional semantic information, represents the i-th frame in an image group, Indicates z i The variational distribution of || represents the connection symbol of KL divergence, p(z i ) represents z i The prior probability of , λ represents the weight parameter, t represents the number of the current frame, represents the future frame numbered t, x t represents the real frame numbered t, ln(·) represents the natural logarithm, represents the probability distribution of the future frame numbered t.
8. A remote data transmitter, characterized in that: Used to: use deep neural network to perform intra-frame encoding on key frames of data frames, so as to perform semantic extraction on I frames to obtain key frame semantic information; Then, the predicted frame is inter-coded with reference to the key frame semantic information to obtain the motion vector semantics and residual semantics of the predicted frame; and then the key frame semantic information, the motion vector semantics and residual semantics of the predicted frame are transmitted to the receiving end.
9. A remote data receiving terminal, characterized in that: Used to: receive semantic information transmitted by the sender, use a deep neural network to refer to the key frame semantic information to restore the complete predicted frame semantic information, and then generate future frames based on the complete predicted frame semantic information and key frame semantic information.
10. A remote data transmission system, characterized in that: It comprises a remote data sending end as described in claim 8 and a remote data receiving end as described in claim 9 which are connected to communicate with each other.