Unmanned aerial vehicle visual information frame distortion prediction method and device, electronic equipment and medium
By training the LSTM model to predict the standard deviation error of the residual distribution of the drone visual information frame and perform error compensation, the problem of insufficient prediction accuracy of the visual information distortion of the drone is solved, and more accurate visual information distortion estimation and stronger visual perception ability are achieved.
Patent Information
- Application Number
- CN202411857560.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing drone visual information distortion prediction methods are difficult to ensure prediction accuracy when the drone moves rapidly and the visual information changes significantly, which often leads to large errors and cannot effectively guide the objective evaluation of the drone's visual information quality and the optimization of encoding parameters.
By training the LSTM model, the encoding mode information of the drone visual information frame is obtained, the standard deviation error of the frame residual distribution is predicted, and error compensation is performed, and the compensation standard deviation is used to predict the distortion value.
It realizes accurate prediction and error compensation for the standard deviation error of the frame residual distribution of drone visual information, improves the accurate estimation ability of drone visual information distortion, and enhances the visual perception prediction ability in harsh environments.
Smart Images

Figure CN120017837A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic information, and in particular relates to a method, device, electronic equipment and medium for predicting distortion of visual information frames of unmanned aerial vehicles. Background Art
[0002] UAV visual information distortion is an important indicator for evaluating the quality of UAV visual information, and is also a key basis for objectively evaluating the quality of visual information. By predicting UAV visual information distortion, it can provide guidance for objective evaluation of UAV visual information quality, optimization adjustment of visual information encoding parameters, and intelligent adjustment of output bit rate.
[0003] The existing UAV visual information distortion prediction method is based on the distribution characteristics of frame transform residuals. First, the frame transform residual distribution is described, and then the distortion prediction is performed based on the relationship between the description information and the distortion. In the process of describing the frame transform residual distribution, since the frame transform residual has the characteristics of Laplace distribution, the standard deviation of the Laplace distribution is used to describe it. In addition, considering that the transform residual distribution of each frame is different, after the visual information begins to be collected, the standard deviation of the transform residual distribution of the first 10 frames or 20 frames is counted and its average value is calculated. Based on the average standard deviation of the transform residual distribution, the visual information distortion is predicted.
[0004] However, existing methods are only applicable to scenarios where the visual information content does not change much, such as visual information collection tasks where the surveillance camera is fixed or moves in a small range. When a drone is used as a visual information collection terminal, the distribution of frame transform residuals changes greatly due to the rapid movement of the drone and the large changes in visual information, as well as interference from factors such as airflow. Therefore, the existing method of distortion prediction based on the mean standard deviation of transform residuals is difficult to guarantee prediction accuracy when applied to drones, and usually produces large errors, thus failing to provide effective guidance for the objective evaluation of drone visual information quality, the optimization and adjustment of visual information coding parameters, and the intelligent adjustment of output bit rate. Summary of the invention
[0005] In a first aspect, an embodiment of the present invention provides a method for predicting distortion of a UAV visual information frame, including: obtaining first coding mode information corresponding to a first UAV visual information frame to be distorted; inputting the first coding mode information into a trained LSTM model, and outputting a first frame residual distribution standard deviation error corresponding to the first UAV visual information frame; determining a first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and a frame residual distribution average standard deviation, wherein the frame residual distribution average standard deviation is calculated based on a second frame residual distribution standard deviation of a preset number of second UAV visual information frames; inputting the first frame residual distribution standard deviation into a distortion model, and outputting a distortion value of the first UAV visual information frame.
[0006] In some embodiments, determining the standard deviation of the residual distribution of the first frame after error compensation based on the standard deviation error of the first frame residual distribution and the average standard deviation of the frame residual distribution includes: obtaining the standard deviation of the residual distribution of the first frame after error compensation based on the sum of the standard deviation error of the residual distribution of the first frame and the average standard deviation of the frame residual distribution.
[0007] In some embodiments, before inputting the first coding mode information into the trained LSTM model, it also includes: obtaining the second coding mode information and the second frame residual distribution standard deviation corresponding to each second UAV visual information frame in a preset number of second UAV visual information frames; determining the frame residual distribution average standard deviation according to the second frame residual distribution standard deviation corresponding to the preset number of second UAV visual information frames; determining the second frame residual distribution standard deviation error corresponding to each second UAV visual information frame according to the frame residual distribution average standard deviation and the second frame residual distribution standard deviation corresponding to each second UAV visual information frame; constructing a training set of the LSTM model formed by the second coding mode information corresponding to a preset number of second UAV visual information frames and the second frame residual distribution standard deviation error; training the LSTM model based on the training set to obtain the trained LSTM model.
[0008] In some embodiments, obtaining the second coding mode information and the second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames includes: setting the encoding parameters of the encoder, and setting the total number of encoded frames to the preset number; inputting the drone visual information into the encoder to obtain the second coding mode information and the second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames.
[0009] In some embodiments, the first encoding mode information and the second encoding mode information both refer to the number of multiple encoding modes of the corresponding visual information frame; wherein the encoding mode includes at least one of the following: 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, SKIP.
[0010] In a second aspect, an embodiment of the present invention provides a distortion prediction device for a drone visual information frame, comprising: an information acquisition module, used to obtain first coding mode information corresponding to a first drone visual information frame to be distorted; an error prediction module, used to input the first coding mode information into a trained LSTM model, and output a first frame residual distribution standard deviation error corresponding to the first drone visual information frame; an error compensation module, used to determine the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation, and the frame residual distribution average standard deviation is calculated based on the second frame residual distribution standard deviation of a preset number of second drone visual information frames; a distortion prediction module, used to input the first frame residual distribution standard deviation into a distortion model, and output the distortion value of the first drone visual information frame.
[0011] In some embodiments, the error compensation module is specifically used to obtain the error compensated first frame residual distribution standard deviation according to the sum of the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation.
[0012] In some embodiments, a model training module is also included; the model training module is used to obtain the second coding mode information and the second frame residual distribution standard deviation corresponding to each second UAV visual information frame in a preset number of second UAV visual information frames; determine the frame residual distribution average standard deviation according to the second frame residual distribution standard deviation corresponding to a preset number of second UAV visual information frames; determine the second frame residual distribution standard deviation error corresponding to each second UAV visual information frame according to the frame residual distribution average standard deviation and the second frame residual distribution standard deviation corresponding to each second UAV visual information frame; construct a training set of an LSTM model formed by the second coding mode information and the second frame residual distribution standard deviation error corresponding to a preset number of second UAV visual information frames; train the LSTM model based on the training set to obtain the trained LSTM model.
[0013] In a third aspect, an embodiment of the present invention provides an electronic device, characterized in that it includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement the steps of the drone visual information frame distortion prediction method described in any one of the first aspects when executing the program stored in the memory.
[0014] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method for predicting distortion of a drone visual information frame as described in any one of the first aspects are implemented.
[0015] The beneficial effects brought by the present invention are as follows:
[0016] It can be seen from the above scheme that the UAV visual information frame distortion prediction method, device, electronic device and medium provided by the embodiment of the present invention obtain the first coding mode information corresponding to the first UAV visual information frame to be distorted; input the first coding mode information into the trained LSTM model, and output the first frame residual distribution standard deviation error corresponding to the first UAV visual information frame; determine the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation, and the frame residual distribution average standard deviation is calculated according to the second frame residual distribution standard deviation of a preset number of second UAV visual information frames; input the first frame residual distribution standard deviation into the distortion model, and output the distortion value of the first UAV visual information frame; that is, the trained LSTM model is used to accurately predict the frame residual distribution standard deviation error, and then the frame residual distribution standard deviation is error compensated, and the error-compensated frame residual distribution standard deviation is used to accurately estimate the UAV visual information distortion. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a method for predicting distortion of visual information frames of unmanned aerial vehicles provided by an embodiment of the present invention;
[0018] Figure 2 A schematic diagram of a training process of an LSTM model provided in an embodiment of the present invention;
[0019] Figure 3 A schematic flow chart of another method for predicting distortion of UAV visual information frames provided by an embodiment of the present invention;
[0020] Figure 4 A schematic diagram of the structure of a device for predicting distortion of visual information frames of unmanned aerial vehicles provided by an embodiment of the present invention;
[0021] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0023] In the related art, the average standard deviation of frame transform residuals is used for distortion prediction. However, when this method is applied to UAV flight platforms, it cannot effectively deal with the distortion prediction error caused by fluctuations in the visual information content of the UAV, resulting in poor prediction accuracy.
[0024] In response to the above technical problems, the technical concept of the present invention is to compensate for the standard deviation error of the frame transform residual by training a neural network model for compensation of the standard deviation error of the frame transform residual, thereby solving the problem that the fluctuation error of the standard deviation of the visual information frame transform residual during the flight of the UAV is difficult to predict, and achieving accurate estimation of the distortion of the visual information.
[0025] Figure 1 The present invention provides a flowchart of a method for predicting the distortion of a UAV visual information frame. Figure 1 As shown, the UAV visual information frame distortion prediction method includes:
[0026] Step S101: obtaining first coding mode information corresponding to a first drone visual information frame for which distortion prediction is to be performed.
[0027] Specifically, the first drone visual information frame refers to a visual information frame to be subjected to distortion prediction, such as the jth frame in the visual information, and the first coding mode information refers to the coding mode of the first drone visual information frame. The coding mode can significantly affect the change of the distribution of the visual information frame transform residual in the time domain, and can be obtained through the output of the encoder. In this step, the jth frame visual information is input into the encoder to obtain the coding mode information of the jth frame visual information.
[0028] In some embodiments, the first coding mode information refers to the number of multiple coding modes of the corresponding visual information frame; wherein the coding mode includes at least one of the following: 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, SKIP.
[0029] Specifically, the coding mode information refers to the number of 8 coding modes used in each frame, namely, 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, and SKIP. For example, a certain visual information frame includes 2 16x16, 10 8x16, 0 8x8, 3 SKIP, etc. The coding mode information can be represented by a vector: M j =[m 1 ,m 2 ,…,m 8 ] T , where M j represents the coding mode vector of the jth frame, m 1 ,…,m 8 Respectively represent the number of 8 coding modes of the current frame.
[0030] Step S102: input the first coding mode information into the trained LSTM model, and output the first frame residual distribution standard deviation error corresponding to the first UAV visual information frame.
[0031] Specifically, considering the execution time of the UAV visual acquisition task, the amount of generated data, and the fact that the visual information error estimation has the characteristics of sequence data, a long short-term memory neural network model (LSTM model for short) is used to predict the standard deviation error of the frame residual distribution. That is, the corresponding 8-dimensional encoding mode information of the visual information frame to be predicted is input into the trained LSTM model to predict the corresponding frame residual distribution standard deviation error σ j,e .
[0032] Step S103, determining the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation, wherein the frame residual distribution average standard deviation is calculated based on the second frame residual distribution standard deviation of a preset number of second drone visual information frames.
[0033] Specifically, a preset number of visual information frames (i.e., the second visual information frame of the drone) in the visual information of the drone can be obtained, the frame residual distribution standard deviation of each visual information frame can be obtained, and the average value can be calculated to obtain the average standard deviation of the frame residual distribution: use and σ j,e Get the standard deviation of the residual distribution of the first frame after error compensation
[0034] In some embodiments, the step S103 includes: obtaining the error-compensated first frame residual distribution standard deviation according to the sum of the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation.
[0035] Specifically, the corresponding calculation formula is as follows:
[0036]
[0037] in, represents the standard deviation of the residual distribution of the jth frame after error compensation, σ j,e represents the standard deviation error of the residual distribution of the jth frame, Represents the mean standard deviation of the frame residual distribution.
[0038] Step S104: input the standard deviation of the residual distribution of the first frame into the distortion model, and output the distortion value of the first UAV visual information frame.
[0039] Specifically, the calculation formula corresponding to the distortion model is as follows:
[0040]
[0041] Where D represents distortion, Q represents quantization step size, γ = 1 / 6, is the standard deviation of the frame residual distribution, that is, the result calculated by formula (1).
[0042] The method for predicting the distortion of drone visual information frames provided by the embodiment of the present invention realizes the prediction of the standard deviation error of the frame transformation residual distribution through the trained LSTM model, and is used for error compensation of the frame residual distribution standard deviation, so as to realize accurate estimation of the distortion of drone visual information according to the standard deviation of the frame residual distribution after error compensation.
[0043] Based on the above embodiments, Figure 2 A schematic diagram of a training process of an LSTM model provided in an embodiment of the present invention, that is, before executing step S102, the following steps are also included:
[0044] Step S201: Obtain second coding mode information and second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames.
[0045] In some embodiments, the second coding mode information refers to the number of multiple coding modes of the corresponding visual information frame; wherein the coding mode includes at least one of the following: 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, SKIP. Specifically, the coding mode information of the second drone visual information frame used in the training phase is the same as the first drone visual information frame, and is also the number of 8-dimensional coding modes.
[0046] In some embodiments, step S201 includes: setting the encoding parameters of the encoder and setting the total number of encoded frames to the preset number; inputting the drone visual information into the encoder to obtain the second encoding mode information and the second frame residual distribution standard deviation corresponding to each second drone visual information frame of a preset number of second drone visual information frames.
[0047] Specifically, the encoding parameters are set, including the quantization step size, search range, and related frame number, and the total number of encoded frames is set to a preset number I, and the encoding parameters and the total number of encoded frames are sent to the encoder, and the collected visual information is encoded based on the encoder. During the encoding process, the encoding mode of each frame is recorded and stored in the form of a vector: M i =[m 1 ,m 2 ,…,m 8 ] T . Where M i represents the coding mode vector of the i-th frame, m 1 ,…,m 8 Represent the number of 8 coding modes in the current frame, and record the standard deviation of the residual distribution of each frame, denoted as σ i , until the total number of encoded frames I is reached.
[0048] Step S202: determining the frame residual distribution average standard deviation according to the second frame residual distribution standard deviations corresponding to a preset number of second drone visual information frames.
[0049] Specifically, determine the standard deviation of the frame transform residual distribution of the I frame: 1 ~σ I , and 1 ~σ I Add and divide by I to get the average standard deviation of the frame residual distribution, denoted as
[0050] It should be noted that when executing step S103, the average standard deviation of the frame residual distribution calculated in step S202 can be directly used.
[0051] Step S203: determining a second frame residual distribution standard deviation error corresponding to each second drone visual information frame according to the frame residual distribution average standard deviation and the second frame residual distribution standard deviation corresponding to each second drone visual information frame.
[0052] Specifically, the calculation formula for the standard deviation error of the transformation residual distribution of each frame is as follows:
[0053]
[0054] Among them, σ i,eRepresents the standard deviation error of the residual distribution of the i-th frame, σ i represents the standard deviation of the residual distribution of the i-th frame transformation, Represents the mean standard deviation of the frame residual distribution.
[0055] Step S204: construct a training set of an LSTM model formed by second coding mode information corresponding to a preset number of second drone visual information frames and a second frame residual distribution standard deviation error.
[0056] Specifically, construct the dataset matrix Among them, M 1 ~M I Respectively represent the coding mode vectors of the 1st frame to the Ith frame, σ 1,e ~σ I,e They represent the standard deviation of the frame residual distribution of the 1st frame to the Ith frame respectively.
[0057] Step S205: Train the LSTM model based on the training set to obtain the trained LSTM model.
[0058] Specifically, during the training process, the LSTM model parameters can be initialized first, such as the maximum number of iterations N, the training set percentage P, etc.; then the data set D can be divided into a training set P and a verification set T according to the training set percentage, and the LSTM model can be iteratively trained using the training set P until convergence or the maximum number of iterations is reached. The verification set T can be used to verify the trained LSTM model. If the verification fails, training can continue.
[0059] On the basis of the above-mentioned embodiment, by training the LSTM model, a well-trained frame transformation residual distribution standard deviation error compensation neural network model is obtained. The LSTM model takes into account the execution time of the UAV visual acquisition task, the amount of generated data, and the sequence data characteristics of the visual information error estimation, thereby improving the prediction accuracy of the UAV visual information distortion.
[0060] Figure 3 The flowchart of another method for predicting the distortion of UAV visual information frames provided by the embodiment of the present invention can be divided into four parts: training preparation, model training, error prediction, and distortion compensation. Figure 3 The embodiments of the present invention are described in detail:
[0061] 1. Training Preparation
[0062] (1) Setting encoding parameters and total number of encoding frames: Setting encoding parameters such as the encoder's quantization step size, search range, and number of related frames, and setting the total number of encoding frames to a preset number I.
[0063] (2) Encoding the visual information: The visual information is input into the encoder, and the encoder encodes the visual information frame by frame in real time according to the set encoding parameters and encoding frame number.
[0064] (3) Record the coding mode and frame residual distribution standard deviation of each frame: After the visual information frame is encoded, store the 8-dimensional coding mode vector and frame transformation residual distribution standard deviation of each frame.
[0065] (4) Determine whether the total number of encoded frames has been reached: When the total number of encoded frames I has not been reached, return to step (2); if the total number of encoded frames I has been reached, execute step (5).
[0066] (5) Calculate the average standard deviation of the frame residual distribution and the standard deviation error of the residual distribution of each frame: add the standard deviations of the frame residual distribution of I frames and divide by I to obtain the average standard deviation of the frame residual distribution; determine the standard deviation error of the residual distribution of each frame based on the difference between the frame residual distribution standard deviation of each frame and the average standard deviation.
[0067] 2. Model Training
[0068] (6) Initialize LSTM model parameters and construct a data set: Initialize the relevant parameters of the LSTM model, such as the maximum number of iterations, the percentage of the training set, etc.; construct the coding mode vector of each frame and the corresponding residual distribution standard deviation error into a matrix as the model data set.
[0069] (7) Iterative training module: The data set is divided into a training set and a validation set according to the set training set percentage. The training set is used for LSTM model training, and the validation set is used for model validation.
[0070] (8) Determine whether the model has converged or reached the maximum number of iterations: When the model has not converged and has not reached the maximum number of iterations, return to execute step (7); when the model has converged or reached the maximum number of iterations, execute step (9).
[0071] (9) LSTM model training is completed.
[0072] 3. Error prediction
[0073] (10) Recording the coding mode of new visual information: When a new visual information frame arrives, the coding mode information of the new visual information frame is recorded from the output of the encoder.
[0074] (11) Prediction of the standard deviation error of the frame residual distribution of the new visual information frame: Input the encoding mode of the new visual information frame into the trained LSTM model and output the standard deviation error of the frame residual distribution.
[0075] 4. Distortion Compensation
[0076] (12) Calculate the standard deviation of the frame residual distribution after error compensation: superimpose the standard deviation error of the frame residual distribution output by the LSTM model with the average standard deviation of the frame residual distribution, and calculate the standard deviation of the visual information frame residual after error compensation.
[0077] (13) Predicting distortion based on the standard deviation of the visual information frame residual after error compensation: inputting the standard deviation of the visual information frame residual after error compensation into the frame residual standard deviation-distortion model to predict distortion.
[0078] In summary, this embodiment predicts the standard deviation error of frame residual distribution by training the LSTM model, and uses the error to compensate for the error of the standard deviation of frame residual distribution, thereby improving the prediction accuracy of the standard deviation of frame residual distribution, and then improving the prediction accuracy of visual information distortion. It solves the problems of insufficient distortion prediction accuracy of traditional visual information distortion prediction methods during the movement of drones, as well as poor adaptability to picture jitter and scene conversion and susceptibility to interference during the movement of drones. This embodiment enhances the visual perception distortion prediction ability of drones when working in severe weather and complex sea conditions.
[0079] Figure 4 A schematic diagram of the structure of a UAV visual information frame distortion prediction device provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the UAV visual information frame distortion prediction device includes:
[0080] An information acquisition module 401 is used to acquire first coding mode information corresponding to a first drone visual information frame to be subjected to distortion prediction;
[0081] The error prediction module 402 is used to input the first coding mode information into the trained LSTM model, and output the first frame residual distribution standard deviation error corresponding to the first UAV visual information frame;
[0082] An error compensation module 403 is used to determine the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation, wherein the frame residual distribution average standard deviation is calculated based on the second frame residual distribution standard deviations of a preset number of second drone visual information frames;
[0083] The distortion prediction module 404 is used to input the standard deviation of the residual distribution of the first frame into the distortion model and output the distortion value of the first drone visual information frame.
[0084] In some embodiments, the error compensation module 403 is specifically used to:
[0085] The first frame residual distribution standard deviation after error compensation is obtained according to the sum of the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation.
[0086] In some embodiments, a model training module 405 is also included;
[0087] The model training module 405 is used to obtain the second coding mode information and the second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames;
[0088] Determine the frame residual distribution average standard deviation according to the second frame residual distribution standard deviation corresponding to a preset number of second drone visual information frames;
[0089] Determine a second frame residual distribution standard deviation error corresponding to each second drone visual information frame according to the frame residual distribution average standard deviation and the second frame residual distribution standard deviation corresponding to each second drone visual information frame;
[0090] Constructing a training set of an LSTM model formed by second coding mode information corresponding to a preset number of second drone visual information frames and a standard deviation error of a second frame residual distribution;
[0091] The LSTM model is trained based on the training set to obtain the trained LSTM model.
[0092] In some embodiments, the model training module 405 is specifically used to:
[0093] Setting encoding parameters of the encoder, and setting the total number of encoded frames to the preset number;
[0094] The drone visual information is input into the encoder to obtain second coding mode information and second frame residual distribution standard deviation corresponding to each second drone visual information frame of a preset number of second drone visual information frames.
[0095] In some embodiments, the first encoding mode information and the second encoding mode information both refer to the number of multiple encoding modes of the corresponding visual information frame;
[0096] The coding mode includes at least one of the following: 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, and SKIP.
[0097] Technicians in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process and corresponding beneficial effects of the drone visual information frame distortion prediction device described above can refer to the corresponding process in the aforementioned method example, and will not be repeated here.
[0098] like Figure 5 As shown, an embodiment of the present invention provides an electronic device, including a processor 501, a communication interface 502, a memory 503 and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.
[0099] Memory 503, used for storing computer programs;
[0100] In one embodiment of the present invention, the processor 501 is used to implement the steps of the drone visual information frame distortion prediction method provided by any one of the aforementioned method embodiments when executing the program stored in the memory 503.
[0101] The implementation principle and technical effect of the electronic device provided by the embodiment of the present invention are similar to those of the above embodiment and will not be described in detail here.
[0102] The above-mentioned memory 503 can be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. The memory 503 has a storage space for program codes for executing any method steps in the above-mentioned method. For example, the storage space for program codes may include various program codes for implementing various steps in the above method respectively. These program codes can be read from or written into one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. Such computer program products are usually portable or fixed storage units. The storage unit may have a storage segment or storage space arranged similarly to the memory 503 in the above-mentioned electronic device. The program code can be compressed, for example, in an appropriate form. Generally, the storage unit includes a program for executing the method steps according to an embodiment of the present invention, that is, a code that can be read by a processor such as 501, which, when run by an electronic device, causes the electronic device to execute various steps in the method described above.
[0103] The embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned method for predicting the distortion of the visual information frame of a drone are implemented.
[0104] The computer-readable storage medium may be included in the device / apparatus described in the above embodiment; or it may exist independently without being assembled into the device / apparatus. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0105] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.
[0106] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0107] The above are preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for predicting the distortion of UAV visual information frames, characterized in that: include: Obtaining first coding mode information corresponding to a first drone visual information frame for which distortion prediction is to be performed; Inputting the first coding mode information into the trained LSTM model, and outputting the first frame residual distribution standard deviation error corresponding to the first UAV visual information frame; Determine the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation, wherein the frame residual distribution average standard deviation is calculated based on the second frame residual distribution standard deviation of a preset number of second drone visual information frames; The standard deviation of the residual distribution of the first frame is input into the distortion model, and the distortion value of the first UAV visual information frame is output.
2. The method according to claim 1, characterized in that The determining the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation includes: The first frame residual distribution standard deviation after error compensation is obtained according to the sum of the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation.
3. The method according to claim 1 or 2, characterized in that: Before inputting the first encoding mode information into the trained LSTM model, the method further includes: Obtain second coding mode information and a second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames; Determine the frame residual distribution average standard deviation according to the second frame residual distribution standard deviation corresponding to a preset number of second drone visual information frames; Determine a second frame residual distribution standard deviation error corresponding to each second drone visual information frame according to the frame residual distribution average standard deviation and the second frame residual distribution standard deviation corresponding to each second drone visual information frame; Constructing a training set of an LSTM model formed by second coding mode information corresponding to a preset number of second drone visual information frames and a standard deviation error of a second frame residual distribution; The LSTM model is trained based on the training set to obtain the trained LSTM model.
4. The method according to claim 3, characterized in that The step of obtaining the second coding mode information and the second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames includes: Setting encoding parameters of the encoder, and setting the total number of encoded frames to the preset number; The drone visual information is input into the encoder to obtain second coding mode information and second frame residual distribution standard deviation corresponding to each second drone visual information frame of a preset number of second drone visual information frames.
5. The method according to claim 4, characterized in that The first coding mode information and the second coding mode information both refer to the number of multiple coding modes of the corresponding visual information frame; The coding mode includes at least one of the following: 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, and SKIP.
6. A device for predicting the distortion of a drone visual information frame, characterized in that: include: An information acquisition module, used to acquire first coding mode information corresponding to a first drone visual information frame to be subjected to distortion prediction; An error prediction module, used for inputting the first coding mode information into the trained LSTM model, and outputting a first frame residual distribution standard deviation error corresponding to the first UAV visual information frame; An error compensation module, used to determine the first frame residual distribution standard deviation after error compensation according to the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation, wherein the frame residual distribution average standard deviation is calculated based on the second frame residual distribution standard deviation of a preset number of second drone visual information frames; The distortion prediction module is used to input the standard deviation of the residual distribution of the first frame into the distortion model and output the distortion value of the first UAV visual information frame.
7. The device according to claim 6, characterized in that The error compensation module is specifically used for: The first frame residual distribution standard deviation after error compensation is obtained according to the sum of the first frame residual distribution standard deviation error and the frame residual distribution average standard deviation.
8. The device according to claim 6 or 7, characterized in that It also includes a model training module; The model training module is used to obtain the second coding mode information and the second frame residual distribution standard deviation corresponding to each second drone visual information frame in a preset number of second drone visual information frames; Determine the frame residual distribution average standard deviation according to the second frame residual distribution standard deviation corresponding to a preset number of second drone visual information frames; Determine a second frame residual distribution standard deviation error corresponding to each second drone visual information frame according to the frame residual distribution average standard deviation and the second frame residual distribution standard deviation corresponding to each second drone visual information frame; Constructing a training set of an LSTM model formed by second coding mode information corresponding to a preset number of second drone visual information frames and a standard deviation error of a second frame residual distribution; The LSTM model is trained based on the training set to obtain the trained LSTM model.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the steps of the drone visual information frame distortion prediction method described in any one of claims 1 to 5 when executing the program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the drone visual information frame distortion prediction method as described in any one of claims 1 to 5 are implemented.