A video display method, device, apparatus and storage medium

CN122783684APending Publication Date: 2026-09-18CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610953438.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]本发明提供一种视频显示方法、装置、设备及存储介质,用以解决现有技术中存在关键视频帧丢失,从而导致风险判定的准确率低的问题

Benefits of technology

本申请实施例提供一种视频显示方法、装置、设备及存储介质,该方法对接收到的目标无人机发送的视频信息进行解码处理,得到多个视频帧;基于TSN特征提取和均值计算,确定多个视频帧的时序特征;以及针对每个视频帧,基于通过目标检测模型得到的视频帧的至少一个目标特征系数集合,确定视频帧的目标语义特征;以及针对每个视频帧,基于视频帧的像素矩阵与视频帧的前一帧的像素矩阵,确定视频帧的运动特征;以及针对每个视频帧,基于通过场景检测模型得到的与视频帧对应的多个预设场景的概率值,确定视频帧的场景特征;基于预设权重系数集合、时序特征、目标语义特征、运动特征和场景特征,确定视频帧的综合性能得分,其中,综合性能得分用于表征视频帧的重要程度;基于综合性能得分对视频的显示帧率进行调节,采用调节后的显示帧率显示视频,从而便于用户对显示视频实时监控,快速开展预警处置,进而提高风险判定的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122783684A_ABST
    Figure CN122783684A_ABST
Patent Text Reader

Abstract

The application discloses a video display method, device and equipment and a storage medium. The method decodes and processes video information sent by a target unmanned aerial vehicle to obtain a plurality of video frames; determines time sequence characteristics of the plurality of video frames based on TSN feature extraction and mean value calculation; and for each video frame, determines target semantic characteristics based on a target detection model, determines motion characteristics based on a pixel matrix of the video frame and a pixel matrix of a previous frame of the video frame, and determines scene characteristics of the video frame based on a scene detection model; determines a comprehensive performance score of the video frame for representing the importance of the video frame based on a preset weight coefficient set, the time sequence characteristics, the target semantic characteristics, the motion characteristics and the scene characteristics; adjusts the display frame rate of the video based on the comprehensive performance score, and displays the video by using the adjusted display frame rate, so that the user can monitor the displayed video in real time, quickly carry out early warning disposal, and thus the accuracy of risk determination is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a video display method, apparatus, device, and storage medium. Background Technology

[0002] With the rapid development of the low-altitude economy and the gradual opening of low-altitude airspace, the widespread use of drones has generated numerous low-altitude security risks. Counter-drone systems, as core equipment for real-time low-altitude security detection, primarily rely on real-time video streams transmitted from drones to assess risks.

[0003] However, after acquiring a real-time video stream, most related technologies extract video frames at a fixed ratio and discard the oldest video frame when the buffer queue is full, displaying and assessing the remaining video frames. This method is prone to losing critical video frames, resulting in low accuracy in risk assessment. Therefore, how to perform refined video processing has become an urgent technical problem to be solved. Summary of the Invention

[0004] This invention provides a video display method, apparatus, device, and storage medium to solve the problem of low accuracy in risk assessment caused by the loss of key video frames in the prior art.

[0005] In a first aspect, embodiments of this application provide a video display method, the method comprising: The received video information from the target drone is decoded to obtain multiple video frames. Based on TSN feature extraction and mean calculation, the temporal features of the plurality of video frames are determined; and for each video frame, the target semantic features of the video frame are determined based on at least one set of target feature coefficients of the video frame obtained through a target detection model; and for each video frame, the motion features of the video frame are determined based on the pixel matrix of the video frame and the pixel matrix of the previous frame of the video frame; and for each video frame, the scene features of the video frame are determined based on the probability values ​​of a plurality of preset scenes corresponding to the video frame obtained through a scene detection model. Based on a preset set of weight coefficients, the temporal features, the target semantic features, the motion features, and the scene features, a comprehensive performance score for the video frame is determined, wherein the comprehensive performance score is used to characterize the importance of the video frame; The video display frame rate is adjusted based on the overall performance score, and the adjusted frame rate is used to display the video.

[0006] In some optional implementations, determining the temporal features of the plurality of video frames based on TSN feature extraction and mean calculation includes: The multiple video frames are divided into multiple video frame groups based on preset values; For each video frame group, the target temporal features of the video frame group are obtained based on TSN feature extraction; The time series features are obtained by averaging multiple target time series features.

[0007] In some optional implementations, determining the target semantic features of the video frame based on at least one set of target feature coefficients obtained through the target detection model includes: The video frames are input into the target detection model to obtain at least one set of target feature coefficients; The coefficients used to characterize the target category in the target feature coefficient set are vectorized, and the coefficients used to characterize the target area in the target feature coefficient set are normalized to obtain the processed target feature coefficient set. Calculate the product of the target feature coefficients in the processed target feature coefficient set; The sum of the products corresponding to each set of target feature coefficients is taken as the target semantic feature.

[0008] In some optional implementations, determining the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame includes: Determine the matrix difference between the pixel matrix of the video frame and the pixel matrix of the previous frame of the video frame, wherein, when the video frame is the first video frame, the pixel matrix of the previous frame of the video frame is a preset matrix. The motion features are obtained by performing norm operations on the matrix differences.

[0009] In some optional implementations, determining the scene features of the video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through a scene detection model includes: The video frames are input into the scene detection model to obtain probability values ​​for multiple preset scenes; The probability values ​​are normalized to obtain the scene features of the video frame.

[0010] In some optional implementations, determining the comprehensive performance score of the video frame based on a preset set of weight coefficients, the temporal features, the target semantic features, the motion features, and the scene features includes: The first weight coefficient corresponding to the temporal feature is determined from the preset weight coefficient set, the second weight coefficient corresponding to the target semantic feature is determined from the preset weight coefficient set, the third weight coefficient corresponding to the motion feature is determined from the preset weight coefficient set, and the fourth weight coefficient corresponding to the scene feature is determined from the preset weight coefficient set. The comprehensive performance score is determined based on the temporal features, the first weight coefficient, the target semantic features, the second weight coefficient, the motion features, the third weight coefficient, the scene features, and the fourth weight coefficient.

[0011] In some optional implementations, adjusting the video display frame rate based on the overall performance score includes: If the overall performance score is less than a first preset threshold, the video display frame rate is reduced. If the overall performance score is greater than the second preset threshold, the video display frame rate is increased; Wherein, the first preset threshold is less than the second preset threshold.

[0012] Secondly, embodiments of this application provide a video display device, the device comprising: The decoding module is used to decode the video information received from the target drone to obtain multiple video frames; The feature extraction module is used to determine the temporal features of the plurality of video frames based on TSN feature extraction and mean calculation; and for each video frame, to determine the target semantic features of the video frame based on at least one set of target feature coefficients of the video frame obtained by the target detection model; and for each video frame, to determine the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; and for each video frame, to determine the scene features of the video frame based on the probability values ​​of a plurality of preset scenes corresponding to the video frame obtained by the scene detection model. The determination module is used to determine the comprehensive performance score of the video frame based on a preset set of weight coefficients, the temporal features, the target semantic features, the motion features, and the scene features, wherein the comprehensive performance score is used to characterize the importance of the video frame; The adjustment module is used to adjust the display frame rate of the video based on the comprehensive performance score, and to display the video using the adjusted display frame rate.

[0013] Thirdly, embodiments of this application provide an electronic device, including: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps included in the method as described in any one of the first aspects, according to the obtained program instructions.

[0014] Fourthly, embodiments of this application provide a computer storage medium storing a computer program for causing a computer to perform the method as described in any one of the first aspects.

[0015] The beneficial effects of this invention are as follows: This application provides a video display method, apparatus, device, and storage medium. The method decodes received video information sent by a target drone to obtain multiple video frames; determines the temporal features of the multiple video frames based on TSN feature extraction and mean calculation; for each video frame, determines the target semantic features of the video frame based on at least one set of target feature coefficients obtained through a target detection model; for each video frame, determines the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; for each video frame, determines the scene features of the video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through a scene detection model; determines the comprehensive performance score of the video frame based on the preset weight coefficient set, temporal features, target semantic features, motion features, and scene features, wherein the comprehensive performance score is used to characterize the importance of the video frame; adjusts the display frame rate of the video based on the comprehensive performance score, and displays the video using the adjusted display frame rate, thereby facilitating real-time monitoring of the displayed video by the user, enabling rapid early warning and response, and improving the accuracy of risk assessment. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1(a) is a schematic diagram of a system architecture provided in an embodiment of this application; Figure 1(b) is a schematic diagram of another system architecture provided in an embodiment of this application; Figure 2 A flowchart illustrating a video display method provided in an embodiment of this application; Figure 3 A flowchart illustrating another video display method provided in an embodiment of this application; Figure 4 A flowchart illustrating another video display method provided in an embodiment of this application; Figure 5 A flowchart illustrating another video display method provided in an embodiment of this application; Figure 6 A flowchart illustrating another video display method provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a video display device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0019] In this application, terms such as "exemplary," "in some embodiments," and "in other embodiments" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the term "exemplary" is used to present the concept in a specific manner.

[0020] It should be noted that the terms "first" and "second" used in the embodiments of this application are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance or order.

[0021] The following explains some of the technical terms used in this application.

[0022] (1) Counter-Unmanned Aerial Vehicle System (C-UAV): A system used to detect, identify, track and deal with target drones.

[0023] (2) TSN (Temporal Segment Network): Divides a complete video into several segments equally, and extracts a small number of key representative frames from each segment to extract features.

[0024] With the rapid development of the low-altitude economy and the gradual opening of low-altitude airspace, the widespread use of drones has generated numerous low-altitude security risks. Drones frequently intrude into no-fly zones or approach key facilities, thus necessitating anti-drone systems to assess the risks associated with drone behavior.

[0025] Anti-drone systems, as core equipment for real-time low-altitude security detection, primarily rely on real-time video streams transmitted by drones for risk assessment. However, after acquiring the real-time video stream, most related technologies extract video frames at a fixed ratio and discard the earliest video frame data when the buffer queue is full, displaying and assessing the remaining video frames. This method is prone to losing critical video frames, resulting in low accuracy in risk assessment.

[0026] The video frame extraction techniques based on deep models such as CNN (Convolutional Neural Network) or Transformer mainly filter video frames by calculating the similarity and semantic differences between frames. However, this method is only applicable to the processing of discrete videos and cannot perform real-time risk assessment of videos, resulting in poor real-time risk assessment for drones.

[0027] Therefore, how to refine video processing to improve the accuracy of risk assessment has become an urgent technical problem to be solved.

[0028] This application provides a video display method, apparatus, device, and storage medium. It decodes received video information sent by a target drone to obtain multiple video frames; determines the temporal features of the multiple video frames; and extracts target semantic features, motion features, and scene features for each video frame. Based on a preset set of weight coefficients, temporal features, target semantic features, motion features, and scene features, it determines a comprehensive performance score for each video frame. Based on the comprehensive performance score, it adjusts the display frame rate of the video and displays the video using the adjusted frame rate, thereby facilitating real-time monitoring of the displayed video, enabling rapid early warning and response, and improving the accuracy of risk assessment.

[0029] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0030] The detection box in Figure 1(a) includes an antenna for acquiring video information sent by the target UAV, a display screen for displaying the video, and an embedded computing motherboard, storage unit, display interaction unit, and power supply unit housed within the detection box. After acquiring the video information via the antenna, the embedded computing motherboard decodes the received video information from the target UAV to obtain multiple video frames. For each video frame, it extracts target semantic features, motion features, and scene features. Then, based on a preset set of weighted coefficients, temporal features, target semantic features, motion features, and scene features, it determines the comprehensive performance score of the video frame. Finally, based on the comprehensive performance score, it adjusts the display frame rate of the video, and displays the adjusted video on the display screen.

[0031] Figure 1(b) illustrates an alternative possible system architecture provided by an embodiment of this application. As shown in Figure 1(b), the system architecture includes a detection box 110 and a display device 120. The detection box 110 includes an antenna for acquiring video information sent by a target drone, and an embedded computing motherboard, a storage unit, a display interaction unit, a power supply unit, etc., disposed in the detection box 110. In a specific implementation, after the detection box 110 acquires video information through the antenna, the embedded computing motherboard decodes the received video information sent by the target drone to obtain multiple video frames. For each video frame, it extracts the target semantic features, motion features, and scene features. Then, based on a preset set of weight coefficients, temporal features, target semantic features, motion features, and scene features, it determines the comprehensive performance score of the video frame. Based on the comprehensive performance score, it adjusts the display frame rate of the video and sends the adjusted display video to the display device 120. The display device 120 displays the adjusted display video.

[0032] In one alternative implementation, the display device 120 may be a portable electronic device, such as a mobile phone, tablet computer, PDA, etc.

[0033] In one alternative implementation, the detection box 110 and the display device 120 can communicate via a communication network.

[0034] In one alternative implementation, the communication network is a wired network or a wireless network.

[0035] It should be noted that Figures 1(a) and 1(b) are merely illustrative examples and are not limited to the above scenarios in the embodiments of this application.

[0036] Figure 2 A flowchart illustrating a video display method according to an embodiment of this application is shown. As shown in the figure, the method includes the following steps: S201. Decode the received video information sent by the target drone to obtain multiple video frames; In a specific embodiment, video information from the target drone is received in real time, and the received video signal is decoded and demodulated to obtain multiple video frames.

[0037] For example, an H.264 universal hardware decoder or an H.265 universal hardware decoder is used to decode and demodulate the received compressed video data, remove the video encapsulation header, eliminate redundant compressed information, and restore the compressed video data to the original video frames, thereby performing feature extraction based on the restored multiple video frames.

[0038] The specific steps for determining the temporal features of multiple video frames based on multiple video frames, and for extracting the target semantic features, motion features, and scene features of each video frame separately, are as follows: S202. Based on TSN feature extraction and mean calculation, determine the temporal features of multiple video frames; and for each video frame, determine the target semantic features of the video frame based on at least one set of target feature coefficients of the video frame obtained through the target detection model; and for each video frame, determine the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; and for each video frame, determine the scene features of the video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through the scene detection model. In one alternative implementation, such as Figure 3 As shown, the steps for determining the temporal features of multiple video frames based on TSN feature extraction and mean calculation are as follows: S301. Divide multiple video frames into multiple video frame groups S based on preset values; In a specific embodiment, a continuous video frame sequence is used. For example, consecutive video frames are divided into K time periods to obtain multiple video frame groups. Each video frame group S includes at least one video frame.

[0039] For example, if the obtained video frame is a 90-frame video frame, that is By setting K=3, the continuous video frames are divided into 3 time periods, resulting in 3 groups of video frames. Among them, video frame group include Video frames, video frame groups include Video frames, video frame groups include The video frames.

[0040] S302. For each video frame group S, obtain the target temporal features of the video frame group based on TSN feature extraction. ; In a specific embodiment, TSN feature extraction is used to extract target behaviors such as continuous approach, loitering, and hovering of a drone in a continuous multi-frame video, and output target temporal features in vector form for subsequent mean calculation.

[0041] For example, to obtain 3 groups of video frames For example, for each group of video frames, the target temporal features are extracted separately, i.e., based on the video frame group. Extract the target temporal features Based on video frame groups Extract the target time series features Based on video frame groups Extract the target temporal features .

[0042] S303, Time-series features of multiple targets Time series features are obtained by calculating the mean. .

[0043] For example, using the aforementioned target time-series features For [0.1, 0.2, 0.5], the target time series features The target time-series features are [0.75, 0.3, 0.1]. Taking [0.8, 0.9, 0.4] as an example, the time series features are obtained by calculating the mean of the three target time series features. .

[0044] Specifically, = =0.55, = =0.47, = =0.33, which is the obtained time series feature. The values ​​are [0.55, 0.47, 0.33].

[0045] This application extracts target temporal features in segments, reducing redundant calculations in consecutive frames and lowering computational consumption. By fusing the average values ​​of segmented features to output temporal features, the accuracy of risk assessment can be improved.

[0046] In one alternative implementation, such as Figure 4 As shown, the steps for determining the target semantic features of a video frame based on at least one set of target feature coefficients obtained through the target detection model are as follows: S401. Input the video frames into the target detection model to obtain at least one set of target feature coefficients. ; Specifically, the object detection model can be a lightweight YOLOv8n model, and the lightweight YOLOv8n model outputs a set of target feature coefficients. This includes the target category in the video frame. Target detection confidence and the bounding box coordinates of the target in the video frame That is, the set of target feature coefficients Where t represents the number of video frames.

[0047] Among them, target category This refers to the categories of targets such as factories, server rooms, military facilities, and key equipment in the video frames transmitted back by the drone. It's understandable that the lightweight YOLOv8n model pre-defines targets such as factories, server rooms, military facilities, and key equipment, along with their categories. The correspondence, for example, the target category corresponding to a factory can be... The target category corresponding to the computer room can be The target category corresponding to military facilities can be: The target category corresponding to key equipment can be: .

[0048] bounding box coordinates The representation of is ( , , , ),in,( , ) represents the top-left pixel coordinates of the bounding box. , () represents the bottom right pixel coordinates of the bounding box. For example, if the target is a computer room, and the top left corner coordinates of the bounding box of the computer room are (12, 8) and the bottom right corner coordinates are (28, 32), then the bounding box of the computer room is (12, 8, 28, 32).

[0049] For example, video frame 1 is input into the object detection model to obtain the set of object feature coefficients. for The video frame 2 is input into the object detection model to obtain the target feature coefficient set. for .

[0050] S402, Target Feature Coefficient Set The coefficients used to characterize the target category are vectorized, and the coefficients used to characterize the target area in the target feature coefficient set are normalized to obtain the processed target feature coefficient set. In practice, based on a pre-defined correspondence, the coefficients representing the target category are... Convert to a vector, where the preset correspondence can be the target category. The corresponding vector coefficient is 0.2, and the target category is... The corresponding vector coefficient is 0.3, and the target category is... The corresponding vector coefficient is 0.8, and the target category is... The corresponding vector coefficient is 0.95.

[0051] For example, for the target category in the target feature coefficient set a After vectorization, the resulting vector coefficient is 0.2, which corresponds to the target category in the target feature coefficient set b. After vectorization, the resulting vector coefficient is 0.8.

[0052] In practice, the coefficients in the target feature coefficient set used to characterize the target area are normalized, that is, the bounding box coordinates are normalized. The target area is normalized. First, the bounding box coordinates are normalized. The area of ​​the bounding box is calculated by using the coordinates in the image. Then, the quotient of the area of ​​the bounding box and the area of ​​the image is calculated to obtain the normalized coefficient.

[0053] For example, the set of target feature coefficients for[ and target feature coefficient set for Let's take an example to illustrate.

[0054] The bounding box coordinates of the target in the target feature coefficient set a for Calculate the area of ​​the bounding box = ( )×( ) = ( )×(1 =84, if the area of ​​the image is 200, then the bounding box coordinates are... The coefficient after normalization is 0.42.

[0055] The bounding box coordinates of the target in the target feature coefficient set b for Calculate the area of ​​the bounding box = ( )×( ) = ( )×(1 If the area of ​​the image is 200, then the bounding box coordinates are 121. The coefficient after normalization is 0.605.

[0056] In summary, the resulting set of processed target feature coefficients The set of target feature coefficients after processing is [0.2, 0.6, 0.42]. The result is [0.8, 0.6, 0.605].

[0057] S403. Calculate the product of the target feature coefficients in the processed target feature coefficient set; S404. The sum of the products corresponding to each set of target feature coefficients is taken as the target semantic feature. .

[0058] For example, for the target feature coefficient set Calculate the set of target feature coefficients after processing. The target semantic features are calculated by multiplying the target feature coefficients in [0.2, 0.6, 0.42]. It is 0.0504; for the target feature coefficient set Calculate the set of target feature coefficients after processing. The target semantic features are calculated by multiplying the target feature coefficients in [0.8, 0.6, 0.605]. It is 0.2904. Where t represents the number of video frames.

[0059] This application uses a lightweight YOLOv8n model for recognition and detection, which improves recognition efficiency and reduces inference computation overhead. By integrating target category embedding, detection confidence, and target image proportion in a multi-dimensional weighted aggregation, the accuracy of target recognition is improved.

[0060] In one alternative implementation, such as Figure 5 As shown, the steps to determine the motion features of a video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame are as follows: S501, Determine the pixel matrix of the video frame. The pixel matrix of the previous frame of the video frame The matrix difference, where, in the case that the video frame is the first video frame, the pixel matrix of the previous frame of the video frame is a preset matrix; For example, using the pixel matrix of the current frame =[30, 75, 85, 110], the pixel matrix of the previous frame of the current frame. Taking [20, 50, 90, 120] as an example, the matrix difference = =[10, 25, -5, -10].

[0061] It is understandable that, when the video frame is the first video frame, the pixel matrix of the previous frame is a preset matrix, for example, the preset matrix can be [10, 10, 10, 10].

[0062] S502. Perform norm operations on the matrix differences to obtain motion characteristics. .

[0063] For example, motion features || ||= .

[0064] This application obtains motion features by solving the inter-frame pixel difference norm. It can effectively distinguish the motion states of drones, such as hovering, slow movement, and rapid scanning, and requires little computation, thus reducing computational overhead while ensuring the accuracy of motion state identification.

[0065] In one alternative implementation, such as Figure 6 As shown, the steps to determine the scene features of a video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through the scene detection model are as follows: S601. Input the video frames into the scene detection model to obtain the probability values ​​of multiple preset scenes; Specifically, the scene detection model can be a lightweight MobileNetV3 model, which is used to classify scenes in drone image transmission footage and output probability values ​​corresponding to each scene.

[0066] For example, if the scene is classified as a general civilian area, a transportation hub area, and a core confidential area, and the probability values ​​output by the lightweight MobileNetV3 model are 12 for general civilian areas, 45 for transportation hub areas, and 83 for core confidential areas, then the probability values ​​of the multiple preset scenes output are [12, 45, 83].

[0067] S602. Normalize the probability values ​​to obtain the scene features of the video frames. .

[0068] The scene features are obtained by normalizing the probability values ​​of multiple preset scenes [12, 45, 83]. The values ​​are [0.09, 0.33, 0.58].

[0069] This application identifies video frame scenes based on a scene detection model, avoiding the problem of limited risk assessment caused by relying solely on screen targets or image motion, and improving the accuracy of risk determination.

[0070] S203, Based on a preset set of weight coefficients and time-series features Target semantic features Motion characteristics and scene features The overall performance score of the video frame is determined, where the overall performance score is used to characterize the importance of the video frame; In one alternative implementation, the time-series features are determined from a preset set of weighting coefficients. The corresponding first weight coefficient And determine the target semantic features from the preset set of weight coefficients. The corresponding second weighting coefficient And determine motion characteristics from a preset set of weight coefficients. The corresponding third weight coefficient And determine the scene features from the preset set of weight coefficients. The corresponding fourth weight coefficient ; Among them, the weighting coefficients and time-series features Target semantic features Motion characteristics and scene features The correspondence is pre-set.

[0071] The comprehensive performance score T is determined based on temporal features, first weight coefficient, target semantic features, second weight coefficient, motion features, third weight coefficient, scene features, and fourth weight coefficient.

[0072] In a specific embodiment, the overall performance score T = × + × + × + × .

[0073] It should be noted that the first weighting coefficient... Second weighting coefficient Third weighting coefficient and the fourth weighting coefficient The sum is 1. Understandably, during initialization or scene switching, the weight coefficients can be adjusted based on a preset scene weight configuration table; for example, in a scene involving a core sensitive area, the fourth weight coefficient can be increased. With the first weighting coefficient The percentage.

[0074] In this embodiment of the application, multiple feature order features are combined. Target semantic features Motion characteristics and scene features The system evaluates video frames to enhance its comprehensive perception of the drone's flight status and scene risks, while employing a lightweight network structure to effectively reduce computational resource consumption.

[0075] S204. Adjust the display frame rate of the video based on the comprehensive performance score T, and display the video using the adjusted display frame rate.

[0076] In one optional implementation, when the overall performance score T is less than a first preset threshold... In cases where the overall performance score exceeds the second preset threshold, the video display frame rate is reduced; In this case, increase the display frame rate of the video; wherein, the first preset threshold Less than the second preset threshold .

[0077] In a specific embodiment, when the overall performance score T is less than a first preset threshold... In cases where the video's display frame rate is reduced, it indicates that the current video frame has a lower risk; when the overall performance score is greater than the second preset threshold... In such cases, increasing the video display frame rate indicates that the current video frame has a higher risk, making it easier for users to monitor the displayed video in real time and quickly carry out early warning and response.

[0078] In another alternative implementation, a medium-risk threshold can be set. With high risk threshold The overall performance score T is normalized to the range of 0 to 1.

[0079] In the range of 0 ≤ overall performance score T < In cases where the drone is deemed to be in a low-risk state, the video display frame rate is reduced, duplicate video frames are filtered out and discarded, and the storage frequency of video frames is reduced, thereby saving energy and reducing the burden on hardware resources.

[0080] exist ≤Comprehensive performance score T< In such cases, if the drone is determined to be in a medium-risk state, the video display frame rate is increased, and key video frames showing changes in the drone's flight status are stored, facilitating real-time monitoring of the displayed video and enabling rapid early warning and response by the user.

[0081] exist ≤Comprehensive performance score T< In cases where the drone is deemed to be in a high-risk state, all video frames are backed up and stored to ensure the integrity and reliability of monitoring in high-risk scenarios.

[0082] This application provides a video display method that decodes received video information sent by a target drone to obtain multiple video frames; determines the temporal features of the multiple video frames based on the multiple video frames, and extracts target semantic features, motion features, and scene features for each video frame; determines the comprehensive performance score of the video frames based on a preset set of weight coefficients, temporal features, target semantic features, motion features, and scene features; adjusts the display frame rate of the video based on the comprehensive performance score, and displays the video using the adjusted display frame rate, thereby facilitating real-time monitoring of the displayed video by the user, enabling rapid early warning and response, and improving the accuracy of risk assessment.

[0083] Based on the same inventive concept, this application also provides a video display device, which is similar to the video display method described above, and the repetitions will not be repeated. Figure 7 The diagram shown is a structural schematic of a video display device provided in an embodiment of this application. The device includes: The decoding module 701 is used to decode the video information sent by the target drone and obtain multiple video frames. The feature extraction module 702 is used to determine the temporal features of multiple video frames based on TSN feature extraction and mean calculation; and for each video frame, to determine the target semantic features of the video frame based on at least one set of target feature coefficients of the video frame obtained by the target detection model; and for each video frame, to determine the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; and for each video frame, to determine the scene features of the video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained by the scene detection model. The determination module 703 is used to determine the comprehensive performance score of a video frame based on a preset set of weight coefficients, temporal features, target semantic features, motion features, and scene features. The comprehensive performance score is used to characterize the importance of the video frame. The adjustment module 704 is used to adjust the display frame rate of the video based on the comprehensive performance score, and to display the video using the adjusted display frame rate.

[0084] This application provides a video display method and apparatus, which decodes received video information sent by a target drone to obtain multiple video frames; determines the temporal features of the multiple video frames based on TSN feature extraction and mean calculation; determines the target semantic features of each video frame based on at least one set of target feature coefficients obtained through a target detection model; determines the motion features of each video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; determines the scene features of each video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through a scene detection model; determines the comprehensive performance score of the video frame based on the preset weight coefficient set, temporal features, target semantic features, motion features, and scene features, wherein the comprehensive performance score is used to characterize the importance of the video frame; adjusts the display frame rate of the video based on the comprehensive performance score, and displays the video using the adjusted display frame rate, thereby facilitating real-time monitoring of the displayed video by the user, enabling rapid early warning and response, and improving the accuracy of risk assessment.

[0085] In an optional embodiment, the feature extraction module 702 is specifically used for: Based on preset values, multiple video frames are divided into multiple video frame groups; For each video frame group, the target temporal features of the video frame group are obtained based on TSN feature extraction; The time series features are obtained by averaging multiple target time series features.

[0086] In an optional embodiment, the feature extraction module 702 is specifically used for: The video frames are input into the target detection model to obtain at least one set of target feature coefficients; The coefficients used to characterize the target category in the target feature coefficient set are vectorized, and the coefficients used to characterize the target area in the target feature coefficient set are normalized to obtain the processed target feature coefficient set. Calculate the product of the target feature coefficients in the processed target feature coefficient set; The sum of the products corresponding to each set of target feature coefficients is taken as the target semantic feature.

[0087] In an optional embodiment, the feature extraction module 702 is specifically used for: Determine the difference between the pixel matrix of the video frame and the pixel matrix of the previous frame, where, if the video frame is the first video frame, the pixel matrix of the previous frame is a preset matrix. Motion characteristics are obtained by performing norm operations on the matrix differences.

[0088] In an optional embodiment, the feature extraction module 702 is specifically used for: The video frames are input into the scene detection model to obtain the probability values ​​of multiple preset scenes; The probability values ​​are normalized to obtain the scene features of the video frames.

[0089] In an optional embodiment, the determining module 703 is specifically used for: The first weight coefficient corresponding to the temporal feature is determined from the preset weight coefficient set, the second weight coefficient corresponding to the target semantic feature is determined from the preset weight coefficient set, the third weight coefficient corresponding to the motion feature is determined from the preset weight coefficient set, and the fourth weight coefficient corresponding to the scene feature is determined from the preset weight coefficient set. The comprehensive performance score is determined based on temporal features, first weight coefficient, target semantic features, second weight coefficient, motion features, third weight coefficient, scene features, and fourth weight coefficient.

[0090] In an optional embodiment, the adjustment module 704 is specifically used for: If the overall performance score is less than the first preset threshold, reduce the video's display frame rate. If the overall performance score is greater than the second preset threshold, increase the video display frame rate; The first preset threshold is less than the second preset threshold.

[0091] Based on the same inventive concept, this application also provides an electronic device. In one embodiment, the structure of the electronic device can be as follows: Figure 8 As shown, it includes a memory 801, one or more processors 802, and a bus 803.

[0092] The memory 801 is used to store computer programs executed by the processor 802. The memory 801 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0093] Memory 801 may be volatile memory, such as random-access memory (RAM); memory 801 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 801 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 801 may be a combination of the above-described memories.

[0094] Processor 802 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 802 is used to implement the above-described flow allocation method when calling computer programs stored in memory 801.

[0095] This application embodiment does not limit the specific connection medium between the memory 801 and the processor 802 described above. This application embodiment... Figure 8 The memory 801 and the processor 802 are connected via a bus 803, and the bus 803 is in Figure 8 The text is rendered in thick lines for illustrative purposes only and should not be construed as limiting. The 803 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 8 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0096] Based on the same inventive concept, this application also provides a computer storage medium storing a computer program for causing a computer to execute any of the above-described video display methods. Since the principle by which the above-described computer-readable storage medium solves the problem is similar to that of the traffic allocation method, the implementation of the above-described computer-readable storage medium can be found in the implementation of the method; repeated details will not be elaborated further.

[0097] This application provides a video display method, apparatus, device, and storage medium. The method decodes received video information sent by a target drone to obtain multiple video frames; determines the temporal features of the multiple video frames based on TSN feature extraction and mean calculation; for each video frame, determines the target semantic features of the video frame based on at least one set of target feature coefficients obtained through a target detection model; for each video frame, determines the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; for each video frame, determines the scene features of the video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through a scene detection model; determines the comprehensive performance score of the video frame based on the preset weight coefficient set, temporal features, target semantic features, motion features, and scene features, wherein the comprehensive performance score is used to characterize the importance of the video frame; adjusts the display frame rate of the video based on the comprehensive performance score, and displays the video using the adjusted display frame rate, thereby facilitating real-time monitoring of the displayed video by the user, enabling rapid early warning and response, and improving the accuracy of risk assessment.

[0098] The present application has been described above with reference to block diagrams and / or flowcharts illustrating methods, apparatus (systems), and / or computer program products according to embodiments of the present application. It should be understood that a block of a block diagram and / or flowchart, as well as combinations of blocks of block diagrams and / or flowcharts, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, and / or other programmable data processing means to produce a machine such that the instructions, executable via the computer processor and / or other programmable data processing means, create methods for implementing the functions / actions specified in the blocks of the block diagrams and / or flowcharts.

[0099] Accordingly, this application can also be implemented using hardware and / or software (including firmware, resident software, microcode, etc.). Furthermore, this application can take the form of a computer program product on a computer-usable or computer-readable storage medium, having computer-usable or computer-readable program code implemented in the medium for use by or in conjunction with an instruction execution system. In the context of this application, a computer-usable or computer-readable medium can be any medium that can contain, store, communicate, transmit, or deliver a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0100] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A video display method, characterized in that, The method includes: The received video information from the target drone is decoded to obtain multiple video frames. Based on TSN feature extraction and mean calculation, the temporal features of the plurality of video frames are determined; and for each video frame, the target semantic features of the video frame are determined based on at least one set of target feature coefficients of the video frame obtained through a target detection model; and for each video frame, the motion features of the video frame are determined based on the pixel matrix of the video frame and the pixel matrix of the previous frame of the video frame; and for each video frame, the scene features of the video frame are determined based on the probability values ​​of a plurality of preset scenes corresponding to the video frame obtained through a scene detection model. Based on a preset set of weight coefficients, the temporal features, the target semantic features, the motion features, and the scene features, a comprehensive performance score for the video frame is determined, wherein the comprehensive performance score is used to characterize the importance of the video frame; The video display frame rate is adjusted based on the overall performance score, and the adjusted frame rate is used to display the video.

2. The method as described in claim 1, characterized in that, The determination of the temporal features of the multiple video frames based on TSN feature extraction and mean calculation includes: The multiple video frames are divided into multiple video frame groups based on preset values; For each video frame group, the target temporal features of the video frame group are obtained based on TSN feature extraction; The time series features are obtained by averaging multiple target time series features.

3. The method as described in claim 1, characterized in that, The step of determining the target semantic features of the video frame based on at least one set of target feature coefficients obtained through the target detection model includes: The video frames are input into the target detection model to obtain at least one set of target feature coefficients; The coefficients used to characterize the target category in the target feature coefficient set are vectorized, and the coefficients used to characterize the target area in the target feature coefficient set are normalized to obtain the processed target feature coefficient set. Calculate the product of the target feature coefficients in the processed target feature coefficient set; The sum of the products corresponding to each set of target feature coefficients is taken as the target semantic feature.

4. The method as described in claim 1, characterized in that, Determining the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame includes: Determine the matrix difference between the pixel matrix of the video frame and the pixel matrix of the previous frame of the video frame, wherein, when the video frame is the first video frame, the pixel matrix of the previous frame of the video frame is a preset matrix. The motion features are obtained by performing norm operations on the matrix differences.

5. The method as described in claim 1, characterized in that, The step of determining the scene features of the video frame based on the probability values ​​of multiple preset scenes corresponding to the video frame obtained through a scene detection model includes: The video frames are input into the scene detection model to obtain probability values ​​for multiple preset scenes; The probability values ​​are normalized to obtain the scene features of the video frame.

6. The method as described in claim 1, characterized in that, The determination of the comprehensive performance score of the video frame based on the preset set of weight coefficients, the temporal features, the target semantic features, the motion features, and the scene features includes: The first weight coefficient corresponding to the temporal feature is determined from the preset weight coefficient set, the second weight coefficient corresponding to the target semantic feature is determined from the preset weight coefficient set, the third weight coefficient corresponding to the motion feature is determined from the preset weight coefficient set, and the fourth weight coefficient corresponding to the scene feature is determined from the preset weight coefficient set. The comprehensive performance score is determined based on the temporal features, the first weight coefficient, the target semantic features, the second weight coefficient, the motion features, the third weight coefficient, the scene features, and the fourth weight coefficient.

7. The method according to any one of claims 1 to 6, characterized in that, The adjustment of the video display frame rate based on the comprehensive performance score includes: If the overall performance score is less than a first preset threshold, the video display frame rate is reduced. If the overall performance score is greater than the second preset threshold, the video display frame rate is increased; Wherein, the first preset threshold is less than the second preset threshold.

8. A video display device, characterized in that, The device includes: The decoding module is used to decode the video information received from the target drone to obtain multiple video frames; The feature extraction module is used to determine the temporal features of the plurality of video frames based on TSN feature extraction and mean calculation; and for each video frame, to determine the target semantic features of the video frame based on at least one set of target feature coefficients of the video frame obtained by the target detection model; and for each video frame, to determine the motion features of the video frame based on the pixel matrix of the video frame and the pixel matrix of the previous frame; and for each video frame, to determine the scene features of the video frame based on the probability values ​​of a plurality of preset scenes corresponding to the video frame obtained by the scene detection model. The determination module is used to determine the comprehensive performance score of the video frame based on a preset set of weight coefficients, the temporal features, the target semantic features, the motion features, and the scene features, wherein the comprehensive performance score is used to characterize the importance of the video frame; The adjustment module is used to adjust the display frame rate of the video based on the comprehensive performance score, and to display the video using the adjusted display frame rate.

9. An electronic device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps of the method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that, The computer storage medium stores a computer program that causes the computer to perform the method according to any one of claims 1 to 7.