Video processing method, device and electronic equipment

By extracting video transmission feature information from the network packet metadata and statistical characteristic data of the video stream and screening out important video frames for analysis, the problems of computational intensity and robustness in video intelligent analysis are solved, and efficient and low-resource consumption video processing is achieved.

CN120416539BActive Publication Date: 2025-09-26ZHEJIANG UNIVIEW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510926559.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-26
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing video intelligent analysis methods have bottlenecks in computational intensity and robustness, resulting in resource waste and difficulty in meeting real-time requirements, especially in complex scenarios where performance and reliability are poor.

Method used

By extracting video transmission feature information from the metadata and statistical characteristic data of the network data packets of the video stream, the dynamic nature of the video content is indirectly represented, and important video frames are screened out for analysis, avoiding indiscriminate processing of all video frames.

Benefits of technology

It reduces computational complexity, improves video processing efficiency and quality, reduces resource consumption, extends equipment life, and enables more efficient video analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416539B_ABST
    Figure CN120416539B_ABST
Patent Text Reader

Abstract

The present invention discloses a video processing method, apparatus, and electronic device. The method comprises: determining video transmission characteristic information of a video stream during network transmission, the video transmission characteristic information being obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream, the video transmission characteristic information being used to characterize the content dynamics of video content corresponding to video frames in the video stream; determining at least one video frame from the video stream based on the video transmission characteristic information; and performing a video analysis task on the at least one video frame. The technical solution of the present application can solve the problem of wasted resources caused by indiscriminate analysis of all video clips.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of video processing technology, and in particular to a video processing method, device, and electronic device. Background Art

[0002] With the application of video surveillance technology in many fields such as security, transportation, and industrial production, intelligent analysis of video content to achieve target detection, behavior recognition, and abnormal warning has become an industry development trend.

[0003] The process of intelligent video analysis involves a large amount of data processing and complex algorithmic operations. Its computational intensity has become a key bottleneck for resource optimization, limiting real-time performance and large-scale deployment capabilities. Related intelligent video analysis methods mainly rely on pixel-domain motion vector extraction or semantic decomposition of image content. By analyzing video pixel changes and semantic information, they filter out video clips that require in-depth processing to reduce the amount of computation. However, direct processing of video content not only consumes a large amount of computing resources, resulting in high computational complexity and difficulty meeting real-time requirements, but also has poor robustness in complex scenes, which greatly affects the performance and reliability of intelligent video analysis. Summary of the Invention

[0004] The present invention provides a video processing method, device, electronic device and storage medium to solve the problem of resource waste caused by indiscriminate analysis of all video clips.

[0005] According to one aspect of the present invention, a video processing method is provided, wherein the method comprises:

[0006] Determining video transmission characteristic information of a video stream during network transmission, the video transmission characteristic information being obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream, the video transmission characteristic information being used to characterize content dynamics of video content corresponding to video frames in the video stream;

[0007] determining at least one video frame from the video stream according to the video transmission characteristic information;

[0008] A video analysis task is performed on the at least one video frame.

[0009] According to another aspect of the present invention, a video processing device is provided, wherein the device comprises:

[0010] a determination module, configured to determine video transmission characteristic information of a video stream during network transmission, wherein the video transmission characteristic information is obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream, and the video transmission characteristic information is used to characterize the content dynamics of video content corresponding to video frames in the video stream;

[0011] a screening module, configured to determine at least one video frame from the video stream according to the video transmission characteristic information;

[0012] A processing module is configured to perform a video analysis task on the at least one video frame.

[0013] According to another aspect of the present invention, an electronic device is provided, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the video processing method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video processing method according to any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the video processing method according to any embodiment of the present invention is implemented.

[0019] The technical solution of the embodiments of the present invention, in the video stream transmission link, contains independently encoded key frames and forward predictive coded frames. By extracting multi-dimensional video transmission feature information from the metadata and statistical characteristics of the video stream network data packets, it is possible to mine content dynamics from the transmission layer. This directly obtains video transmission feature information without decoding, efficiently captures the dynamic changes of the video stream during network transmission, avoids direct analysis of complex video content, and greatly reduces computational complexity. In terms of video frame determination, the dynamics of video content are indirectly characterized based on video transmission feature information, and video frames that are critical to subsequent processing tasks can be accurately screened. Compared with blindly processing all video frames, this avoids the waste of resources caused by indiscriminate analysis of all video clips, significantly reduces unnecessary processing, and improves video processing efficiency. When performing video analysis tasks, only the useful video frames that are screened need to be processed in a targeted manner, making the video processing results more in line with actual needs and improving video processing quality. At the same time, accurately processing useful video frames can also reduce system resource consumption, reduce the burden on computing equipment, and extend the service life of equipment, allowing more tasks to be analyzed with the same amount of computing resources.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 is a flowchart of a video processing method provided according to an embodiment of the present invention;

[0023] Figure 2 is a flowchart of another video processing method provided according to an embodiment of the present invention;

[0024] Figure 3 is a flowchart of another video processing method provided according to an embodiment of the present invention;

[0025] Figure 4 2 is a schematic structural diagram of a video processing method and apparatus provided by an embodiment of the present invention;

[0026] Figure 5 The figure is a schematic structural diagram of an electronic device for implementing the video processing method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] Figure 1 A flowchart of a video processing method is provided for an embodiment of the present invention. This embodiment is applicable to situations where useful video frames in a video stream are screened and then a video analysis task is performed. The method can be executed by a video processing device, which can be implemented in the form of hardware and / or software. The video processing device can be configured in any electronic device with network communication capabilities.

[0030] like Figure 1 As shown, the video processing method provided by the embodiment of the present invention may include the following processes:

[0031] S110. Determine video transmission characteristic information of the video stream during network transmission. The video transmission characteristic information is obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream. The video transmission characteristic information is used to characterize the content dynamics of the video content corresponding to the video frames in the video stream.

[0032] The multiple video frames in a video stream include independently encoded key frames and forward-predicted coded frames. Essentially, a video stream consists of a series of continuous video frames. A video frame is a single static image within a video stream and is the smallest unit of video. A video stream is a dynamic image formed by the sequential transmission of multiple video frames in time. Video streaming refers to the transmission of video content over a network in the form of a continuous data stream. Video streaming uses streaming protocols and encoding compression technologies to segment video into small data units (such as video frames or GOP groups), which are then transmitted sequentially for real-time decoding and playback.

[0033] For video streams encoded using a preset inter-frame compression technology, the corresponding video sequence will consist of several GOPs (Group of Pictures). The preset inter-frame compression technology may include, but is not limited to, at least one of the following: H.264, H.265, VP9, ​​AV1, MPEG-2, and H.266 / VVC. Furthermore, given the high real-time requirements of the security field, bidirectional predictive coding frames will not be used. As a result, the first video frame in each GOP corresponding to the video stream is a key frame, and all video frames following the key frame in the same GOP in the playback order are forward-predictive coded frames. On this basis, the multiple video frames of the video stream transmitted over the network in this case include independently encoded key frames and forward-predictive coded frames.

[0034] Some solutions require video analysis of a subset of frames in a video stream. This typically involves extracting pixel-domain motion vectors from each frame or analyzing the semantics of the image content. By analyzing pixel variations and semantic information, these techniques filter out frames requiring deep processing from the multiple frames in the stream. However, directly processing video content consumes significant computing resources, resulting in high computational complexity and difficulty meeting real-time requirements. Furthermore, direct processing of video content suffers from poor robustness in complex scenarios.

[0035] To this end, the present application scheme takes into account that there are multiple dimensions of video transmission characteristics during network transmission of video streams, which contain the content dynamics of the video content corresponding to the video frames in the video stream. Therefore, this part of the video transmission feature information can be directly obtained. Although the video transmission feature information does not directly describe the video changes corresponding to the video frames, the metadata and statistical characteristic data of the network data packets corresponding to the video stream contain information on the dynamics of the video content. Therefore, the video transmission feature information can reflect the dynamic change characteristics of the video content without relying on the characteristics of content semantics or visual or auditory feature analysis. In other words, the present application adopts the above scheme without the need to parse the video content semantics or visual features of the video stream. It can achieve real-time analysis of the dynamics of the video content only through the metadata and statistical characteristic data of the network data packets, providing a solution for lightweight processing of video stream processing.

[0036] Optionally, the metadata of the network data packet can be described using at least one of the packet header information and the protocol field. The statistical characteristic data of the network data packet can be described using at least one of the traffic fluctuation and the time series distribution. The dynamics of the video content corresponding to the video frames in the video stream can be described using the video complexity and scene switching frequency generated by the temporal changes of the video content.

[0037] Optionally, the dynamics of the video content corresponding to the video frames in the video stream refers to the degree of change in the video content in time or space, which may include but is not limited to the degree of change in the following: the movement speed and range of objects in the video screen, the sudden change in the screen content caused by the lens switching (such as switching from an indoor scene to an outdoor scene), and the fluctuation amplitude of the pixel value in the video screen.

[0038] Optionally, the dynamics of the video content corresponding to the video frame in the video stream may refer to the change, movement or uncertainty characteristics of the video content corresponding to the video frame (such as image, video, text, audio, etc.) at the temporal, spatial or semantic level, which can be used to measure the degree of change of the video content caused by the change of the video content over time.

[0039] Optionally, the video transmission feature information may reflect changes in the temporal sequence of video frames, such as changes in at least one of optical flow, inter-frame difference, and motion vector. The video transmission feature information may also describe statistical characteristic data of the video corresponding to the video frame, such as at least one of brightness mean, color distribution variance, and texture complexity.

[0040] Video transmission feature information can be indirect data representing the dynamic content of video frames in a video stream at the transport and coding layers of the video stream. The coding layer data representing the video stream can be encoded data representing the video content (e.g., pixels, sound waves) compressed into a binary stream. The transport layer data representing the video stream can be metadata representing the data packets encapsulated from the encoded video stream and transmitted over the network. This data is not concerned with content semantics and only addresses data segmentation, timing, and reliability.

[0041] S120: Determine at least one video frame from the video stream according to the video transmission characteristic information.

[0042] When filtering multiple video frames included in a video stream, there are redundant invalid video frames in the multiple video frames, and since the video content corresponding to the video frames that need to be filtered out from the multiple video frames is usually video content with some content dynamics, the content dynamics reflected by the video transmission feature information can be used as a basis to implement filtering processing on the multiple video frames of the video stream, so as to filter out video frames that match the content dynamics reflected by the video transmission feature information from the multiple video frames, thereby reducing the video frames that need to be processed.

[0043] Furthermore, in the field of video processing, video transmission feature information, reflecting the dynamic nature of video content, no longer relies on decoding video frames. Instead, it captures indirect representations of content dynamics through non-decoding methods, avoiding the high computational resource consumption and poor real-time performance of this process. Specifically, by avoiding direct decoding of video frames, the focus is instead on the tightly coupled relationship between the data packets generated by the video stream at the transport layer and the video encoding behavior. During the video encoding process, different content dynamics lead to differences in encoding behavior, which are then reflected in the characteristics of the transport layer data. For example, when a video scene changes dramatically, contains a large number of moving objects, or has frequent scene switching, the encoder generates more key frames or increases the amount of encoded data to maintain video quality. These changes are directly reflected in statistical characteristics such as the size, number, and transmission frequency of network data packets. In contrast, video streams with static scenes have relatively stable data packets with uniform transmission intervals.

[0044] Based on this coupling relationship, while video transmission feature information doesn't directly reveal the video's content, through mathematical modeling and self-learning algorithm analysis, it can effectively reflect the underlying patterns of scene changes and indirectly characterize the dynamic nature of the video content. For example, by analyzing the variance of packet sizes over a period of time, it can be determined whether the video is rapidly switching; and based on the regularity of packet transmission intervals, it can be inferred whether the video is static or slowly changing.

[0045] The above-mentioned indirect representation method through non-decoding not only greatly reduces the computational complexity, enabling it to run quickly on resource-constrained devices (such as mobile terminals and embedded devices), but also can perceive dynamic changes in video content at a near real-time speed, providing strong support for subsequent video frame screening, and greatly improving the overall performance and efficiency of the video processing system.

[0046] S130: Perform a video analysis task on at least one video frame.

[0047] After obtaining at least one representative video frame from the video stream, a video analysis task can be performed on the at least one selected video frame. The video analysis task can refer to a series of standardized or customized operations that are pre-defined in video processing based on specific application requirements and automatically analyze the video content in the video frame.

[0048] Optionally, the video analysis task can parse the visual information of the video frame (for example, including but not limited to image pixels, target features, motion trajectories, etc.) through an algorithm or model, and output structured data or decision results.

[0049] In one optional example, from the perspective of visual information extraction, video analysis tasks can include target detection and recognition tasks, such as using computer vision algorithms to accurately locate and identify at least one of the captured targets, including pedestrians, vehicles, and objects, in a video frame. For example, in traffic surveillance videos, vehicles violating traffic regulations can be quickly detected, while image segmentation tasks can segment different captured objects or captured areas in a video frame. The video analysis task is performed using at least one deep learning model such as YOLO and FasterR-CNN. In another optional example, from the perspective of content understanding, the video analysis task can be used to continuously analyze the movements, postures, and trajectories of captured targets in video frames to determine whether the captured targets have abnormal behavior. For example, in security monitoring, it can identify captured targets with suspicious movements such as climbing over a wall. The video analysis task can also be used to determine the scene category of a video frame, such as whether the video frame is indoors or outdoors, in a shopping mall or on the street. The video analysis task can be combined with spatiotemporal feature extraction algorithms, such as 3D convolutional neural networks.

[0050] In one optional example, a video analysis task can include video quality assessment, which quantitatively evaluates the subjective and objective quality of a video by analyzing at least one of the following metrics: clarity, color reproduction, and noise level. Video quality assessment can be performed using at least one of the following metrics: PSNR and SSIM. A video analysis task can be used to extract key information from multiple video frames and generate a concise video summary, allowing users to quickly understand the core content of the video. This is commonly used in the processing of news videos and conference recordings.

[0051] Understandably, the core objective of this application is to accurately extract and output highly salient video frames. Highly salient video frames can be defined as those with prominent visual features, high information density, and critical value for understanding the content within a video sequence. The aforementioned optional solutions for the video analysis task are provided as examples only, demonstrating the technical feasibility of this application's solution.

[0052] The technical solution of the embodiment of the present invention, in the video stream transmission link, the video stream contains independently encoded key frames and forward predictive coded frames. By extracting multi-dimensional video transmission feature information from the metadata and statistical characteristics of the video stream network data packets, it is possible to mine content dynamics from the transmission layer. This directly obtains video transmission feature information without decoding, efficiently captures the dynamic changes of the video stream during network transmission, avoids direct analysis of complex video content, and greatly reduces computational complexity. In terms of video frame determination, the video content dynamics are characterized based on video transmission feature information, and video frames that are critical to subsequent processing tasks can be accurately screened. Compared with blindly processing all video frames, this avoids the waste of resources caused by indiscriminate analysis of all video clips, significantly reduces unnecessary processing, and improves video processing efficiency. When performing video analysis tasks, only the useful video frames that are screened need to be processed in a targeted manner, making the video processing results more in line with actual needs and improving video processing quality. At the same time, accurately processing useful video frames can also reduce system resource consumption, reduce the burden on computing equipment, and extend the service life of equipment, so that more tasks can be analyzed with the same amount of computing resources.

[0053] Figure 2 A flow chart of another video processing method provided for an embodiment of the present invention. The technical solution of this embodiment further optimizes the process of determining the video transmission feature information of the video stream during network transmission in the aforementioned embodiment on the basis of the technical solution of the aforementioned embodiment. This embodiment can be combined with various optional solutions in one or more of the aforementioned embodiments.

[0054] like Figure 2 As shown, the video processing method provided by the embodiment of the present invention may include the following processes:

[0055] S210: Determine video transmission parameter information of a video stream during network transmission. The video transmission parameter information includes at least one of the following: a byte size of a data packet during network transmission of the video stream, real-time protocol metadata during network transmission of the video stream, and a number of video frame fragments during network transmission of the video stream. The video transmission parameter information is parameter information extracted from metadata and statistical characteristic data of network data packets corresponding to the video stream transmitted over the network.

[0056] S220: Determine video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information.

[0057] Among them, the multiple video frames of the video stream include independently encoded key frames and forward prediction encoded frames. The video transmission feature information is obtained based on the metadata and statistical characteristic data of the network data packets corresponding to the video stream. The video transmission feature information is used to characterize the content dynamics of the video content corresponding to the video frames in the video stream.

[0058] Based on the above embodiment, optionally, the video transmission characteristic information is represented by at least one of data packet byte quantity disturbance information, real-time transport protocol metadata anomaly information, bit rate dynamic evolution information and fragment entropy state change information: wherein, the data packet byte quantity disturbance information is used to indicate the unexpected fluctuation change of the data packet byte quantity per unit time when the video stream is transmitted over the network; the real-time transport protocol metadata anomaly information is used to indicate the data packet loss per unit time when the video stream is transmitted over the network, and the data packet loss is manifested as discontinuous sequence number of the data packet or timestamp jump; the bit rate dynamic evolution information is used to indicate the dynamic change of the data transmission bit rate per unit time over time when the video stream is transmitted over the network; the fragment entropy state change information is used to indicate the fluctuation of the number of video frame fragments per unit time when the video stream is transmitted over the network.

[0059] Packet byte volume disturbances refer to abnormal fluctuations or unexpected changes in the packet byte volume during network transmission. These fluctuations or unexpected changes can be caused by a variety of factors, ranging from normal fluctuations in the network environment to failures, attacks, or human intervention. Packet byte volume refers to the total number of data bytes contained in a single packet, typically measured in bytes. The variance or mutation frequency of packet byte volume indirectly reflects the dynamic nature of the video content corresponding to a video frame. For example, the greater the variance of packet byte volume, the more dynamic the video content corresponding to the video frame. Conversely, the smaller the variance of packet byte volume, the less dynamic the video content corresponding to the video frame.

[0060] The packet byte size is directly related to the data output of the video encoder. In static scenes, the encoder tends to generate a stable low-data-volume video stream with small byte-volume fluctuations; however, in dynamic scenes (such as object movement or scene switching), the encoder needs to generate more data (such as I-frames or complex P-frames), resulting in significant byte-volume fluctuations. Therefore, the packet byte size perturbation information As a primary feature, it can quickly screen out potentially dynamic segments. From an information theory perspective, fluctuations in byte size can be viewed as changes in information entropy, with the high entropy of dynamic scenes contrasting sharply with the low entropy of static scenes. Alternatively, when video scenes change dramatically (such as high-speed motion or complex textures), the encoded frame data volume is larger (requiring more bits to describe details), resulting in an increase in the byte size of a single data packet or an increase in the number of data packets per unit time.

[0061] Real-time Transport Protocol metadata includes timestamps and sequence numbers. The timestamp identifies the starting point of the sampled data in the packet; the sequence number is an incremental number associated with each packet corresponding to the real-time transport protocol metadata and is used to detect packet loss and reassembly. The volatility of the timestamp intervals and the frequency of sequence number jumps associated with the real-time transport protocol metadata are correlated with the dynamic nature of the video content corresponding to the video frames. Real-time transport protocol metadata anomalies can indicate packet loss per unit time during network transmission of the video stream.

[0062] Real-time Transport Protocol metadata reflects the real-time transmission characteristics of the video stream. Dynamic scenes are prone to packet loss, which is manifested as discontinuous sequence numbers or timestamp jumps. Extracting abnormal information from the Real-time Transport Protocol metadata It can detect transport layer anomalies directly related to content dynamics. Packet loss occurs when some packets fail to reach the receiving end during transmission. By checking the sequence numbers or timestamps of the packets, the receiving end can detect data transmission anomalies, which manifest as discontinuous sequence numbers or timestamp jumps.

[0063] Bitrate dynamic evolution can be used to describe the dynamic change of the bitrate of video data transmission over time when the video stream is transmitted over the network. Bitrate is a comprehensive representation of the data throughput when the video stream is transmitted over the network. The content dynamics of the video frame corresponding to the video content (such as fast motion or scene switching) will cause significant fluctuations in the amount of data output by the encoder. Assuming that variable bitrate encoding is used by default, the bitrate dynamic evolution information can be extracted by The purpose is to quantify this dynamic evolution and provide a more macro trend indicator than byte volume disturbance.

[0064] Video frame fragmentation refers to the process of packetizing and transmitting video frames due to data packet size limitations or protocol stack processing. The number of video frame fragments during network transmission of a video stream is closely related to the data volume and encoding complexity of the video stream. In dynamic scenes, an increase in data volume may cause a significant change in the number or distribution of video frame fragments, while the fragmentation pattern of static scenes tends to be stable. For example, dynamically complex video content can contain areas of different complexity (such as a moving subject and a static background in the picture at the same time), resulting in large differences in fragment sizes (some fragments contain high-dynamic data and a large number of bytes; some contain low-dynamic data and a small number of bytes). Therefore, by extracting the entropy state change information of the fragments It can reflect the deep response of the transport layer to the dynamic content of the video content corresponding to the video frame.

[0065] Shard entropy is a measure of the degree of disorder or uncertainty in a system. In the context of data sharding, shard entropy can be expressed as the degree of disorder or uncertainty in the distribution of data within each shard and across shards. If data is evenly and regularly distributed across the shards, the shard entropy is low, indicating good organization and relative order. If data is haphazardly distributed, with significant differences between shards and no discernible pattern, the entropy is high, indicating greater uncertainty.

[0066] Increasing the number of shards will increase the entropy of the shards. This is because as the number of shards increases, the data is dispersed into more sub-parts, and the possibilities and combinations of data distribution will also increase significantly. This makes the overall distribution of data more complex and uncertain, leading to an increase in entropy. Reducing the number of shards may reduce the entropy. Fewer shards means that the data is more concentrated, and the correlation and regularity between data are easier to discover and grasp, which will reduce the degree of chaos and uncertainty in the system.

[0067] The parameters mentioned above, such as packet byte disturbance information, real-time transport protocol metadata anomaly information, bit rate dynamic evolution information, and fragment entropy state change information, imply the changing patterns of the video content itself. There is no need to analyze the content semantics or visual features (such as target detection and speech recognition). Real-time acquisition of content dynamics can be achieved only through the metadata and statistical characteristic data of the network data packets.

[0068] As an optional but non-limiting implementation, when the video transmission parameter information includes the byte size of data packets of the video stream during network transmission, determining the video transmission characteristic information of the video stream during network transmission based on the video transmission parameter information may include the following steps:

[0069] Determine the first parameter information based on the video transmission parameter information, and the first parameter information is used to indicate the packet byte quantity of each data packet per unit time when the video stream is transmitted over the network; determine the second parameter information based on the first parameter information, and the second parameter information is used to indicate the average packet byte quantity per unit time when the video stream is transmitted over the network; determine the packet byte quantity disturbance information of the video stream during network transmission based on the first parameter information and the second parameter information.

[0070] By capturing the number of bytes in each packet of a video stream during network transmission in real time, we can obtain the packet byte count S(t). Optionally, we can use a sliding timer to dynamically set the video frame rate (here, the default is N=25, and the commonly used frame rate is 25fps) to collect time series data. This will determine the packet byte count S(t) per unit time of the video stream during network transmission. The packet byte count S(t) is a direct representation of the video stream at the transport layer. Changes in the packet byte count S(t) reflect the data output behavior of the video encoder in different scenarios.

[0071] In an optional solution, before determining the second parameter information according to the first parameter information, the following step may be further included: performing exponential smoothing filtering on the packet byte amount of each data packet per unit time when the video stream indicated by the first parameter information is transmitted over the network.

[0072] Specifically, exponential smoothing filtering can be used to suppress the instantaneous jitter noise in network transmission when the video stream is transmitted over the network to improve the accuracy. is the smoothed value of the data packet byte volume at the current moment, S(t) is the original value of the data packet byte volume at the current moment, is the smoothed value of the packet byte size at the previous moment, is the smoothing factor, which is dynamically set according to the encoding format and scene complexity. Here, α can be set to 0.3. The calculation formula of the exponential smoothing filter is as follows:

[0073] .

[0074] After obtaining the packet byte volume of each packet per unit time when the video stream is transmitted over the network, calculate the average packet byte volume per unit time , and then we can calculate the deviation of the data packet byte volume of each data packet per unit time from the average value, and square the calculated deviation; sum and average the squares of all deviations to get the standard deviation of the byte volume fluctuation per unit time , namely, the packet byte quantity disturbance information, the packet byte quantity disturbance information This provides a low-cost preliminary saliency indicator, avoiding high-entropy analysis of all fragments. The calculation formula for the packet byte volume disturbance information is as follows:

[0075] .

[0076] As an optional but non-limiting implementation, when the video transmission parameter information includes real-time transport protocol metadata when the video stream is transmitted over the network, determining the video transmission characteristic information of the video stream during network transmission based on the video transmission parameter information may include the following steps:

[0077] The third parameter information is determined based on the video transmission parameter information, where the third parameter information is used to indicate the timestamp and sequence number in the header of the real-time transport protocol metadata of multiple consecutive data packets of the video stream extracted from the real-time transport protocol metadata during video stream transmission; the fourth parameter information is determined based on the third parameter information, where the fourth parameter information is used to indicate the number of timestamp jumps and the number of discontinuous sequence numbers per unit time during network transmission of the video stream; and the real-time transport protocol metadata abnormality information of the video stream during network transmission is determined based on the fourth parameter information.

[0078] The real-time transport protocol metadata headers of multiple consecutive data packets of the video stream indicated by the real-time transport protocol metadata are parsed to extract the timestamp TS(t) and sequence number Seq(t), and the continuity and regularity of the extracted multiple timestamps TS(t) and sequence numbers Seq(t) are detected. In this way, the number of timestamp jumps and the number of discontinuous sequence numbers per unit time in the network transmission of the video stream can be obtained, and the abnormal information of the real-time transport protocol metadata of the video stream during the network transmission can be determined.

[0079] In an optional solution, determining the fourth parameter information according to the third parameter information may include the following steps:

[0080] For the timestamps and sequence numbers in the real-time transport protocol metadata headers of two adjacent data packets among the multiple consecutive data packets indicated by the third parameter information, fifth parameter information is determined based on the timestamps and sequence numbers in the real-time transport protocol metadata headers of the two adjacent data packets, and the fifth parameter information is used to indicate the timestamp increment in the real-time transport protocol metadata headers of the two adjacent data packets and the sequence number difference in the real-time transport protocol metadata headers of the two adjacent data packets; and fourth parameter information is determined based on the fifth parameter information.

[0081] The timestamp and sequence number in the Real-time Transport Protocol metadata header of two adjacent data packets are used to calculate the timestamp increment in the Real-time Transport Protocol metadata header of two adjacent data packets. The difference in sequence numbers between the real-time transport protocol metadata headers of two adjacent data packets , detect anomalies (such as jumps or packet loss). Count the number of timestamp jumps per unit time (like , 25fps standard interval 40ms) and the number of discontinuous sequence numbers (like ),Will Divide it by the length of time per unit time to get the abnormal frequency per unit time. , you can get the real-time transport protocol metadata anomaly information of the video stream during network transmission, As a secondary feature, it reflects the abnormal characteristics of the video clip and is used to reduce the significant features of the analysis. The calculation formula of the real-time transport protocol metadata abnormal information is expressed as follows:

[0082] .

[0083] As an optional but non-limiting implementation, when the video transmission parameter information includes a video data transmission bit rate when the video stream is transmitted over a network, determining the video transmission characteristic information of the video stream during network transmission based on the video transmission parameter information may include the following steps:

[0084] First parameter information is determined based on the video transmission parameter information, where the first parameter information is used to indicate the packet byte amount of each data packet per unit time when the video stream is transmitted over the network; sixth parameter information is determined based on the first parameter information, where the sixth parameter information is used to indicate the video data transmission bit rate when the video stream is transmitted over the network; and seventh parameter information is determined based on the sixth parameter information to obtain dynamic evolution information of the bit rate when the video stream is transmitted over the network, where the seventh parameter information is used to indicate the rate of change of the video data transmission bit rate per unit time when the video stream is transmitted over the network.

[0085] When calculating the instantaneous bit rate based on the number of data packets, the total number of data packets transmitted per unit time, S(i), is calculated. The total number of data packets transmitted per unit time, S(i), is divided by the length of the unit time to obtain the instantaneous bit rate. The formula is: (Unit: kbps).

[0086] In an optional solution, before determining the seventh parameter information according to the sixth parameter information, the method further includes the following steps: performing a sliding average filtering on the video data transmission bit rate when the video stream indicated by the sixth parameter information is transmitted over the network.

[0087] Because the video stream may have some short-term jitter during network transmission, in order to improve the accuracy, a sliding average filter can be used to perform a sliding average filter on the video data transmission bit rate when the video stream is transmitted over the network. (Smooth the results of the last 5 seconds, that is, the window The window size of the sliding average filter can be dynamically adjusted according to the scene and task type of the video frame. The calculation formula of the sliding credential filter is as follows:

[0088] .

[0089] In order to quantify the video data transmission bit rate when the smoothed video stream is transmitted over the network The bit rate evolution coefficient can be calculated by calculating the rate of change of the bit rate. Specifically, the bit rate deviation of the video data transmission during network transmission of the video stream in each unit time and the previous unit time of each unit time can be calculated and divided by the time length of the unit time to obtain the bit rate change rate characteristic. , you can get the bit rate dynamic evolution information of the video stream when it is transmitted over the network, the bit rate dynamic evolution information of the video stream when it is transmitted over the network As the third-level feature, it is used to verify the trend consistency of the first two levels of features to avoid local noise misleading decision-making. The calculation formula for the dynamic evolution of the bit rate of the video stream during network transmission is expressed as follows:

[0090] .

[0091] As an optional but non-limiting implementation, when the video transmission parameter information includes the number of video frame fragments when the video stream is transmitted over the network, determining the video transmission characteristic information of the video stream during network transmission based on the video transmission parameter information may include the following steps:

[0092] The eighth parameter information is determined based on the video transmission parameter information, and the eighth parameter information is used to indicate the number of video frame fragments transmitted at each time point per unit time when the video stream is transmitted over the network; the ninth parameter information is determined based on the eighth parameter information, and the ninth parameter information is used to indicate the average number of video frame fragments transmitted per unit time when the video stream is transmitted over the network; and the fragment entropy state change information of the video stream during network transmission is determined based on the eighth parameter information and the ninth parameter information.

[0093] Count the number of video frame fragments P(t) transmitted per unit time at each time point during network transmission of the video stream (based on the fragmentation rule of maximum transmission unit MTU = 1500 bytes). In order to obtain the characteristics of the fragment entropy change , it is necessary to process the number of video frame fragments that the video stream is divided into during network transmission step by step, and obtain the change of the number of video frame fragments over a period of time. First, select the unit time (the default N=25 in one second), and calculate the number of video frame fragments at each time point in the unit time , calculate the average number of video frame fragments per unit time , calculate the shard variance at all time points in unit time , you can get the fragment entropy state change information when the video stream is transmitted over the network, the fragment entropy state change information As the final verification feature, the entropy state change of the fragmentation mode is comprehensively evaluated. The calculation formula of the fragment entropy state change information when the video stream is transmitted over the network is expressed as follows:

[0094] .

[0095] S230: Determine at least one video frame from the video stream according to the video transmission characteristic information.

[0096] Optionally, after obtaining video transmission characteristic information represented by at least one of packet byte volume disturbance information, real-time transport protocol metadata anomaly information, bit rate dynamic evolution information, and slice entropy state transition information, the video transmission characteristic information of the video stream during network transmission can be stored. The video transmission characteristic information of the video stream during network transmission can be represented as: .

[0097] The purpose of storing the video transmission feature information during network transmission of the video stream is to preserve historical data and support subsequent feedback correction. The 24-hour retention period balances storage costs and correction requirements. The sliding window ensures real-time performance and adapts to the continuous operation characteristics of the monitoring system. Specifically, the video transmission feature information F(t) during network transmission of the video stream can be distributed stored according to the timestamp t, with a default retention period of 24 hours. The data structure is , records older than 24 hours are automatically cleaned up to free up storage space; sliding windows (here the default is 30 seconds) are dynamically maintained based on the type of motion event for real-time analysis, and data within the window is indexed by timestamp.

[0098] S240: Perform a video analysis task on at least one video frame.

[0099] The technical solution of the embodiments of the present invention, in the video stream transmission link, contains independently encoded key frames and forward predictive coded frames. By extracting multi-dimensional video transmission feature information from the metadata and statistical characteristics of the video stream network data packets, it is possible to mine content dynamics from the transmission layer. This directly obtains video transmission feature information without decoding, efficiently captures the dynamic changes of the video stream during network transmission, avoids direct analysis of complex video content, and greatly reduces computational complexity. In terms of video frame determination, the dynamics of video content are indirectly characterized based on video transmission feature information, and video frames that are critical to subsequent processing tasks can be accurately screened. Compared with blindly processing all video frames, this avoids the waste of resources caused by indiscriminate analysis of all video clips, significantly reduces unnecessary processing, and improves video processing efficiency. When performing video analysis tasks, only the useful video frames that are screened need to be processed in a targeted manner, making the video processing results more in line with actual needs and improving video processing quality. At the same time, accurately processing useful video frames can also reduce system resource consumption, reduce the burden on computing equipment, and extend the service life of equipment, allowing more tasks to be analyzed with the same amount of computing resources.

[0100] Figure 3 A flow chart of another video processing method provided for an embodiment of the present invention. The technical solution of this embodiment further optimizes the process of determining at least one video frame from a video stream based on video transmission feature information in the aforementioned embodiment on the basis of the technical solution of the aforementioned embodiment. This embodiment can be combined with various optional solutions in one or more of the aforementioned embodiments.

[0101] like Figure 3 As shown, the video processing method provided by the embodiment of the present invention may include the following processes:

[0102] S310. Determine video transmission feature information of a video stream during network transmission. The multiple video frames of the video stream include independently encoded key frames and forward prediction encoded frames. The video transmission feature information is obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream. The video transmission feature information is used to characterize the content dynamics of the video content corresponding to the video frames in the video stream.

[0103] S320: Determine the task type corresponding to the video analysis task.

[0104] S330. Determine a preset feature matrix adapted for the video analysis task according to the task type corresponding to the video analysis task. The preset feature matrix is ​​used to indicate the weights of transmission feature parameters of different dimensions in the video transmission feature information when screening video frames from the video stream. The weights indicated by the preset feature matrix are used to reflect the contribution of transmission feature parameters of different dimensions in the video transmission feature information when measuring the content dynamics of the video content corresponding to the video frames in the video stream.

[0105] The task type of a video analysis task can refer to the classification of the task functions to be performed by the video analysis task. For example, the video analysis task corresponding to target detection belongs to the target recognition category, and the video analysis task corresponding to behavior analysis belongs to the dynamic behavior analysis category. The values ​​of each element in the preset feature matrix represent the weights of transmission feature parameters of different dimensions in the video transmission feature information. The weights reflect the contribution of the feature parameters to measuring the dynamic nature of the video content. Video transmission feature information is extracted from the metadata and statistical characteristic data of the video stream network data packets through non-decoding methods to characterize the dynamic nature of the video content.

[0106] After determining the video analysis task, the task type corresponding to the specific video analysis task can be determined, and a preset feature matrix that matches the task type corresponding to the video analysis task can be selected from a pre-configured database. Different task types have different focuses on the dynamics of video content, and therefore require different feature parameter weights. For example, for target detection tasks, you may be more concerned about the sudden appearance or movement of objects in the video. In this case, the weights of feature parameters related to sudden changes in data packets should be set higher; while for scene classification tasks, you may be more concerned about the overall stability of the video, and the weights of parameters related to the long-term statistical characteristics of data packets are more important. According to the task type, a preset feature matrix that is adapted to it is selected or generated. The matrix specifies the weights of the transmission feature parameters of each dimension in the video transmission feature information when screening video frames.

[0107] By associating the task type with a preset feature matrix using the above solution, video frames can be targeted and screened based on specific task requirements, avoiding blindly processing all video frames, reducing unnecessary computation, and improving processing efficiency. Different video analysis tasks have different requirements for computing resources. Reasonable setting of the feature matrix weights allows the system to focus resources on video frames that are important to the task, avoiding resource waste. On resource-constrained devices (such as edge computing devices), this approach can reduce device load and extend device life while ensuring task processing effectiveness. Furthermore, because the weights of the preset feature matrix reflect the contribution of feature parameters to measuring the dynamic nature of video content, selecting an appropriate feature matrix based on the task type can more accurately capture the video content features relevant to the task, thereby improving the accuracy and reliability of video analysis tasks.

[0108] S340: Determine at least one video frame from the video stream according to the preset feature matrix adapted for the video analysis task and the video transmission feature information.

[0109] S350: Perform a video analysis task on at least one video frame.

[0110] As an optional but non-limiting implementation scheme, determining at least one video frame from a video stream based on a preset feature matrix adapted for a video analysis task and video transmission feature information may include the following steps:

[0111] According to the weights in the preset feature matrix adapted for the video analysis task, the decision results of the video frames in the video stream are obtained by weighted processing of the transmission feature parameters of different dimensions in the video transmission feature information. The decision results of the video frames in the video stream are used to indicate the possibility of the video frames in the video stream participating in the video analysis task; at least one video frame is determined from the video stream according to the decision results of the video frames in the video stream.

[0112] The preset feature matrix is ​​a matrix set in advance for different video analysis tasks. The elements in the matrix represent the weights of the transmission feature parameters of each dimension of the video transmission feature information, reflecting the importance of the transmission feature parameters of each dimension of the video transmission feature information in measuring the dynamics of the video content.

[0113] After determining the weights in the preset feature matrix suitable for the video analysis task, the transmission feature parameters of each dimension in the video transmission feature information can be weighted according to the corresponding weights in the preset feature matrix. The weighted processing is then used to obtain the decision results for the video frames in the video stream to comprehensively evaluate the transmission characteristics of the video frames. The decision results for the video frames in the video stream can quantify the content dynamics of each video frame in the video stream through weighted processing, which is used to measure the likelihood of the video frame participating in the video analysis task. The higher the value, the greater the likelihood.

[0114] The transmission feature parameters of different dimensions are extracted from the video transmission feature information. Then, referring to the preset feature matrix adapted for the video analysis task, the transmission feature parameters of each dimension in the video transmission feature information are multiplied by their corresponding weight values. The weighted values ​​of the transmission feature parameters of all dimensions are then added together to obtain the decision result for each video frame. This decision result is a quantitative indicator that comprehensively reflects the relevance of the video frame to the video analysis task.

[0115] After obtaining the decision results for all video frames, an appropriate filtering strategy is set, such as selecting a certain number of video frames with the highest decision results, or selecting video frames with decision results greater than a certain threshold, and determining them as target video frames for the video analysis task. For example, when performing a vehicle target detection task, the decision results can be used to filter out video frames with a high probability of vehicle movement or presence for subsequent analysis.

[0116] Feature Matrix By weighted fusion of multi-dimensional features, the significance is comprehensively evaluated, and the initial value can be set based on experience to reflect the contribution of each feature to the dynamics. For example, the preset feature matrix adapted to the video analysis task is selected according to the task type corresponding to the video analysis task. : .in, is the weight of the perturbation information of the byte amount of the data packet; The weight of the real-time transport protocol metadata anomaly information; The weight of the bit rate dynamic evolution information; is the weight of the slice entropy state change information; the feature vector corresponding to the video transmission feature information when the video stream is transmitted over the network within a unit time The decision result scores of the video frames in the video stream obtained by weighted calculation are as follows: , S(t) is the decision result of the video frame in the video stream, the threshold Control the trigger sensitivity to ensure that only high-saliency video clips are screened out as at least one video frame for analysis. Then, the decision result score threshold can be pre-associated with the task type corresponding to the video analysis task, for example ,like , then trigger the execution of the video analysis task for at least one video frame, that is, push the video segment of at least one video frame to execute the video analysis task, otherwise skip.

[0117] This optional solution, which weights video transmission feature information using a preset feature matrix, fully considers the specific video content characteristics of different tasks and accurately selects frames from the video stream that are highly relevant to the task. Compared to randomly or indiscriminately processing all video frames, this significantly improves the accuracy and specificity of the selection process. Furthermore, it reduces the processing of irrelevant video frames, significantly reducing the computational load and processing time. Furthermore, the selected video frames are more suitable for the video analysis task, helping to improve the accuracy and reliability of the analysis results.

[0118] As an optional but non-limiting implementation solution, after determining at least one video frame from the video stream according to the video transmission characteristic information, the following steps are further included:

[0119] Determine a feedback result of at least one video frame, where the feedback result of at least one video frame is used to indicate whether to increase the sensitivity of action information related to the task type corresponding to the video analysis task in the video content included in the video frame, or whether to reduce the sensitivity of action information unrelated to the task type corresponding to the video analysis task in the video content included in the video frame; and update a preset feature matrix adapted for the video analysis task in response to the feedback result of at least one video frame.

[0120] After analyzing at least one screened video frame, an evaluation conclusion on the relevance of the video content to the video analysis task is obtained, which is used to indicate whether the sensitivity to the specified action information needs to be adjusted. Task-type-related action information can be dynamic change information of objects or scenes in the video frame, including motion trajectory, posture changes, etc., which can be divided into two categories: relevant to the task type and irrelevant to the task type. For example, in a pedestrian detection task, the movement of pedestrians is relevant action information, while the movement of leaves in the background is irrelevant action information. The sensitivity of task-type-related action information can refer to the sharpness of the video analysis task in recognizing the specified action information in the video content. The higher the sensitivity of task-type-related action information, the easier it is to capture the specified action information related to the task type; the lower the sensitivity of task-type-related action information, the more action information irrelevant to the task type will be retained.

[0121] The feature matrix is ​​iteratively optimized through a positive and negative feedback mechanism, and the feedback result of at least one video frame is configured according to the task execution result of performing a video analysis task on at least one video frame. Then, the weight of the preset feature matrix is ​​dynamically adjusted according to the feedback result of at least one video frame. A method for self-learning and dynamic adjustment of the feature matrix is ​​implemented through a feature storage module and a feedback correction module to improve the accuracy of saliency discrimination and adaptability to different scenarios.

[0122] Positive feedback is reinforced by alert features , making the system more sensitive to similar dynamic behaviors. Learning rate Control the adjustment range to avoid overfitting, and repeatedly perform positive feedback correction steps by combining historical data. Operation steps: query the intelligent analysis results (such as the "fight" event alarm), and obtain the timestamp of the alarm video clip ; Extract the feature vector corresponding to the video transmission feature information of the video stream during network transmission from the storage module , update the feature matrix, add the deviation between the positive feedback result and the old feature evidence to the old matrix and multiply it by the preset positive learning rate to get the new feature matrix, which is expressed as follows: .

[0123] Negative feedback weakens the influence of misjudged features, reduces sensitivity to static or irrelevant segments, and optimizes discrimination ability. Less than To maintain stability. Operation steps: query the intelligent analysis results (such as "fighting" behavior alarm), obtain the alarm timestamp ; Extract the feature vector corresponding to the video transmission feature information of the video stream during network transmission from the storage module , adjust the adjustment matrix, subtract the deviation between the negative feedback result and the old feature evidence from the old matrix and multiply it by the preset negative feedback learning rate to obtain the new feature matrix, which is expressed as follows: .

[0124] Updating the preset feature matrix adapted to the video analysis task may include: adjusting the weights of the transmission feature parameters of each dimension in the preset feature matrix according to the feedback results of at least one video frame, and optimizing the strategy for subsequent video frame screening and analysis. Adjust the weights in the preset feature matrix according to the feedback results of at least one video frame. If it is necessary to increase the sensitivity to relevant action information, the weights of the feature parameters associated with such action information can be increased. For example, in a vehicle detection task, the weights of the feature parameters of the data packets related to the vehicle's motion trajectory and speed changes can be increased; if the sensitivity to irrelevant action information is to be reduced, the weights of its associated feature parameters can be reduced accordingly. In this way, the preset feature matrix is ​​made more in line with actual task requirements, and the subsequent video frame screening and analysis process is optimized.

[0125] This solution, through positive and negative feedback mechanisms, dynamically adjusts the weights indicated by the preset feature matrix based on the actual video content and analysis results. This makes screening decisions adaptive, removing the need for fixed screening and analysis modes. This allows for better handling of complex and changing video scenarios, improving the flexibility and accuracy of task processing. Furthermore, targeted adjustments to sensitivity to relevant and irrelevant motion information effectively filter out interference factors, highlight key information, and reduce misjudgments and missed detections. Furthermore, the preset feature matrix is ​​continuously updated based on feedback, enabling the system to maintain robust analysis results in diverse environments and with varying types of video data.

[0126] The technical solution of the embodiments of the present invention, in the video stream transmission link, contains independently encoded key frames and forward predictive coded frames. By extracting multi-dimensional video transmission feature information from the metadata and statistical characteristics of the video stream network data packets, it is possible to mine content dynamics from the transmission layer. This directly obtains video transmission feature information without decoding, efficiently captures the dynamic changes of the video stream during network transmission, avoids direct analysis of complex video content, and greatly reduces computational complexity. In terms of video frame determination, the dynamics of video content are indirectly characterized based on video transmission feature information, and video frames that are critical to subsequent processing tasks can be accurately screened. Compared with blindly processing all video frames, this avoids the waste of resources caused by indiscriminate analysis of all video clips, significantly reduces unnecessary processing, and improves video processing efficiency. When performing video analysis tasks, only the useful video frames that are screened need to be processed in a targeted manner, making the video processing results more in line with actual needs and improving video processing quality. At the same time, accurately processing useful video frames can also reduce system resource consumption, reduce the burden on computing equipment, and extend the service life of equipment, allowing more tasks to be analyzed with the same amount of computing resources.

[0127] Figure 4 A schematic structural diagram of a video processing device is provided for an embodiment of the present invention. This embodiment is applicable to situations where useful video frames in a video stream are screened and then video analysis tasks are performed. The video processing device can be implemented in the form of hardware and / or software and can be configured in any electronic device with network communication capabilities.

[0128] like Figure 4 As shown, the video processing device provided by the embodiment of the present invention may include the following:

[0129] Determination module 410, configured to determine video transmission characteristic information of a video stream during network transmission, wherein the plurality of video frames of the video stream include independently encoded key frames and forward predictive encoded frames, the video transmission characteristic information being obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream, and the video transmission characteristic information being used to characterize the content dynamics of video content corresponding to the video frames in the video stream;

[0130] a screening module 420, configured to determine at least one video frame from the video stream according to the video transmission characteristic information;

[0131] The processing module 430 is configured to perform a video analysis task on the at least one video frame.

[0132] Based on the above embodiment, optionally, determining video transmission feature information of a video stream during network transmission includes:

[0133] Determining video transmission parameter information of the video stream during network transmission, the video transmission parameter information including at least one of the following: the number of bytes of data packets when the video stream is transmitted over the network, real-time transport protocol metadata when the video stream is transmitted over the network, and the number of video frame fragments when the video stream is transmitted over the network;

[0134] Video transmission characteristic information of the video stream during network transmission is determined according to the video transmission parameter information.

[0135] Based on the above embodiment, optionally, the video transmission characteristic information is represented by at least one of data packet byte quantity disturbance information, real-time transport protocol metadata anomaly information, bit rate dynamic evolution information and fragment entropy state change information: wherein the data packet byte quantity disturbance information is used to indicate unexpected fluctuations in the data packet byte quantity per unit time when the video stream is transmitted over the network; the real-time transport protocol metadata anomaly information is used to indicate data packet loss per unit time when the video stream is transmitted over the network, and the data packet loss is manifested as discontinuous sequence numbers of data packets or timestamp jumps; the bit rate dynamic evolution information is used to indicate dynamic changes in the data transmission bit rate per unit time over time when the video stream is transmitted over the network; and the fragment entropy state change information is used to indicate fluctuations in the number of video frame fragments per unit time when the video stream is transmitted over the network.

[0136] Based on the above embodiment, optionally, when the video transmission parameter information includes the byte size of a data packet of a video stream during network transmission, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes:

[0137] Determining first parameter information according to the video transmission parameter information, where the first parameter information is used to indicate the packet byte amount of each data packet per unit time when the video stream is transmitted over the network;

[0138] Determining second parameter information according to the first parameter information, where the second parameter information is used to indicate an average number of bytes of data packets per unit time when the video stream is transmitted over the network;

[0139] Data packet byte quantity disturbance information of the video stream during network transmission is determined according to the first parameter information and the second parameter information.

[0140] Based on the above embodiment, optionally, before determining the second parameter information according to the first parameter information, the method further includes:

[0141] An exponential smoothing filter process is performed on the data packet byte amount of each data packet per unit time when the video stream indicated by the first parameter information is transmitted over the network.

[0142] Based on the above embodiment, optionally, when the video transmission parameter information includes real-time transport protocol metadata when the video stream is transmitted over the network, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes:

[0143] Determine third parameter information according to the video transmission parameter information, wherein the third parameter information is used to indicate a timestamp and a sequence number in a header of a plurality of consecutive data packets of the video stream extracted from the real-time transport protocol metadata during video stream transmission;

[0144] Determine fourth parameter information according to the third parameter information, where the fourth parameter information is used to indicate the number of timestamp jumps and the number of discontinuous sequence numbers per unit time during network transmission of the video stream;

[0145] Real-time transport protocol metadata exception information of the video stream during network transmission is determined according to the fourth parameter information.

[0146] Based on the above embodiment, optionally, determining the fourth parameter information according to the third parameter information includes:

[0147] determining, for the timestamps and sequence numbers in the Real-time Transport Protocol metadata headers of two adjacent data packets among the plurality of consecutive data packets indicated by the third parameter information, fifth parameter information, the fifth parameter information being used to indicate a timestamp increment in the Real-time Transport Protocol metadata headers of the two adjacent data packets and a sequence number difference in the Real-time Transport Protocol metadata headers of the two adjacent data packets;

[0148] The fourth parameter information is determined according to the fifth parameter information.

[0149] Based on the above embodiment, optionally, when the video transmission parameter information includes a video data transmission bit rate when the video stream is transmitted over the network, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes:

[0150] Determining first parameter information according to the video transmission parameter information, where the first parameter information is used to indicate the packet byte amount of each data packet per unit time when the video stream is transmitted over the network;

[0151] Determining sixth parameter information based on the first parameter information, the sixth parameter information is used to indicate the video data transmission bit rate when the video stream is transmitted over the network;

[0152] The seventh parameter information is determined based on the sixth parameter information to obtain the bit rate dynamic evolution information of the video stream during network transmission. The seventh parameter information is used to indicate the rate of change of the video data transmission bit rate per unit time during network transmission.

[0153] Based on the above embodiment, optionally, before determining the seventh parameter information according to the sixth parameter information, the method further includes:

[0154] Perform sliding average filtering on the video data transmission bit rate when the video stream indicated by the sixth parameter information is transmitted over the network.

[0155] Based on the above embodiment, optionally, when the video transmission parameter information includes the number of video frame fragments when the video stream is transmitted over the network, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes:

[0156] Determining eighth parameter information according to the video transmission parameter information, the eighth parameter information being used to indicate the number of video frame fragments transmitted at each time point per unit time when the video stream is transmitted over the network;

[0157] Determine ninth parameter information according to the eighth parameter information, where the ninth parameter information is used to indicate an average number of video frame fragments transmitted per unit time when the video stream is transmitted over the network;

[0158] The fragment entropy state transition information of the video stream during network transmission is determined according to the eighth parameter information and the ninth parameter information.

[0159] Based on the above embodiment, optionally, determining at least one video frame from the video stream according to the video transmission characteristic information includes:

[0160] Determine the task type corresponding to the video analysis task;

[0161] Determining a preset feature matrix adapted for the video analysis task according to a task type corresponding to the video analysis task, the preset feature matrix being used to indicate weights of transmission feature parameters of different dimensions in the video transmission feature information when screening video frames from the video stream, and the weights indicated by the preset feature matrix being used to reflect contributions of transmission feature parameters of different dimensions in the video transmission feature information when measuring content dynamics of video content corresponding to video frames in the video stream;

[0162] At least one video frame is determined from the video stream according to a preset feature matrix adapted for the video analysis task and the video transmission feature information.

[0163] Based on the above embodiment, optionally, determining at least one video frame from the video stream according to a preset feature matrix adapted for the video analysis task and the video transmission feature information includes:

[0164] According to the weights in the preset feature matrix adapted for the video analysis task, a decision result of the video frame in the video stream is obtained by weighting the transmission feature parameters of different dimensions in the video transmission feature information, and the decision result of the video frame in the video stream is used to indicate the possibility of the video frame in the video stream participating in the video analysis task;

[0165] At least one video frame is determined from the video stream according to a decision result of the video frame in the video stream.

[0166] Based on the above embodiment, optionally, after determining at least one video frame from the video stream according to the video transmission characteristic information, the method further includes:

[0167] Determining a feedback result of the at least one video frame, the feedback result of the at least one video frame being used to indicate whether to increase sensitivity to motion information related to a task type corresponding to the video analysis task in the video content included in the video frame, or whether to decrease sensitivity to motion information unrelated to the task type corresponding to the video analysis task in the video content included in the video frame;

[0168] In response to the feedback result of the at least one video frame, a preset feature matrix adapted for the video analysis task is updated.

[0169] The technical solution of the embodiments of the present invention, in the video stream transmission link, contains independently encoded key frames and forward predictive coded frames. By extracting multi-dimensional video transmission feature information from the metadata and statistical characteristics of the video stream network data packets, it is possible to mine content dynamics from the transmission layer. This directly obtains video transmission feature information without decoding, efficiently captures the dynamic changes of the video stream during network transmission, avoids direct analysis of complex video content, and greatly reduces computational complexity. In terms of video frame determination, the dynamics of video content are indirectly characterized based on video transmission feature information, and video frames that are critical to subsequent processing tasks can be accurately screened. Compared with blindly processing all video frames, this avoids the waste of resources caused by indiscriminate analysis of all video clips, significantly reduces unnecessary processing, and improves video processing efficiency. When performing video analysis tasks, only the useful video frames that are screened need to be processed in a targeted manner, making the video processing results more in line with actual needs and improving video processing quality. At the same time, accurately processing useful video frames can also reduce system resource consumption, reduce the burden on computing equipment, and extend the service life of equipment, allowing more tasks to be analyzed with the same amount of computing resources.

[0170] The video processing device provided in the embodiment of the present invention can execute the video processing method provided in any embodiment of the present invention, and has the corresponding functions and beneficial effects of executing the video processing method. For detailed process, please refer to the relevant operations of the video processing method in the above embodiment.

[0171] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of the present invention.

[0172] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0173] like Figure 5As shown, electronic device 10 includes at least one processor 11 and memory, such as read-only memory (ROM) 12 and random access memory (RAM) 13, communicatively connected to at least one processor 11. The memory stores computer programs executable by the at least one processor. Processor 11 can perform various appropriate actions and processes based on the computer programs stored in ROM 12 or loaded from storage unit 18 into RAM 13. RAM 13 can also store various programs and data required for the operation of electronic device 10. Processor 11, ROM 12, and RAM 13 are interconnected via bus 14. An input / output (I / O) interface 15 is also connected to bus 14.

[0174] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0175] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the video processing method.

[0176] In some embodiments, the video processing method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video processing method in any other suitable manner (e.g., via firmware).

[0177] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0178] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0179] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or apparatus. A computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0181] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0182] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0183] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0184] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A video processing method, characterized in that: The method comprises: Determining video transmission characteristic information of a video stream during network transmission, the video transmission characteristic information being obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream, the video transmission characteristic information being used to characterize content dynamics of video content corresponding to video frames in the video stream; Determine at least one video frame from the video stream according to the video transmission characteristic information; wherein, determining at least one video frame from the video stream according to the video transmission characteristic information includes: determining a task type corresponding to a video analysis task; determining a preset feature matrix adapted for the video analysis task according to the task type corresponding to the video analysis task, the preset feature matrix being used to indicate the weights of transmission feature parameters of different dimensions in the video transmission characteristic information when screening video frames from the video stream, the weights indicated by the preset feature matrix being used to reflect the contribution of transmission feature parameters of different dimensions in the video transmission characteristic information when measuring the content dynamics of video content corresponding to video frames in the video stream; determine at least one video frame from the video stream according to the preset feature matrix adapted for the video analysis task and the video transmission characteristic information; A video analysis task is performed on the at least one video frame.

2. The method according to claim 1, characterized in that Determine the video transmission characteristics of the video stream during network transmission, including: Determining video transmission parameter information of the video stream during network transmission, the video transmission parameter information including at least one of the following: the number of bytes of data packets when the video stream is transmitted over the network, real-time transport protocol metadata when the video stream is transmitted over the network, and the number of video frame fragments when the video stream is transmitted over the network; Video transmission characteristic information of the video stream during network transmission is determined according to the video transmission parameter information.

3. The method according to claim 1 or 2, characterized in that The video transmission characteristic information is represented by at least one of data packet byte quantity disturbance information, real-time transport protocol metadata anomaly information, bit rate dynamic evolution information and fragment entropy state change information: wherein the data packet byte quantity disturbance information is used to indicate the unexpected fluctuation change of the data packet byte quantity per unit time when the video stream is transmitted over the network; the real-time transport protocol metadata anomaly information is used to indicate the data packet loss per unit time when the video stream is transmitted over the network, and the data packet loss is manifested as discontinuous sequence number of the data packet or timestamp jump; the bit rate dynamic evolution information is used to indicate the dynamic change of the data transmission bit rate per unit time over time when the video stream is transmitted over the network; the fragment entropy state change information is used to indicate the fluctuation of the number of video frame fragments per unit time when the video stream is transmitted over the network.

4. The method according to claim 2, characterized in that In a case where the video transmission parameter information includes the amount of bytes of data packets when the video stream is transmitted over the network, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes: Determining first parameter information according to the video transmission parameter information, where the first parameter information is used to indicate the packet byte amount of each data packet per unit time when the video stream is transmitted over the network; Determining second parameter information according to the first parameter information, where the second parameter information is used to indicate an average number of bytes of data packets per unit time when the video stream is transmitted over the network; Data packet byte quantity disturbance information of the video stream during network transmission is determined according to the first parameter information and the second parameter information.

5. The method according to claim 2, characterized in that In a case where the video transmission parameter information includes real-time transport protocol metadata when the video stream is transmitted over a network, determining video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes: Determine third parameter information according to the video transmission parameter information, wherein the third parameter information is used to indicate a timestamp and a sequence number in a header of a plurality of consecutive data packets of the video stream extracted from the real-time transport protocol metadata during video stream transmission; Determine fourth parameter information according to the third parameter information, where the fourth parameter information is used to indicate the number of timestamp jumps and the number of discontinuous sequence numbers per unit time during network transmission of the video stream; Real-time transport protocol metadata exception information of the video stream during network transmission is determined according to the fourth parameter information.

6. The method according to claim 5, characterized in that Determining fourth parameter information according to the third parameter information includes: determining, for the timestamps and sequence numbers in the Real-time Transport Protocol metadata headers of two adjacent data packets among the plurality of consecutive data packets indicated by the third parameter information, fifth parameter information, the fifth parameter information being used to indicate a timestamp increment in the Real-time Transport Protocol metadata headers of the two adjacent data packets and a sequence number difference in the Real-time Transport Protocol metadata headers of the two adjacent data packets; The fourth parameter information is determined according to the fifth parameter information.

7. The method according to claim 2, characterized in that In a case where the video transmission parameter information includes a video data transmission bit rate when the video stream is transmitted over a network, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes: Determining first parameter information according to the video transmission parameter information, where the first parameter information is used to indicate the packet byte amount of each data packet per unit time when the video stream is transmitted over the network; Determining sixth parameter information based on the first parameter information, the sixth parameter information is used to indicate the video data transmission bit rate when the video stream is transmitted over the network; The seventh parameter information is determined based on the sixth parameter information to obtain the bit rate dynamic evolution information of the video stream during network transmission. The seventh parameter information is used to indicate the rate of change of the video data transmission bit rate per unit time during network transmission.

8. The method according to claim 2, characterized in that In a case where the video transmission parameter information includes the number of video frame fragments when the video stream is transmitted over a network, determining the video transmission characteristic information of the video stream during network transmission according to the video transmission parameter information includes: Determining eighth parameter information according to the video transmission parameter information, the eighth parameter information being used to indicate the number of video frame fragments transmitted at each time point per unit time when the video stream is transmitted over the network; Determine ninth parameter information according to the eighth parameter information, where the ninth parameter information is used to indicate an average number of video frame fragments transmitted per unit time when the video stream is transmitted over the network; The fragment entropy state transition information of the video stream during network transmission is determined according to the eighth parameter information and the ninth parameter information.

9. The method according to claim 1, characterized in that After determining at least one video frame from the video stream according to the video transmission characteristic information, the method further includes: Determining a feedback result of the at least one video frame, the feedback result of the at least one video frame being used to indicate whether to increase sensitivity to motion information related to a task type corresponding to the video analysis task in the video content included in the video frame, or whether to decrease sensitivity to motion information unrelated to the task type corresponding to the video analysis task in the video content included in the video frame; In response to the feedback result of the at least one video frame, a preset feature matrix adapted for the video analysis task is updated.

10. A video processing device, characterized in that: The device comprises: A determination module is used to determine video transmission characteristic information of a video stream during network transmission, wherein the video transmission characteristic information is obtained based on metadata and statistical characteristic data of network data packets corresponding to the video stream, and the video transmission characteristic information is used to characterize the content dynamics of the video content corresponding to the video frame in the video stream; wherein, determining at least one video frame from the video stream based on the video transmission characteristic information comprises: determining a task type corresponding to a video analysis task; determining a preset characteristic matrix adapted for the video analysis task based on the task type corresponding to the video analysis task, wherein the preset characteristic matrix is ​​used to indicate the weights of transmission characteristic parameters of different dimensions in the video transmission characteristic information when screening video frames from the video stream, and the weights indicated by the preset characteristic matrix are used to reflect the contribution of transmission characteristic parameters of different dimensions in the video transmission characteristic information when measuring the content dynamics of the video content corresponding to the video frame in the video stream; determining at least one video frame from the video stream based on the preset characteristic matrix adapted for the video analysis task and the video transmission characteristic information; a screening module, configured to determine at least one video frame from the video stream according to the video transmission characteristic information; A processing module is configured to perform a video analysis task on the at least one video frame.

11. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the video processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • High-speed recognizing method of change degree of video content

    CN103200419A

  • Video data transmission method

    CN119255057A