Video coding method, video decoding method, video coding device, video decoding device and storage medium

By calculating the inter-frame optical flow variation index of video frames to divide them into key frames and redundant frames, and performing semantic text recognition on the redundant frames to generate bitstreams, the problem that existing video coding methods are difficult to adapt to different tasks is solved, and communication costs are reduced and analysis accuracy is improved.

CN120915950AActive Publication Date: 2025-11-07PEKING UNIV

Patent Information

Application Number
CN202510967450.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-07
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing video coding methods rely on a uniform and fixed model structure, which makes it difficult to flexibly adapt to different video temporal tasks. This leads to redundant video frame transmission, which increases communication costs and reduces the efficiency of subsequent analysis.

Method used

By calculating the inter-frame optical flow variation index of video frames, the video frame sequence is divided into key frames and redundant frames. Image encoding is performed on key frames, and semantic text recognition is performed on redundant frames to generate semantic text code streams for transmission, replacing the transmission of redundant frame images.

Benefits of technology

It effectively reduces communication overhead, extracts semantic information from different dimensions, flexibly adapts to different downstream tasks, and improves the accuracy of subsequent analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915950A_ABST
    Figure CN120915950A_ABST
Patent Text Reader

Abstract

The invention provides a video coding method and device, a video decoding method and device and a storage medium, and the video coding method comprises the steps: obtaining a video frame sequence; for any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents the motion change degree between the video frame and the previous video frame; dividing the video frame sequence into a key frame set and a redundant frame set; performing image coding on each key frame in the key frame set to obtain an image code stream; performing semantic text recognition on each redundant frame in the redundant frame set through various recognition models, and generating a semantic text code stream according to a plurality of semantic text recognition results corresponding to the plurality of redundant frames one by one; and sending the image code stream and the semantic text code stream to a cloud server. According to the embodiment of the invention, different downstream tasks can be flexibly adapted by extracting information of different dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to a video encoding method, a video decoding method, a device and a storage medium. BACKGROUND

[0002] In an end-edge-cloud integrated intelligent sensing system, an end-side device continuously generates a large amount of video data, such as security monitoring video and real-time video stream of a traffic scene. These video data are usually collected and transmitted at high resolution and high frame rate, which brings great pressure on transmission and calculation of the edge side and the cloud side. Traditional video encoding methods, such as H.264 and H.265, mainly focus on pixel-level compression within each frame or segment, but lack effective mining of video semantic information and timing information. Therefore, these methods often result in transmission of a large number of redundant video frames, which not only increases communication cost but also reduces the efficiency of subsequent analysis.

[0003] To solve this problem, the prior art attempts to introduce a feature encoding method. By deploying a visual model on the edge side to extract video features and transmit them, the cost of video transmission is reduced. However, this feature encoding method has obvious limitations: it relies on a uniform and fixed model structure and feature representation, and is difficult to flexibly adapt to different video timing tasks, such as video event detection, behavior recognition and anomaly detection. SUMMARY

[0004] Therefore, the present application proposes a video encoding method, a video decoding method, a device and a storage medium to solve the problem that the existing feature encoding method relies on a uniform and fixed model structure and feature representation and is difficult to flexibly adapt to different video timing tasks.

[0005] The first aspect embodiment of the present application proposes a video encoding method, comprising:

[0006] obtaining a video frame sequence;

[0007] For any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents the degree of motion change between the video frame and a previous video frame;

[0008] dividing the video frame sequence into a key frame set and a redundant frame set; the inter-frame optical flow change index of a key frame in the key frame set is greater than or equal to an optical flow change threshold, and the inter-frame optical flow change index of a redundant frame in the redundant frame set is less than the optical flow change threshold;

[0009] respectively performing image encoding on each key frame in the key frame set to obtain an image code stream;

[0010] perform semantic text recognition on each of the redundant frame set by multiple recognition models, and generate a semantic text code stream from the one-to-one corresponding multiple semantic text recognition results of the multiple redundant frames;

[0011] send the image code stream and the semantic text code stream to a cloud server.

[0012] The embodiment of the application replaces the transmission of redundant frame images with a semantic text code stream generated from multiple semantic text recognition results, which can effectively reduce communication overhead. By performing semantic text recognition on each of the redundant frame set by multiple recognition models, different dimensions of semantic information can be extracted, different downstream tasks can be flexibly adapted, and the accuracy of subsequent analysis tasks can also be improved.

[0013] In the embodiment of the application, the inter-frame optical flow change indicator of the video frame is calculated, including:

[0014] For any pixel point in the multiple pixel points in the video frame, the optical flow vector of the pixel point between the video frame and the previous video frame is calculated; the optical flow vector represents the motion direction and distance of the pixel point between two adjacent video frames;

[0015] The Euclidean norm of the optical flow vector is calculated; the Euclidean norm represents the motion amplitude of the pixel point between two adjacent video frames;

[0016] The multiple Euclidean norms corresponding to the multiple pixel points are summed to obtain a summation result;

[0017] According to the summation result and the total area of the image of the video frame, the inter-frame optical flow change indicator of the video frame is calculated.

[0018] In the embodiment of the application, the video frame sequence is divided into a key frame set and a redundant frame set, including:

[0019] For any video frame in the video frame sequence, the size relationship between the inter-frame optical flow change indicator of the video frame and the optical flow change threshold is determined;

[0020] If the inter-frame optical flow change indicator of the video frame is greater than or equal to the optical flow change threshold, the video frame is determined as a key frame;

[0021] If the inter-frame optical flow change indicator of the video frame is less than the optical flow change threshold, the video frame is determined as a redundant frame;

[0022] The key frame set is established according to multiple key frames, and the redundant frame set is established according to multiple redundant frames.

[0023] In the embodiments of the present application, the plurality of recognition models include a scene recognition model, an object detection model, and a text information extraction model; semantic text recognition is performed on each of the redundant frame set by using the plurality of recognition models, including:

[0024] For any redundant frame, scene label information of the redundant frame is extracted according to the scene recognition model;

[0025] Object information of the redundant frame is extracted according to the object detection model; the object information includes object attribute information and object position information;

[0026] Text information of the redundant frame is extracted according to the text information extraction model;

[0027] Timestamp information of the redundant frame for semantic text recognition is determined;

[0028] The scene label information, the object information, the text information, and the timestamp information are combined to obtain a semantic text recognition result of the redundant frame.

[0029] In the embodiments of the present application, the inter-frame optical flow change index of the video frame is calculated, including:

[0030]

[0031] Wherein, represents the inter-frame optical flow change index of the i th video frame; HW represents the total area of the i th video frame, H represents the height of the image, and W represents the width of the image; F i (x, y) represents the optical flow vector of the pixel point (x, y) between the i th video frame and the (i-1) th video frame; ‖F i (x, y)‖2 represents the Euclidean norm of the optical flow vector.

[0032] The second aspect of the present application provides a video decoding method, including:

[0033] An image code stream and a semantic text code stream are received; the image code stream is obtained by respectively performing image encoding on each key frame in a key frame set, the key frame set is composed of video frames in a video frame sequence whose inter-frame optical flow change index is greater than or equal to an optical flow change threshold value, the inter-frame optical flow change index represents the degree of motion change between the video frame and the previous video frame; the semantic text code stream is generated according to a plurality of semantic text recognition results corresponding to a plurality of redundant frames, each semantic text recognition result is obtained by performing semantic text recognition on a corresponding redundant frame in a redundant frame set by using a plurality of recognition models; the redundant frame set is composed of video frames in the video frame sequence whose inter-frame optical flow change index is less than the optical flow change threshold value;

[0034] decoding the image code stream to obtain a plurality of key frames, and decoding the semantic text code stream to obtain a plurality of semantic text recognition results;

[0035] combining the plurality of key frames and the plurality of semantic text recognition results to obtain reconstructed video sequence information.

[0036] In the embodiments of the present application, after obtaining the reconstructed video sequence information, the method further comprises:

[0037] inputting the reconstructed video sequence information into the trained abnormal event recognition model to output an abnormal event recognition result.

[0038] An embodiment of the third aspect of the present application provides a video encoding device, comprising:

[0039] a sequence acquisition module configured to acquire a video frame sequence;

[0040] an index calculation module configured to calculate, for any video frame in the video frame sequence, an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a degree of motion change between the video frame and a previous video frame;

[0041] a sequence division module configured to divide the video frame sequence into a key frame set and a redundant frame set; a key frame in the key frame set has an inter-frame optical flow change index greater than or equal to an optical flow change threshold, and a redundant frame in the redundant frame set has an inter-frame optical flow change index less than the optical flow change threshold;

[0042] an image code stream generation module configured to respectively perform image encoding on each key frame in the key frame set to obtain an image code stream;

[0043] a semantic text code stream generation module configured to respectively perform semantic text recognition on each redundant frame in the redundant frame set by using a plurality of recognition models, and generate a semantic text code stream from a plurality of semantic text recognition results corresponding to a plurality of redundant frames;

[0044] a code stream sending module configured to send the image code stream and the semantic text code stream to a cloud server.

[0045] An embodiment of the fourth aspect of the present application provides an electronic device, which comprises a memory and a processor, the memory and the processor are connected with each other in communication, the memory stores computer instructions, and the processor executes the computer instructions to perform the method of the first aspect or the second aspect.

[0046] The embodiment of the fifth aspect of the present application provides a computer readable storage medium, and the computer readable storage medium stores computer instructions, and the computer instructions are used to make a computer execute the method in the first aspect or the second aspect.

[0047] Additional aspects and advantages of the present application will be made apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0048] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description, with reference to the drawings in which:

[0049] In the drawings:

[0050] Figure 1 A flow diagram of a video encoding method according to an embodiment of the present application is shown;

[0051] Figure 2 A structural diagram of a video encoding apparatus according to an embodiment of the present application is shown;

[0052] Figure 3 A structural diagram of an electronic device according to an embodiment of the present application is shown;

[0053] Figure 4 A schematic diagram of a storage medium according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0054] Exemplary embodiments of the present application will be described herein below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thoroughly and completely understood, and will fully convey the scope of the application to those skilled in the art.

[0055] It should be noted that unless otherwise specified, technical or scientific terms used in the present application should be understood as their common meaning to those skilled in the art to which the present application pertains.

[0056] According to an embodiment of the present application, a video encoding method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0057] This embodiment provides a video encoding method, which is applied to a client. Figure 1 This is a flowchart of a video encoding method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:

[0058] Step S101: Obtain the video frame sequence.

[0059] Specifically, video data of real-time traffic scenes can be collected using camera sensors and parsed into a video frame sequence V = {I1, I2, ..., I...} i}; where each video frame I i This represents the image acquired at time point i.

[0060] Step S102: For any video frame in the video frame sequence, calculate the inter-frame optical flow variation index of the video frame.

[0061] Specifically, the inter-frame optical flow variation index represents the degree of motion change between the video frame and the previous video frame.

[0062] In some specific embodiments, step S102 above includes steps S1021-S1024:

[0063] Step S1021: For any pixel among multiple pixels in the video frame, calculate the optical flow vector of the pixel between the video frame and the previous video frame.

[0064] Specifically, the optical flow vector F i (x,y) represents the direction and distance of movement of the pixel (x,y) between two adjacent video frames (i.e., video frame i and the previous video frame i-1).

[0065] Step S1022: Calculate the Euclidean norm of the optical flow vector.

[0066] Specifically, for each optical flow vector, its Euclidean norm is calculated, i.e., ||F||. i (x,y)‖2; The Euclidean norm (also known as the L2 norm) represents the motion amplitude of the pixel between two adjacent video frames.

[0067] Step S1023: Summing the multiple Euclidean norms corresponding one-to-one with the multiple pixels to obtain a summation result.

[0068] Specifically, for all pixel positions in the i-th frame, the L2 norm of the optical flow vectors of all pixels is summed. That is, for all pixels (x, y) within the image width W and height H, the following calculation is performed:

[0069]

[0070] Step S1024, according to the summation result and the total image area of the video frame, the inter-frame optical flow change index of the video frame is calculated.

[0071] Specifically, the summation result can be divided by the total image area of the video frame i to obtain the inter-frame optical flow change index of the video frame i.

[0072] In the above steps S1021-S1024, the inter-frame optical flow change index of the video frame i can be calculated by the following formula:

[0073]

[0074] wherein, represents the inter-frame optical flow change index of the i-th video frame; HW represents the total image area of the i-th video frame, H represents the height of the image, and W represents the width of the image; F i (x, y) represents the optical flow vector of the pixel point (x, y) between the i-th video frame and the (i-1)-th video frame; ‖F i (x, y)‖2 represents the Euclidean norm of the optical flow vector.

[0075] Step S103, the video frame sequence is divided into a key frame set and a redundant frame set.

[0076] Specifically, the inter-frame optical flow change index of the key frame in the key frame set is greater than or equal to the optical flow change threshold, and the inter-frame optical flow change index of the redundant frame in the redundant frame set is less than the optical flow change threshold. Wherein, the optical flow change threshold can be set according to the actual situation, which is not limited here.

[0077] In some embodiments, the above step S103 includes steps S1031-S1034:

[0078] Step S1031, for any video frame in the video frame sequence, the size relationship between the inter-frame optical flow change index of the video frame and the optical flow change threshold is judged.

[0079] Step S1032, if the inter-frame optical flow change index of the video frame is greater than or equal to the optical flow change threshold, the video frame is determined as a key frame.

[0080] Step S1033, if the inter-frame optical flow change index of the video frame is less than the optical flow change threshold, the video frame is determined as a redundant frame.

[0081] Step S1034, establishing the key frame set according to the plurality of key frames, and establishing the redundant frame set according to the plurality of redundant frames.

[0082] In the above steps S1031-S1034, the key frame set K and the redundant frame set R can be determined by the following way:

[0083]

[0084] Wherein, I i represents the i-th video frame in the video frame sequence, represents the inter-frame optical flow change indicator of the i-th video frame, τ flow represents the optical flow change threshold;

[0085] That is, when the inter-frame optical flow change indicator of any video frame i is greater than or equal to the optical flow change threshold τ flow , the i-th video frame is determined as a key frame;

[0086] Conversely, when the inter-frame optical flow change indicator of any video frame i is less than the optical flow change threshold τ flow , the i-th video frame is determined as a redundant frame, that is, R=V\K; the video frames in the video frame sequence V except the key frames K are all redundant frames.

[0087] Step S104, respectively performing image encoding on each key frame in the key frame set to obtain an image code stream.

[0088] Specifically, each frame in the key frame set can be high-quality image encoded by a video encoding standard (such as H.264 / H.265) to generate an image code stream B img :

[0089]

[0090] Step S105, respectively performing semantic text recognition on each redundant frame in the redundant frame set by a plurality of recognition models, and generating a semantic text code stream from the plurality of semantic text recognition results corresponding to the plurality of redundant frames.

[0091] Specifically, the plurality of recognition models include a scene recognition model, an object detection model, and a text information extraction model; wherein the scene recognition model includes but is not limited to the RAM++ model, the object detection model includes but is not limited to the EVA02 model, the OWL-ViTv2 model, and the text information extraction model includes but is not limited to the PaddleOCR model and the LLaVA-NEXT model.

[0092] In some embodiments, the step S105 includes steps S1051-S1055:

[0093] In step S1051, scene label information of the redundant frame is extracted according to the scene recognition model.

[0094] In step S1052, object information of the redundant frame is extracted according to the object detection model.

[0095] The object information includes object attribute information and object position information.

[0096] In step S1053, text information of the redundant frame is extracted according to the text information extraction model.

[0097] Specifically, the specific implementation of the scene recognition model, the object detection model and the text information extraction model is not limited, and the corresponding functions of the corresponding models can be realized, for example:

[0098] Suppose there is a video frame in a traffic scene, the scene label of the video frame can be identified as urban street through the scene recognition model; a red car exists in one image position of the video frame and a blue bicycle exists in another image position of the video frame through the object detection model; the text information "no parking" in the video frame can be identified through the text information extraction model.

[0099] In step S1054, the timestamp information of the redundant frame for semantic text recognition is determined.

[0100] In step S1055, the scene label information, the object information, the text information and the timestamp information are merged to obtain the semantic text recognition result of the redundant frame.

[0101] Specifically, the obtained semantic text recognition result is as follows:

[0102] T feat (I i )=Concat(T env ,T pos ,T text ,Timestamp)

[0103] Wherein, T feat (I i ) represents the semantic text recognition result based on the i-th redundant frame, T env represents the scene label information, T pos represents the object information, T text represents the text information, and Timestam represents the timestamp information.

[0104] In step S105, the semantic text recognition result corresponding to each redundant frame is encoded by the BPE tokenizer to obtain an encoded result; and the multiple encoded results corresponding to the multiple redundant frames are merged to obtain a semantic text code stream.

[0105] As shown in the following formula:

[0106] B Token (I i )=BPE(T feat (I i ))

[0107] Wherein, B Token (I i ) represents the encoded result of the i-th redundant frame, BPE represents the BPE tokenizer, T feat (I i ) represents the semantic text recognition result corresponding to the i-th redundant frame.

[0108] Step S106, the image code stream and the semantic text code stream are sent to a cloud server.

[0109] Embodiment 2:

[0110] Corresponding to the implementation of the above video encoding method, the embodiment of the application also provides a video decoding method applied to a cloud server. The video decoding method comprises:

[0111] Step S201, receiving an image code stream and a semantic text code stream.

[0112] Specifically, the image code stream is obtained by respectively performing image encoding on each key frame in a key frame set, the key frame set is composed of video frames in a video frame sequence whose inter-frame optical flow change index is greater than or equal to an optical flow change threshold, and the inter-frame optical flow change index represents the degree of motion change between the video frame and the previous video frame; the semantic text code stream is generated according to multiple semantic text recognition results corresponding to multiple redundant frames, each semantic text recognition result is obtained by respectively performing semantic text recognition on a corresponding redundant frame in a redundant frame set by multiple recognition models; and the redundant frame set is composed of video frames in a video frame sequence whose inter-frame optical flow change index is less than the optical flow change threshold.

[0113] Step S202, decoding the image code stream to obtain multiple key frames, and decoding the semantic text code stream to obtain multiple semantic text recognition results.

[0114] Step S203, combining the multiple key frames and the multiple semantic text recognition results to obtain reconstructed video sequence information.

[0115] Specifically, the reconstructed video sequence information can be obtained according to the plurality of key frames, the timestamp information corresponding to each key frame, the plurality of semantic text recognition results, and the timestamp information corresponding to each semantic text recognition result.

[0116] In some embodiments, after obtaining the reconstructed video sequence information, the method further comprises:

[0117] inputting the reconstructed video sequence information into the trained abnormal event recognition model, and outputting an abnormal event recognition result.

[0118] In the embodiments of the present application, the abnormal event recognition model includes but is not limited to a behavior recognition model, a traffic event detection model, and a long-time question and answer model, so as to complete a video-level multi-modal reasoning task, and the cloud server can optimize the key frame determination strategy and the expert model configuration on the edge side based on the time sequence semantic feedback.

[0119] Embodiment 3:

[0120] Corresponding to the implementation mode of the above video encoding method, the embodiments of the present application also provide an intelligent video analysis device for real-time traffic event detection. The device comprises:

[0121] The edge-side device (such as an intelligent camera) collects real-time traffic scene video data to obtain an original video sequence.

[0122] The edge-side device (such as an edge computing node) calculates the inter-frame optical flow change according to the formula, determines a key frame set and a redundant frame set, performs efficient video compression encoding (such as H.265) on the key frame set to generate a key frame code stream, deploys scene understanding (RAM++), object detection (EVA02), and text recognition (PaddleOCR) models to extract semantic information from the redundant frame set, adds a timestamp field, uses a BPE algorithm to construct a structured semantic text code stream, and combines the image code stream and the structured semantic text code stream to send to the cloud.

[0123] The cloud-side device (such as a cloud analysis server) receives and decodes the key frame image and the structured semantic text, reconstructs the video time sequence information, uses a multi-modal video understanding model (such as LLaVA) to detect traffic abnormal events, and outputs a detection result.

[0124] Compared with the prior art, the intelligent video analysis device for real-time traffic event detection provided by the application, under an end-edge-cloud collaborative architecture, fuses frame-level structured semantic compression and timestamp alignment mechanism, dynamically extracts high-value semantic information of redundant frames in the video through edge-side deployment of multiple types of perception expert models, and uniformly organizes the high-value semantic information into a structured text code stream with timing information, thereby replacing the transmission demand of most of the redundant frame images. This method not only greatly compresses the overall video code rate, but also effectively preserves the semantic integrity and timing structure of the video content, and has good real-time performance and scalability. It is suitable for the semantic understanding needs of multi-modal large models in video question and answer, event analysis, behavior recognition and other timing perception tasks, while ensuring the understanding accuracy, significantly reducing the transmission pressure between the edge and the cloud.

[0125] Embodiment 4:

[0126] Corresponding to the implementation manner of the above video encoding method, the embodiment of the application further provides a video encoding device for executing the video encoding method described in the above embodiments. As shown in the figure, the video encoding device comprises: Figure 2

[0127] a sequence acquisition module, configured to acquire a video frame sequence;

[0128] an index calculation module, configured to calculate, for any video frame in the video frame sequence, an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a degree of motion change between the video frame and a previous video frame;

[0129] a sequence division module, configured to divide the video frame sequence into a key frame set and a redundant frame set; a key frame in the key frame set has an inter-frame optical flow change index greater than or equal to an optical flow change threshold, and a redundant frame in the redundant frame set has an inter-frame optical flow change index less than the optical flow change threshold;

[0130] an image code stream generation module, configured to respectively perform image encoding on each key frame in the key frame set to obtain an image code stream;

[0131] a semantic text code stream generation module, configured to respectively perform semantic text recognition on each redundant frame in the redundant frame set through multiple recognition models, and generate a semantic text code stream from multiple semantic text recognition results corresponding to the multiple redundant frames one by one;

[0132] a code stream sending module, configured to send the image code stream and the semantic text code stream to a cloud server.

[0133] ​Optionally, the index calculation module is further configured to: calculate, for any pixel point in the plurality of pixel points in the video frame, an optical flow vector of the pixel point between the video frame and a previous video frame; the optical flow vector represents a motion direction and distance of the pixel point between two adjacent video frames; calculate a Euclidean norm of the optical flow vector; the Euclidean norm represents a motion amplitude of the pixel point between two adjacent video frames; sum a plurality of Euclidean norms corresponding to the plurality of pixel points to obtain a summation result; and calculate the inter-frame optical flow change index of the video frame according to the summation result and a total image area of the video frame.

[0134] Optionally, the sequence division module is further configured to: determine, for any video frame in the video frame sequence, a size relationship between the inter-frame optical flow change index of the video frame and an optical flow change threshold; if the inter-frame optical flow change index of the video frame is greater than or equal to the optical flow change threshold, determine the video frame as a key frame; if the inter-frame optical flow change index of the video frame is less than the optical flow change threshold, determine the video frame as a redundant frame; establish the key frame set according to a plurality of key frames, and establish the redundant frame set according to a plurality of redundant frames.

[0135] Optionally, the plurality of recognition models include a scene recognition model, an object detection model, and a text information extraction model; and the semantic text code stream generation module is further configured to: extract, for any redundant frame, scene tag information of the redundant frame according to the scene recognition model; extract object information of the redundant frame according to the object detection model; the object information includes object attribute information and object position information; extract text information of the redundant frame according to the text information extraction model; determine timestamp information of the redundant frame for semantic text recognition; and combine the scene tag information, the object information, the text information, and the timestamp information to obtain a semantic text recognition result of the redundant frame.

[0136] The video encoding apparatus provided by the above embodiments of the present application and the video encoding method provided by the embodiments of the present application have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0137] The embodiments of the present application further provide an electronic device for executing the above video encoding method or video decoding method. Please refer to Figure 3 which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As shown in Figure 3As shown, the electronic device 3 comprises a processor 300, a memory 301, a bus 302 and a communication interface 303, the processor 300, the communication interface 303 and the memory 301 are connected through the bus 302; the memory 301 stores a computer program capable of running on the processor 300, and the processor 300 runs the computer program to execute the video encoding method or the video decoding method provided in the foregoing embodiments of the present application.

[0138] The memory 301 can include a high-speed random access memory (RAM) and can also include a non-volatile memory such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 303 (which can be wired or wireless), and the Internet, a wide area network, a local network, a metropolitan area network, etc. can be used.

[0139] The bus 302 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory 301 is used to store programs, and the processor 300 executes the programs after receiving execution instructions. The video encoding method or the video decoding method disclosed in the foregoing embodiments can be applied to the processor 300 or implemented by the processor 300.

[0140] The processor 300 can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 300 or the instruction in the form of software. The processor 300 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory 301, and the processor 300 reads the information in the memory 301 and combines the hardware to complete the steps of the above method.

[0141] The electronic device provided by the embodiments of the present application has the same beneficial effects as the method it adopts, runs or implements.

[0142] The embodiments of the present application also provide a computer readable storage medium corresponding to the video encoding method or video decoding method provided by the preceding embodiments. Please refer to Figure 4 The computer readable storage medium shown is an optical disc 30, and a computer program (i.e. program product) is stored on the optical disc 30. When the computer program is run by a processor, the video encoding method or video decoding method provided by any of the preceding embodiments is executed.

[0143] It should be noted that examples of the computer readable storage medium can also include, but are not limited to, a phase change memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), other types of random access memory (RAM), a read only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a flash memory or other optical, magnetic storage medium, which will not be described one by one here.

[0144] The computer readable storage medium provided by the above embodiments of the present application has the same beneficial effects as the video encoding method or the video decoding method provided by the embodiments of the present application, and has the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0145] It should be noted that:

[0146] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail in order not to obscure the understanding of this description.

[0147] Similarly, it is to be understood that the above description is illustrative of the various aspects of the present application and that changes can be made to the embodiments described while still obtaining a desired outcome. As such, the scope of the present application is not to be unduly limited to such specific embodiments merely because of the language used herein, but is to be controlled solely by the appended claims, and their equivalents.

[0148] Further, those skilled in the art will appreciate that the features of the various embodiments described herein are not mutually exclusive, but can be combined in different embodiments. For example, in the following claims, any of the embodiments claimed can be used in any combination.

[0149] The above descriptions are only the preferred embodiments of the present application, but the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of video coding, the method comprising: The method comprises: obtaining a video frame sequence; for any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a degree of motion change between the video frame and a previous video frame; dividing the video frame sequence into a key frame set and a redundant frame set; the inter-frame optical flow change index of a key frame in the key frame set is greater than or equal to an optical flow change threshold, and the inter-frame optical flow change index of a redundant frame in the redundant frame set is less than the optical flow change threshold; respectively performing image encoding on each key frame in the key frame set to obtain an image code stream; respectively performing semantic text recognition on each redundant frame in the redundant frame set by using multiple recognition models, and generating a semantic text code stream from multiple semantic text recognition results corresponding to the multiple redundant frames; sending the image code stream and the semantic text code stream to a cloud server.

2. The method of claim 1, wherein, The method comprises: obtaining a video frame sequence; for any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a degree of motion change between the video frame and a previous video frame; dividing the video frame sequence into a key frame set and a redundant frame set; the inter-frame optical flow change index of a key frame in the key frame set is greater than or equal to an optical flow change threshold, and the inter-frame optical flow change index of a redundant frame in the redundant frame set is less than the optical flow change threshold; respectively performing image encoding on each key frame in the key frame set to obtain an image code stream; 3. The method according to claim 1 or 2, characterized in that, respectively performing semantic text recognition on each redundant frame in the redundant frame set by using multiple recognition models, and generating a semantic text code stream from multiple semantic text recognition results corresponding to the multiple redundant frames; sending the image code stream and the semantic text code stream to a cloud server. The method comprises: obtaining a video frame sequence; for any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a degree of motion change between the video frame and a previous video frame; 4. The method of claim 1 or 2, wherein the plurality of recognition models comprises a scene recognition model, an object detection model, and a text information extraction model. dividing the video frame sequence into a key frame set and a redundant frame set; the inter-frame optical flow change index of a key frame in the key frame set is greater than or equal to an optical flow change threshold, and the inter-frame optical flow change index of a redundant frame in the redundant frame set is less than the optical flow change threshold; respectively performing image encoding on each key frame in the key frame set to obtain an image code stream; respectively performing semantic text recognition on each redundant frame in the redundant frame set by using multiple recognition models, and generating a semantic text code stream from multiple semantic text recognition results corresponding to the multiple redundant frames; sending the image code stream and the semantic text code stream to a cloud server. The method comprises: obtaining a video frame sequence; 5. The method according to claim 1 or 2, characterized in that, for any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a degree of motion change between the video frame and a previous video frame; wherein, represents the inter-frame optical flow variation indicator of the i-th video frame; HW represents the total area of the image of the i-th video frame, H represents the height of the image, and W represents the width of the image; F i (x,y) represents the optical flow vector of the pixel point (x,y) between the i-th video frame and the i-1-th video frame; ‖F i (x,y)‖2 represents the Euclidean norm of the optical flow vector.

6. A method of video decoding, comprising: dividing the video frame sequence into a key frame set and a redundant frame set; the inter-frame optical flow change index of a key frame in the key frame set is greater than or equal to an optical flow change threshold, and the inter-frame optical flow change index of a redundant frame in the redundant frame set is less than the optical flow change threshold; respectively performing image encoding on each key frame in the key frame set to obtain an image code stream; respectively performing semantic text recognition on each redundant frame in the redundant frame set by using multiple recognition models, and generating a semantic text code stream from multiple semantic text recognition results corresponding to the multiple redundant frames; sending the image code stream and the semantic text code stream to a cloud server. receiving an image code stream and a semantic text code stream; the image code stream is obtained by respectively performing image encoding on each key frame in a key frame set, the key frame set is composed of video frames in a video frame sequence, and an inter-frame optical flow change index of each video frame in the key frame set is greater than or equal to an optical flow change threshold; the inter-frame optical flow change index represents a motion change degree between the video frame and a previous video frame; the semantic text code stream is generated according to a plurality of semantic text recognition results corresponding to a plurality of redundant frames; each semantic text recognition result is obtained by respectively performing semantic text recognition on a corresponding redundant frame in a redundant frame set through a plurality of recognition models; the redundant frame set is composed of video frames in the video frame sequence, and an inter-frame optical flow change index of each video frame in the redundant frame set is less than the optical flow change threshold; decoding the image code stream to obtain a plurality of key frames, and decoding the semantic text code stream to obtain a plurality of semantic text recognition results; combining the plurality of key frames and the plurality of semantic text recognition results to obtain reconstructed video sequence information.

7. The method of claim 6, wherein, After obtaining the reconstructed video sequence information, the method further includes: inputting the reconstructed video sequence information into a trained abnormal event recognition model to output an abnormal event recognition result.

8. A video encoding apparatus, comprising: The apparatus includes: a sequence acquisition module configured to acquire a video frame sequence; an index calculation module configured to calculate, for any video frame in the video frame sequence, an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents a motion change degree between the video frame and a previous video frame; a sequence division module configured to divide the video frame sequence into a key frame set and a redundant frame set; a key frame in the key frame set has an inter-frame optical flow change index greater than or equal to an optical flow change threshold, and a redundant frame in the redundant frame set has an inter-frame optical flow change index less than the optical flow change threshold; an image code stream generation module configured to respectively perform image encoding on each key frame in the key frame set to obtain an image code stream; a semantic text code stream generation module configured to respectively perform semantic text recognition on each redundant frame in the redundant frame set through a plurality of recognition models, and generate a semantic text code stream according to a plurality of semantic text recognition results corresponding to a plurality of redundant frames; a code stream sending module configured to send the image code stream and the semantic text code stream to a cloud server.

9. A computer device, comprising: includes: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make a computer execute the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video transmission method and apparatus

    CN104320669A

  • Target detection model compression and acceleration method based on pruning and knowledge distillation

    CN112699958A

  • Semantic segmentation-oriented cloud edge collaborative video compression and uploading method and device

    CN114143541A

  • Behavior recognition method and device based on multi-semantic and multi-modal information

    CN115641642A

  • Session video semantic compression framework and method for adaptively selecting reference frame

    CN116405684A

Cited By

  • Video compression method and video decoding method

    CN122027825A

  • Method of video compression and method of video decoding

    CN122027825B