Semantic video coding method, security video system and storage medium

By performing hierarchical semantic analysis and encoding on the original video frames, semantic keyframes, ordinary semantic frames, and background keyframes are generated, which solves the problems of poor encoding flexibility and low compression rate in existing technologies, and achieves more efficient video data transmission and encoding quality adjustment.

CN118803263BActive Publication Date: 2025-10-24CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410010515.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-10-24
Estimated Expiration
2044-01-02

AI Technical Summary

Technical Problem

Existing semantic encoding and decoding methods have poor coding flexibility and fail to make full use of the temporal redundancy information between video frames, resulting in limited improvement in compression ratio.

Method used

By performing target detection on the original video frames, a first basic semantic layer, a target enhancement layer, and a background layer are generated. Semantic keyframes, ordinary semantic frames, and background keyframes are generated according to preset time values. These are then encoded in conjunction with experience quality parameters, and semantic analysis is used to transmit changes in the target and background in layers.

Benefits of technology

It improves encoding flexibility and compression rate, reduces bandwidth usage during upload at the sending end, and can adjust encoding quality based on feedback to meet user experience requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118803263B_ABST
    Figure CN118803263B_ABST
Patent Text Reader

Abstract

The application provides a semantic video coding method and system and a storage medium, and relates to the technical field of video coding. The method comprises the following steps: collecting original video data to perform target detection, obtaining a first basic semantic layer, a target enhancement layer and a background layer; generating a semantic key frame based on the original video frame and the first basic semantic layer every interval of a preset time value; obtaining a quality of experience parameter to generate a common semantic frame based on the first basic semantic layer and the target enhancement layer; generating a background key frame according to the difference degree of adjacent background layers; generating a semantic video stream according to the semantic key frame, the common semantic frame and the background key frame; and sending the semantic video stream. Based on semantic analysis, the target and the background are taken as different layers, and are encoded by different ways, so that the effect of layered transmission is realized. The coding flexibility is improved, the change between the target and the background is considered for coding, the time sequence redundant information is removed, the compression rate is improved, and the bandwidth occupation of the sending end during uploading is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video coding, in particular to a semantic video coding method and system and a storage medium. BACKGROUND

[0002] With the growth of urban construction and the demand for security in families, security monitoring cameras have been widely used, producing a large amount of video data. Traditional coding methods based on H.264 or H.265 standards face performance bottlenecks and require a large amount of computation to improve compression rates. With the development of technology, video coding systems based on artificial intelligence have emerged, bringing new opportunities. There are two main types of video coding methods based on artificial intelligence. One is to combine artificial intelligence technology with traditional coding and replace some modules. However, this method does not completely break away from the traditional coding framework, and the compression rate is not significantly improved. The other method is to completely replace the traditional coding framework with artificial intelligence technology.

[0003] In the prior art, the video coding method based on artificial intelligence analyzes the semantics of the video frame image, divides the image into target and background images, and realizes compression based on the content of the video image, which can improve the compression rate.

[0004] However, the existing semantic coding is still based on the coding of the target and background of the same video frame, i.e., the target and background are associated, and the coding flexibility is poor, which is not conducive to adjusting according to the feedback quality of experience, and the temporal redundancy information between video frames is not considered, which is not conducive to further improving the compression rate. SUMMARY

[0005] The present application provides a semantic video coding method, system and storage medium to solve the defects of poor coding flexibility and inability to further improve the compression rate in the prior art.

[0006] The present application provides a semantic video coding method applied to a sending end, comprising:

[0007] Collecting original video data;

[0008] Semantically encoding the original video data to obtain a semantic video stream;

[0009] Sending the semantic video stream to a cloud platform;

[0010] Wherein, the semantic encoding of the original video data to obtain a semantic video stream comprises:

[0011] According to the original video data, target detection is performed on each original video frame to obtain a corresponding first basic semantic layer, a target enhancement layer and a background layer, the first basic semantic layer representing the category, position and key points of the target, and the target enhancement layer representing the edge details of the target;

[0012] Every interval preset time value, the corresponding original video frame and the first basic semantic layer are combined as a semantic key frame;

[0013] An experience quality parameter is obtained, and according to the experience quality parameter, a normal semantic frame is generated based on the first basic semantic layer and the target enhancement layer or a normal semantic frame is generated based on the first basic semantic layer;

[0014] According to the difference degree of the background layers corresponding to adjacent original video frames, a background key frame is generated;

[0015] According to the semantic key frame, the normal semantic frame and the background key frame, a semantic video stream is generated.

[0016] According to the semantic video coding method provided by the application, further comprising:

[0017] Obtaining experience quality feedback information, and updating the experience quality parameter according to the experience quality feedback information;

[0018] The experience quality feedback information is generated by the cloud platform or the receiving end according to the result of semantic decoding of the semantic video stream.

[0019] The application also provides a semantic video coding method applied to a cloud platform, comprising:

[0020] Obtaining a semantic video stream;

[0021] Semantically decoding the semantic video stream to obtain first decoded video data and first experience quality feedback information;

[0022] Sending the first experience quality feedback information to a sending end, and the first experience quality feedback information is used to update the experience quality parameter of the semantic encoding of the sending end;

[0023] Generally encoding the first decoded video data to obtain a general video stream;

[0024] Sending the general video stream to a receiving end or sending the semantic video stream to the receiving end;

[0025] The semantic video stream comprises a semantic key frame, a normal semantic frame and a background key frame, and the semantic decoding of the semantic video stream to obtain the first decoded video data comprises:

[0026] According to the semantic key frame, a target image is acquired from an original video frame and a first basic semantic layer;

[0027] According to the common semantic frame, the target image is updated from the first basic semantic layer and a target enhancement layer;

[0028] According to the background key frame, a background image is acquired;

[0029] The target image and the background image are combined to generate the first decoded video data.

[0030] According to the semantic video coding method provided by the application, the first experience quality feedback information is acquired by:

[0031] The first decoded video data is subjected to target detection to acquire a second basic semantic layer;

[0032] A first semantic evaluation value is acquired from the first basic semantic layer and the corresponding second basic semantic layer, and the first semantic evaluation value is determined by a target category confidence difference value, an average position deviation distance value and an average key point distance value;

[0033] A first experience evaluation value is acquired, and the first experience evaluation value represents the quality of frame rate, resolution and bit rate;

[0034] The first experience quality feedback information is generated from the first semantic evaluation value and the first experience evaluation value.

[0035] According to the semantic video coding method provided by the application, after the semantic video stream is acquired, the method further includes:

[0036] According to the semantic video stream, a semantic label is acquired, and the semantic label is used as an index for retrieval;

[0037] The semantic video stream is associated with the semantic label and stored;

[0038] In response to an on-demand instruction of a receiving end, the semantic label is retrieved according to the on-demand instruction to confirm the corresponding semantic video stream, the semantic video stream is subjected to semantic decoding to acquire first decoded video data, the first decoded video data is subjected to general encoding to acquire a general video stream, and the general video stream is sent to the receiving end;

[0039] Alternatively, in response to an on-demand instruction of a receiving end, the semantic label is retrieved according to the on-demand instruction to confirm the corresponding semantic video stream, and the semantic video stream is sent to the receiving end.

[0040] The application further provides a semantic video coding method applied to a receiving end, comprising:

[0041] acquiring a semantic video stream;

[0042] performing semantic decoding on the semantic video stream to acquire second decoded video data and second quality of experience feedback information;

[0043] sending the second quality of experience feedback information to a sending end, wherein the second quality of experience feedback information is used to update a quality of experience parameter used for semantic encoding by the sending end;

[0044] playing the second decoded video data;

[0045] alternatively, acquiring a general video stream, performing general decoding on the general video stream to acquire third decoded video data and general quality of experience feedback information, sending the general quality of experience feedback information to the sending end, and playing the third decoded video data;

[0046] wherein the semantic video stream comprises semantic key frames, general semantic frames and background key frames, and the semantic decoding on the semantic video stream to acquire the second decoded video data comprises:

[0047] acquiring a target image according to the semantic key frames and according to original video frames and a first basic semantic layer;

[0048] acquiring and updating the target image according to the general semantic frames and according to the first basic semantic layer and a target enhancement layer;

[0049] acquiring a background image according to the background key frames;

[0050] combining the target image and the background image to generate the second decoded video data.

[0051] According to the semantic video coding method provided by the application, in the semantic decoding on the semantic video stream to acquire the second decoded video data and the second quality of experience feedback information, the acquiring of the second quality of experience feedback information comprises:

[0052] performing target detection on the second decoded video data to acquire a third basic semantic layer;

[0053] acquiring a second semantic evaluation value according to the first basic semantic layer and the corresponding third basic semantic layer, wherein the second semantic evaluation value is determined by a target category confidence difference value, an average position deviation distance value and an average key point distance value;

[0054] obtaining a second experience evaluation value, the second experience evaluation value representing quality of frame rate, resolution and bit rate;

[0055] generating the second experience quality feedback information according to the second semantic evaluation value and the second experience evaluation value.

[0056] The application further provides a security video system, comprising a sending terminal device, a cloud platform and a receiving terminal device, the cloud platform being in communication connection with the sending terminal device and the receiving terminal device respectively;

[0057] The sending terminal device is used for implementing the semantic video coding method applied to the sending terminal;

[0058] The cloud platform is used for implementing the semantic video coding method applied to the cloud platform;

[0059] The receiving terminal device is used for implementing the semantic video coding method applied to the receiving terminal.

[0060] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, the processor implementing the semantic video coding method when executing the program.

[0061] The application further provides a non-transient computer readable storage medium, which stores a computer program, the computer program being executable on a processor to implement the semantic video coding method.

[0062] The semantic video coding method, system and storage medium provided by the application have at least the following beneficial effects: the original video data collected by the sending end, such as security video data, is semantically encoded, specifically, target detection is performed on the original video frame, a first basic semantic layer related to the target, a target enhancement layer and a background layer outside the target are obtained, and the original video frame and the first basic semantic layer are combined as a semantic key frame every interval preset time value, the semantic key frame provides complete information as a reference frame. For ordinary semantic frames that do not meet the interval preset time value condition, according to the quality of experience parameter, the first semantic basic layer and the target enhancement layer are used to generate an ordinary semantic frame, or the first semantic basic layer is used to generate an ordinary semantic frame. According to the time sequence change of the original video frame, when the difference degree of adjacent background layers exceeds the preset range, that is, the background layer changes greatly, a background key frame is generated. The semantic video stream is generated according to the semantic key frame, the ordinary semantic frame and the background key frame. By sending the semantic key frame as the basic reference frame of the video at a preset time value interval, for the video change of the target between the semantic key frames, the ordinary semantic frame is used for updating, and for the video change of the background, the background key frame is used for updating when the background change degree is large. In this way, based on semantic analysis, the target and the background are taken as different layers and encoded in different ways, so as to realize the effect of layered transmission, which is beneficial to improve the coding flexibility, facilitate adjustment according to the feedback quality of experience, and consider the change between the target and the background for coding, which is beneficial to remove the time sequence redundant information, improve the compression rate and reduce the bandwidth occupation of the sending end during uploading. BRIEF DESCRIPTION OF DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0064] Figure 1 is one of the flowcharts of the semantic video coding method applied to the sending end provided by the present application;

[0065] Figure 2 is the second flowchart of the semantic video coding method applied to the sending end provided by the present application;

[0066] Figure 3 is one of the flowcharts of the semantic video coding method applied to the cloud platform provided by the present application;

[0067] Figure 4 is the second flowchart of the semantic video coding method applied to the cloud platform provided by the present application;

[0068] Figure 5 Figure 3 is a flowchart of a semantic video coding method applied to a cloud platform according to an embodiment of the present application;

[0069] Figure 6 Figure 1 is a flowchart of a semantic video coding method applied to a receiving end according to an embodiment of the present application;

[0070] Figure 7 Figure 2 is a flowchart of a semantic video coding method applied to a receiving end according to another embodiment of the present application;

[0071] Figure 8 Figure 7 is a structural diagram of an embodiment of a security video system according to an embodiment of the present application;

[0072] Figure 9 Figure 8 is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0073] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0074] The semantic video coding method of the present application will be described below with reference to Figure 1 and Figure 2 applicable to a sending end, comprising:

[0075] S100: collecting original video data;

[0076] S110: performing semantic coding on the original video data to obtain a semantic video stream;

[0077] S120: sending the semantic video stream to a cloud platform;

[0078] The S110 comprises:

[0079] S111: performing target detection on each original video frame according to the original video data to obtain a corresponding first basic semantic layer, a target enhancement layer and a background layer, the first basic semantic layer representing the category, position and key points of the target, and the target enhancement layer representing the edge details of the target;

[0080] S112: combining the corresponding original video frame and the first basic semantic layer as a semantic key frame every interval of a preset time value;

[0081] S113: Obtain a quality of experience parameter, and generate a common semantic frame based on the first basic semantic layer and the target enhancement layer or generate a common semantic frame based on the first basic semantic layer according to the quality of experience parameter;

[0082] S114: Generate a background key frame according to a difference degree of the background layer corresponding to adjacent original video frames;

[0083] S115: Generate a semantic video stream according to the semantic key frame, the common semantic frame, and the background key frame.

[0084] The original video data collected by the sending end, such as security video data, is semantically encoded. Specifically, target detection is performed on the original video frame to obtain a first basic semantic layer related to the target, a target enhancement layer, and a background layer outside the target. Based on a preset time value, the original video frame and the first basic semantic layer are combined as a semantic key frame every interval of the preset time value, and the semantic key frame provides complete information as a reference frame. For common semantic frames that do not meet the interval preset time value condition, according to a quality of experience parameter, a common semantic frame is generated based on the first semantic basic layer and the target enhancement layer, or a common semantic frame is generated based on the first semantic basic layer. According to the time sequence change of the original video frame, when the difference degree of adjacent background layers exceeds a preset range, that is, the background layer changes greatly, a background key frame is generated. A semantic video stream is generated according to the semantic key frame, the common semantic frame, and the background key frame.

[0085] By sending the semantic key frame as the basic reference frame of the video at a preset time interval, for the video change of the target between the semantic key frames, the common semantic frame is updated, and for the video change of the background, the background key frame is updated when the background change degree is large. In this way, based on semantic analysis, the target and the background are treated as different layers and are encoded in different ways, achieving the effect of layered transmission, which is beneficial to improve the coding flexibility, facilitate adjustment according to the feedback quality of experience, and consider the change between the target and the background for encoding, which is beneficial to remove the time sequence redundant information, improve the compression rate, and reduce the bandwidth occupation of the sending end during uploading.

[0086] It can be understood that in the semantic video stream, a key semantic frame is generated at a preset time value, the key semantic frame includes an original video frame and a first basic semantic layer obtained through semantic analysis, and the first basic semantic layer includes information such as a category, a position, and a key point of a target (such as a person, an object, or the like). The common semantic frame between two key semantic frames necessarily includes a basic semantic layer, and based on a quality of experience parameter, can also include a target enhancement layer, and the target enhancement layer includes information such as an edge detail of the target. The semantic key frame is taken as a reference frame, target image changes are updated based on the common semantic frame, and video changes of the target image are realized. At the same time of realizing the video changes of the target image, a background image is updated based on a background key frame, and video changes of the background image are realized. In this way, between two semantic key frames, the original video frame does not need to be transmitted, the target image is updated with the common semantic frame, which is beneficial to improve a compression rate, and at the same time, the background key frame is generated only when a background change degree is greater than a preset range, which is beneficial to further improve the compression rate.

[0087] In some embodiments of the present application, the generation of the background key frame according to the difference degree of the background layers corresponding to the adjacent original video frames includes: obtaining a first background value and a second background value of the adjacent original video frames, and calculating the difference between the first background value and the second background value, which can be calculated through the following formula:

[0088] ΔBG=|BG r -BG r-1 |

[0089] wherein, ΔBG is the difference value, BG t is the first background value, and BG t-1 is the second background value.

[0090] When the difference value is greater than a preset background threshold, the background layer corresponding to the original video frame is taken as the background key frame. The first background value and the second background value can be determined according to the RGB value, the brightness value, or the like of the image.

[0091] In some embodiments of the present application, the semantic encoding process can be implemented through a CGAN (Conditional Generative Adversarial Network) model. The CGAN model introduces conditional information, so that the generator is more controllable, can generate more meaningful data according to a given condition, improves the quality of the generated data, and supports multi-modal generation.

[0092] Reference Figure 2 In some embodiments of the present application applied to the semantic video encoding and decoding method of the sending end, the method further includes:

[0093] S116: Obtain quality of experience feedback information, and update the quality of experience parameter according to the quality of experience feedback information.

[0094] The experience quality feedback information is generated by a result of semantic decoding of the semantic video stream by a cloud platform or a receiving end.

[0095] The experience quality parameter can be updated according to the feedback experience quality information, which can be a result of semantic decoding of the semantic video stream by the cloud platform or the receiving end, evaluating the semantic analysis effect when the semantic encoding of the sending end, and generating the experience quality feedback information. In addition, the experience quality feedback information can also be integrated by combining the traditional experience quality information based on frame rate, resolution and bit rate. In this way, based on the feedback information of the semantic analysis effect and the feedback information of the video quality based on the frame rate, resolution and bit rate, the sending end updates the experience quality parameter, adjusts the general semantic frame, and further adjusts the quality of the video to meet the demand.

[0096] It should be emphasized that the traditional conventional experience quality (QoE, Quality of Experience) only considers frame rate, resolution, bit rate and other parameters, and does not involve the effect of semantic analysis. The present application comprehensively considers the effect of semantic analysis and the traditional experience quality information to generate the experience quality feedback information, and then the sending end updates the experience quality parameter based on the obtained experience quality feedback information, adjusts the generated semantic video stream, which is beneficial to adapt to the way of semantic encoding and improve the quality of semantic encoding.

[0097] Reference Figure 3 The present application also provides a semantic video encoding and decoding method applied to a cloud platform, comprising:

[0098] S200: obtaining a semantic video stream;

[0099] S210: performing semantic decoding on the semantic video stream to obtain first decoded video data and first experience quality feedback information;

[0100] S220: sending the first experience quality feedback information to a sending end, wherein the first experience quality feedback information is used to update an experience quality parameter of the sending end for semantic encoding;

[0101] S230: performing general encoding on the first decoded video data to obtain a general video stream;

[0102] S240: sending the general video stream to a receiving end, or sending the semantic video stream to the receiving end;

[0103] The semantic video stream comprises semantic key frames, general semantic frames and background key frames, and the S210 comprises:

[0104] S211: obtaining a target image according to the semantic key frame and according to the original video frame and the first basic semantic layer;

[0105] S212: updating the target image according to the normal semantic frame and according to the first basic semantic layer and the target enhancement layer;

[0106] S213: obtaining a background image according to the background key frame;

[0107] S214: combining the target image and the background image to generate the first decoded video data.

[0108] The cloud platform obtains the semantic video stream from the sending end, and performs semantic decoding on the semantic video stream. Specifically, the cloud platform is based on the frame timing of the semantic video stream. When it is a semantic key frame, a target image is obtained from the original video frame based on the corresponding first basic semantic layer. When it is a semantic key frame, the target image is updated based on the corresponding first basic semantic layer and the target enhancement layer, that is, the position, key point, edge detail, etc. of the target image are changed. When it is a background key layer, a background image is obtained. Based on the target image and the background image, the two are combined to form a complete video image, and the target image and the background image are continuously updated based on the above process to realize video image change and obtain first decoded video data. In this way, the target image and the background image are obtained through layered transmission, which is conducive to improving flexibility. At the same time, the normal semantic layer does not include the original video frame when transmitting the semantic video stream, and the background key frame will only be generated when the background change degree is greater than a preset range, which is conducive to reducing the bandwidth occupied when transmitting the semantic video stream and improving the compression rate.

[0109] After the cloud platform obtains the first decoded video data, the cloud platform performs semantic analysis on the first decoded video data to obtain first quality of experience feedback information. The first quality of experience feedback information can indicate the difference between the first decoded video data and the original video data to measure the quality of semantic encoding. The first quality of experience feedback information is sent to the sending end to update the quality of experience parameter of the sending end, which is conducive to improving the quality of semantic encoding of the sending end.

[0110] Different receiving ends have different processing performances, such as different models of mobile phones, tablet computers, etc. Therefore, for receiving end devices with strong processing performance, the semantic video stream is directly forwarded to the receiving end. Since the compression rate of the semantic video stream is high, it is conducive to reducing the occupied bandwidth and improving the transmission efficiency. For receiving ends with weak processing performance, the first decoded video data is normally encoded, such as the normal encoding and decoding mode of H.264 or H.265 standard, to obtain a normal video stream sent to the receiving end, so that the receiving end can decode and play the normal video stream through the normal decoding mode. In this way, the performance difference of the receiving end can be adapted.

[0111] In some embodiments of the present application, in the S214, when the target image is combined with the background image, feathering processing between the target edge and the background edge can be performed to soften the transition edge and improve the image effect.

[0112] In some embodiments of the present application, the process of semantic decoding performed by the cloud platform can be implemented by a CGAN model. The CGAN model introduces conditional information, so that the generator is more controllable, can generate more meaningful data according to the given conditions, improves the quality of the generated data, and supports multi-modal generation.

[0113] Reference Figure 4 In some embodiments of the present application applied to the semantic video coding and decoding method of the cloud platform, in the S210, the step of obtaining the first quality of experience feedback information comprises:

[0114] S215: performing target detection on the first decoded video data to obtain a second basic semantic layer;

[0115] S216: obtaining a first semantic evaluation value according to the first basic semantic layer and the corresponding second basic semantic layer, the first semantic evaluation value being determined by a category confidence difference value, an average position deviation distance value, and an average key point distance value;

[0116] S217: obtaining a first experience evaluation value, the first experience evaluation value representing the quality of frame rate, resolution, and bit rate;

[0117] S218: generating the first quality of experience feedback information according to the first semantic evaluation value and the first experience evaluation value.

[0118] By performing target detection on the first decoded video data similar to the sending end, the effect of semantic analysis is achieved, and a second basic semantic layer is obtained. The second basic semantic layer and the first basic semantic layer have the same effect, both representing the category, position, and key point of the target. By comparing the second basic semantic layer with the first basic semantic layer, the category confidence difference value and the average position deviation distance are obtained, and then the first semantic evaluation value is obtained. At the same time, the first experience evaluation value (corresponding to the traditional quality of experience information) obtained according to the frame rate, resolution, and bit rate of the video is combined to comprehensively generate the first quality of experience feedback information. In this way, the first semantic evaluation value based on semantic analysis and the first experience evaluation value obtained by the traditional quality of experience are combined, so that the generated first quality of experience feedback information is more suitable for the environment of semantic coding and decoding, and the semantic encoding adjustment of the sending end is more accurate.

[0119] It can be understood that the first and second basic semantic layers are compared for the corresponding frame. The first decoded video data is obtained by semantic decoding based on the semantic video stream. The first decoded video data has a difference caused by coding and decoding from the original video data, but the video image content of the first decoded video data is basically consistent with the video image content of the original video data. The semantic basic layers obtained by semantic analysis of the two are different, and the semantic encoding of the sending end is adjusted according to the difference to reduce the difference and improve the quality of semantic encoding.

[0120] In some embodiments of the application, the category confidence difference value in S216 can be calculated by the following formula:

[0121] Δscore i =|score i -score rec_i |

[0122] Wherein, Δscore i is the category confidence difference value, score i is the target category confidence of the first basic semantic layer, score rec_i is the target category confidence of the second basic semantic layer, and i represents the ith target. In the case of multiple targets, the average value of each target is taken.

[0123] In some embodiments of the application, the average position deviation distance value in S216 can be calculated by the following formula:

[0124]

[0125] Wherein, Δkps i is the average position deviation distance value, kps i is the target position corresponding to the first basic semantic layer, kps rec_i is the target position corresponding to the second basic semantic layer, d is the Euclidean distance calculation symbol, i represents the ith target, and N is the total number of targets.

[0126] In some embodiments of the application, the average key point distance value in S216 can be calculated by the following formula:

[0127] Δloc i =|center i -center rec_i |

[0128] Wherein, Δloc i is the average key point distance value, center i is the target recognition frame center point corresponding to the first semantic basic layer, centerrec_i The target recognition frame center point corresponding to the second semantic base layer is represented as i, and i represents the i-th target. In the case of multiple targets, the average value of each target is taken.

[0129] Reference Figure 5 In some embodiments of the semantic video coding method applied to the cloud platform, after S200, the method further comprises:

[0130] S250: obtaining a semantic label according to the semantic video stream, the semantic label being used as an index for retrieval;

[0131] S260: associating and storing the semantic video stream with the semantic label;

[0132] S270: in response to an on-demand instruction from a receiving end, retrieving the semantic label according to the on-demand instruction to identify the corresponding semantic video stream, performing semantic decoding on the semantic video stream to obtain first decoded video data, performing general encoding on the first decoded video data to obtain a general video stream, and sending the general video stream to the receiving end;

[0133] Alternatively, S280: in response to an on-demand instruction from a receiving end, retrieving the semantic label according to the on-demand instruction to identify the corresponding semantic video stream, and sending the semantic video stream to the receiving end.

[0134] In the security scenario, the video captured by the camera device, i.e., the sending end device, is encoded and decoded and sent to the terminal device of the user, i.e., the receiving end, to achieve the effect of real-time live streaming. At the same time, in some cases, it is also necessary to trace back the historical video, i.e., to perform on-demand. Therefore, the cloud platform stores the semantic video stream and associates it with the semantic label, so as to facilitate retrieval based on the semantic label when on-demand. In response to the on-demand instruction, after obtaining the required semantic video stream, based on the processing performance of the receiving end, when the processing performance of the receiving end is weak, the semantic video stream is decoded to obtain first decoded video data, the first decoded video data is generally encoded to obtain a general video stream, and the general video stream is sent to the receiving end; when the processing performance of the receiving end is strong, the semantic video stream is directly forwarded to the receiving end. Since the compression rate of the semantic video stream is high, it is beneficial to reduce the transmission bandwidth occupation and improve the transmission efficiency. In this way, the on-demand function is realized, and at the same time, the performance difference of the receiving end can be adapted to the processing when on-demand.

[0135] Reference Figure 6 The present application also provides a semantic video coding method applied to a receiving end, comprising:

[0136] S300: obtaining a semantic video stream;

[0137] S310: performing semantic decoding on the semantic video stream to obtain second decoded video data and second quality of experience feedback information;

[0138] S320: sending the second quality of experience feedback information to the sending end, the second quality of experience feedback information being used to update a quality of experience parameter used by the sending end for semantic encoding;

[0139] S330: playing the second decoded video data;

[0140] Alternatively, S340: obtaining a general video stream, performing general decoding on the general video stream to obtain third decoded video data and general quality of experience feedback information, sending the general quality of experience feedback information to the sending end, and playing the third decoded video data;

[0141] The semantic video stream comprises semantic key frames, general semantic frames and background key frames, and the S310 comprises:

[0142] S311: obtaining a target image according to the semantic key frames and according to an original video frame and a first basic semantic layer;

[0143] S312: updating the target image according to the general semantic frames and according to the first basic semantic layer and a target enhancement layer;

[0144] S313: obtaining a background image according to the background key frames;

[0145] S314: combining the target image and the background image to generate the second decoded video data.

[0146] When the processing performance of the receiving end is strong, the receiving end obtains a semantic video stream from a cloud platform and performs semantic decoding on the semantic video stream. Specifically, the receiving end is based on the frame timing of the semantic video stream. When it is a semantic key frame, a target image is obtained from an original video frame based on a corresponding first basic semantic layer. When it is a semantic key frame, the target image is updated based on the corresponding first basic semantic layer and a target enhancement layer, i.e., the position, key points, edge details, etc. of the target image are changed. When it is a background key layer, a background image is obtained. Based on the target image and the background image, the two are combined to form a complete video image. The target image and the background image are constantly updated based on the above process to realize video image changes and obtain second decoded video data for playing. In this way, the target image and the background image are obtained through layered transmission, which is conducive to improving flexibility. At the same time, the general semantic layer does not include the original video frame when transmitting, and the background key frame is only generated when the background change degree is greater than a preset range, which is conducive to reducing the bandwidth occupied when transmitting the semantic video stream and improving the compression rate.

[0147] In the case that the processing performance of the receiving end is weak, the receiving end obtains the general video stream from the cloud platform, and plays the third decoded video data obtained through general decoding.

[0148] In some embodiments of the present application, in the S314, when the target image is combined with the background image, feather processing between the target edge and the background edge can be performed to soften the transition edge and improve the image effect.

[0149] It can be understood that the process of semantic decoding of the semantic video stream by the cloud platform and the receiving end is essentially the same, except that the process of semantic decoding is completed in the cloud platform or in the receiving end according to the processing performance of the receiving end. The semantic decoding is completed in the cloud platform, which can reduce the requirement on the processing performance of the receiving end; the semantic decoding is completed in the receiving end, which can simplify the processing process of the cloud platform, without the need for semantic decoding and general encoding, thereby improving the processing efficiency, and based on the high compression rate of the semantic video stream, it is also beneficial to improve the transmission efficiency.

[0150] In some embodiments of the present application, the process of semantic decoding of the receiving end can be implemented through a CGAN model. The CGAN model introduces conditional information, so that the generator is more controllable, can generate more meaningful data according to the given conditions, improves the quality of the generated data, and supports multi-modal generation.

[0151] Reference Figure 7 In some embodiments of the present application applied to the semantic video coding and decoding method of the receiving end, in the S310, the step of obtaining the second quality of experience feedback information comprises:

[0152] S315: performing target detection on the second decoded video data to obtain a third basic semantic layer;

[0153] S316: obtaining a second semantic evaluation value according to the first basic semantic layer and the corresponding third basic semantic layer, the second semantic evaluation value being determined by a category confidence difference value, an average position deviation distance value and an average key point distance value;

[0154] S317: obtaining a second experience evaluation value, the second experience evaluation value representing the quality of frame rate, resolution and bit rate;

[0155] S318: generating the second quality of experience feedback information according to the second semantic evaluation value and the second experience evaluation value.

[0156] By performing target detection on the second decoded video data similar to the sending end, the effect of semantic analysis is achieved, and a third basic semantic layer is obtained. The third basic semantic layer and the first basic semantic layer have the same effect of representing the category, position and key point of the target. By comparing the third basic semantic layer and the first basic semantic layer, a category confidence difference value and an average position deviation distance are obtained, and then a second semantic evaluation value is obtained. At the same time, a second experience evaluation value (corresponding to the traditional experience quality information) is obtained according to the frame rate, resolution and bit rate of the video, and a second experience quality feedback information is generated by comprehensively combining the second experience quality feedback information. In this way, the second semantic evaluation value based on semantic analysis and the second experience evaluation value obtained by the traditional experience quality are combined, so that the generated second experience quality feedback information is more suitable for the environment of semantic coding and decoding, and the semantic encoding adjustment of the sending end is more accurate.

[0157] In some embodiments of the present application, the category confidence difference value in S316 can be calculated by the following formula:

[0158] Δscore i =|score i -score rec_i |

[0159] Wherein, Δscore i is the category confidence difference value, score i is the target category confidence corresponding to the first basic semantic layer, score rec_i is the target category confidence corresponding to the third basic semantic layer, and i represents the ith target. In the case of multiple targets, the average value of each target is taken.

[0160] In some embodiments of the present application, the average position deviation distance value in S316 can be calculated by the following formula:

[0161]

[0162] Wherein, Δkps i is the average position deviation distance value, kps i is the target position corresponding to the first basic semantic layer, kps rec_i is the target position corresponding to the third basic semantic layer, d is the Euclidean distance calculation symbol, i represents the ith target, and N is the total number of targets.

[0163] In some embodiments of the present application, the average key point distance value in S316 can be calculated by the following formula:

[0164] Δloc i =|center i -center rec_i |

[0165] wherein, Delta loc i is the average keypoint distance value, center i is the target recognition box center point corresponding to the first semantic base layer, center rec_i is the target recognition box center point corresponding to the third semantic base layer, i represents the i-th target. In the case of multiple targets, the average value of each target is taken.

[0166] It can be understood that the process of the cloud platform generating the first quality of experience feedback information and the process of the receiving end generating the second quality of experience feedback information are essentially the same. Whether to generate the first quality of experience feedback information or the second quality of experience feedback information depends on whether the semantic decoding process is performed in the cloud platform or in the receiving end. It should be noted that the process of the receiving end generating the quality of experience feedback information is different from the process of generating the second quality of experience feedback information. Since the quality of experience feedback information is generated based on general decoding, it does not involve semantic analysis. The quality of experience feedback information (which is essentially traditional quality of experience information) includes traditional video-based frame rate, resolution, bit rate, and other generated quality of experience information.

[0167] With reference to Figure 8 , the application further provides a security video system, comprising: a sending end device, a cloud platform, and a receiving end device, the cloud platform is in communication connection with the sending end device and the receiving end device respectively;

[0168] The sending end device is used to implement the above-mentioned semantic video coding and decoding method applied to the sending end.

[0169] The cloud platform is used to implement the above-mentioned semantic video coding and decoding method applied to the cloud platform.

[0170] The receiving end device is used to implement the above-mentioned semantic video coding and decoding method applied to the receiving end.

[0171] The security video system provided by the application can be correspondingly referred to each other with the above-mentioned semantic video coding and decoding methods applied to the sending end, the cloud platform, and the receiving end, and will not be described again.

[0172] The working process of the security video system is as follows: the original video data collected by the sending end device, such as security video data, is semantically encoded to generate a semantic video stream uploaded to the cloud platform. For a receiving end device with strong processing performance, the cloud platform directly forwards the semantic video stream to the receiving end device; for a receiving end device with weak processing performance, the semantic video stream is semantically decoded to obtain first decoded video data and generate first quality of experience feedback information sent to the sending end device, and the first decoded video data is normally encoded to generate a normal video stream sent to the receiving end device. The receiving end device with strong performance acquires the semantic video stream for semantic decoding to obtain second decoded video data and generate second quality of experience feedback information sent to the sending end device, and plays the second decoded video data; the receiving end device with weak performance acquires the normal video stream for normal decoding to obtain third decoded video data and generate normal quality of experience feedback information sent to the sending end device, and plays the third decoded video data. The sending end updates the quality of experience parameters according to the acquired quality of experience feedback information to adjust the semantic encoding. When the receiving end sends an on-demand instruction to the cloud platform, the cloud platform retrieves the semantic tags according to the on-demand instruction to acquire the corresponding semantic video stream, and performs the corresponding sending process according to the performance of the receiving end.

[0173] In the security scenario, the sending end device is usually a smart camera with video acquisition capability and semantic encoding; the cloud platform, i.e., the cloud server, can perform tasks such as semantic decoding, quality of experience feedback information calculation, H.264 / H.265 encoding, video file storage, video semantic retrieval, etc.; the receiving end can be a mobile phone, a platform computer, a computer, etc., which interacts with the user and can realize applications such as on-demand, live broadcast, search, recommendation, etc., and has H.264 / H.265 decoding capability, and some high-performance devices also have the capability of semantic decoding and quality of experience feedback information calculation.

[0174] Figure 9 An example of an entity structure diagram of an electronic device is shown in Figure 9 The electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communications bus 840. The processor 810 can invoke the logical instructions in the memory 830 to execute the semantic video encoding and decoding method described above.

[0175] The electronic device, which can be used as a sending end device, a cloud platform device, or a receiving end device, implements the corresponding semantic video encoding and decoding method.

[0176] In addition, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0177] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the semantic video coding method provided by the above-mentioned methods.

[0178] In the description of the present application, it should be understood that the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically limited.

[0179] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present description, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the skilled in the art can combine and combine the different embodiments or examples described in the present description and the features of the different embodiments or examples without contradiction.

[0180] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of semantic video coding, characterized by, Applied to a sending end, comprising: Collecting original video data; Semantic coding the original video data to obtain a semantic video stream; Sending the semantic video stream to a cloud platform; Wherein, the semantic coding the original video data to obtain a semantic video stream comprises: According to the original video data, target detection is performed on each original video frame to obtain a corresponding first basic semantic layer, a target enhancement layer and a background layer, the first basic semantic layer representing the category, position and key points of the target, and the target enhancement layer representing the edge details of the target; Every interval preset time value, the corresponding original video frame and the first basic semantic layer are combined as a semantic key frame; An experience quality parameter is obtained, and according to the experience quality parameter, an ordinary semantic frame is generated based on the first basic semantic layer and the target enhancement layer or based on the first basic semantic layer; According to the difference degree of the background layers corresponding to adjacent original video frames, a background key frame is generated; According to the semantic key frame, the ordinary semantic frame and the background key frame, a semantic video stream is generated; Experience quality feedback information is obtained, and the experience quality parameter is updated according to the experience quality feedback information, and the experience quality feedback information is generated by the result obtained by the cloud platform or the receiving end performing semantic decoding on the semantic video stream.

2. A method of semantic video coding, characterized by, Applied to a cloud platform, comprising: Obtaining a semantic video stream; Semantic decoding the semantic video stream to obtain first decoded video data and first experience quality feedback information; Sending the first experience quality feedback information to the sending end, which is used to update the experience quality parameter of the sending end for semantic coding; Ordinary coding the first decoded video data to obtain an ordinary video stream; Sending the ordinary video stream to the receiving end or sending the semantic video stream to the receiving end; Wherein, the semantic video stream comprises a semantic key frame, an ordinary semantic frame and a background key frame, and the semantic decoding the semantic video stream to obtain first decoded video data comprises: According to the semantic key frame, a target image is obtained according to the original video frame and the first basic semantic layer; According to the ordinary semantic frame, the target image is updated according to the first basic semantic layer and the target enhancement layer, the first basic semantic layer representing the category, position and key points of the target, and the target enhancement layer representing the edge details of the target; According to the background key frame, a background image is obtained, and the background key frame is generated according to the difference degree of the background layers corresponding to adjacent original video frames; The target image and the background image are combined to generate the first decoded video data; The first experience quality feedback information obtaining step comprises: Target detection is performed on the first decoded video data to obtain a second basic semantic layer; According to the first basic semantic layer and the corresponding second basic semantic layer, a first semantic evaluation value is obtained, and the first semantic evaluation value is determined by a target category confidence difference value, an average position deviation distance value and an average key point distance value. obtaining a first experience evaluation value, the first experience evaluation value representing quality of frame rate, resolution and bit rate; generating the first experience quality feedback information according to the first semantic evaluation value and the first experience evaluation value.

3. The method of Claim 2, wherein, After the semantic video stream is obtained, further comprising: obtaining a semantic label according to the semantic video stream, the semantic label being used as an index for retrieval; associating and storing the semantic video stream with the semantic label; in response to an on-demand instruction of a receiving end, retrieving the semantic label according to the on-demand instruction to confirm the corresponding semantic video stream, performing semantic decoding on the semantic video stream, obtaining first decoded video data, performing general encoding on the first decoded video data, obtaining a general video stream, and sending the general video stream to the receiving end; or, in response to an on-demand instruction of a receiving end, retrieving the semantic label according to the on-demand instruction to confirm the corresponding semantic video stream, and sending the semantic video stream to the receiving end.

4. A method of semantic video coding, characterized by, applied to a receiving end, comprising: obtaining a semantic video stream; performing semantic decoding on the semantic video stream to obtain second decoded video data and second experience quality feedback information; sending the second experience quality feedback information to a sending end, the second experience quality feedback information being used to update experience quality parameters used for semantic encoding by the sending end; playing the second decoded video data; or, obtaining a general video stream, performing general decoding on the general video stream to obtain third decoded video data and experience quality general feedback information, sending the experience quality general feedback information to a sending end, and playing the third decoded video data; wherein the semantic video stream comprises semantic key frames, general semantic frames and background key frames, and the step of performing semantic decoding on the semantic video stream to obtain second decoded video data comprises: obtaining a target image according to the semantic key frames and according to original video frames and a first basic semantic layer; updating the target image according to the general semantic frames and according to the first basic semantic layer and a target enhancement layer, the first basic semantic layer representing a category, a position and key points of a target, and the target enhancement layer representing edge details of the target; obtaining a background image according to the background key frames, the background key frames being generated according to a difference degree of a corresponding background layer of adjacent original video frames; combining the target image and the background image to generate the second decoded video data; the step of obtaining the second experience quality feedback information comprises: performing target detection on the second decoded video data to obtain a third basic semantic layer; obtaining a second semantic evaluation value according to the first basic semantic layer and the corresponding third basic semantic layer, the second semantic evaluation value being determined by a category confidence difference value, an average position deviation distance value and an average key point distance value; obtaining a second experience evaluation value, the second experience evaluation value representing quality of frame rate, resolution and bit rate; generating the second experience quality feedback information according to the second semantic evaluation value and the second experience evaluation value.

5. A security video system characterized by, comprising: A sending terminal device, a cloud platform, and a receiving terminal device, wherein the cloud platform is communicatively connected with the sending terminal device and the receiving terminal device respectively; The sending terminal device is configured to implement the semantic video coding method according to claim 1. The cloud platform is configured to implement the semantic video coding method according to claim 2 or 3. The receiving terminal device is configured to implement the semantic video coding method according to claim 4.

6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the semantic video coding method according to any one of claims 1 to 4 when executing the program.

7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the semantic video coding method according to any one of claims 1 to 4 when executed by the processor.

Citation Information

Patent Citations

  • Apparatus and method for video coding and decoding and computer program

    CN111543060A

  • Hierarchical video coding method, device and product based on semantic information

    CN116723333A