Video Compression Method, Device, Electronic Device, and Machine-readable Storage Medium

By using semantic analysis model to perform semantic analysis on videos and dynamically adjusting the compression ratio of video frames and regions, the problem of poor flexibility of traditional video compression methods is solved, and a more efficient and flexible video compression effect is achieved.

CN115022645BActive Publication Date: 2025-05-27HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210731325.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-05-27
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

Traditional video compression methods are poor in flexibility and cannot dynamically adjust the compression ratio according to the video content.

Method used

Using the trained semantic analysis model, semantic analysis of the compressed video is performed to obtain the semantic information of the video, including the correspondence between the timestamp and the region position information. Based on this information, the video is compressed adaptively, and the compression ratio of the specified target/behavior area is different from that of other areas.

Benefits of technology

Improves the flexibility of video compression, and can dynamically adjust the compression ratio according to the video content, ensuring image quality of the specified target/behavior area and the compression efficiency of the overall video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115022645B_ABST
    Figure CN115022645B_ABST
Patent Text Reader

Abstract

The present application provides a video compression method, apparatus, electronic device, and machine-readable storage medium. The method includes: using a trained semantic analysis model to perform semantic analysis on a video to be compressed to obtain semantic information of the video to be compressed; and compressing the video to be compressed according to the semantic information of the video to be compressed. This method can improve the flexibility of video compression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent processing technology, and in particular, to a video compression method, apparatus, electronic device, and machine-readable storage medium. Background Art

[0002] Video compression can reduce the size of video data, thereby saving the storage space required for storing video data and saving the bandwidth resources consumed during video data transmission, and improving the transmission efficiency of video data.

[0003] However, in traditional video compression schemes, when performing video compression, a unified compression ratio is usually used to compress the entire video data, resulting in poor flexibility in video compression. Summary of the Invention

[0004] In view of this, this application provides a video compression method, apparatus, electronic device, and machine-readable storage medium to solve the problem of poor flexibility in traditional video compression methods.

[0005] Specifically, this application is implemented through the following technical solutions:

[0006] According to the first aspect of the embodiments of this application, a video compression method is provided, including:

[0007] Using a trained semantic analysis model, perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed, where the semantic information includes the correspondence between timestamps and regional location information; the timestamp is the timestamp corresponding to the first video frame, and the first video frame is the video frame in the video to be compressed that has a specified target / behavior; the regional location information is the location information of the first region, and the first region is the region of the specified target / behavior in the first video frame;

[0008] Compress the video to be compressed according to the semantic information of the video to be compressed; wherein, the compression ratio of the first region is different from that of other regions.

[0009] According to the second aspect of the embodiments of this application, a video compression apparatus is provided, including:

[0010] A semantic analysis unit, configured to use a trained semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed, where the semantic information includes the correspondence between timestamps and regional location information; the timestamp is the timestamp corresponding to the first video frame, and the first video frame is the video frame in the video to be compressed that has a specified target / behavior; the regional location information is the location information of the first region, and the first region is the region of the specified target / behavior in the first video frame;

[0011] A compression unit for compressing the video to be compressed according to the semantic information of the video to be compressed; wherein, the compression ratio of the first region is different from that of other regions.

[0012] According to the third aspect of the embodiments of the present application, there is provided an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the method provided in the first aspect.

[0013] According to the fourth aspect of the embodiments of the present application, there is provided a machine-readable storage medium, where machine-executable instructions are stored in the machine-readable storage medium, and when the machine-executable instructions are executed by a processor, the method provided in the first aspect is implemented.

[0014] The technical solution provided by the present application can at least bring the following beneficial effects:

[0015] By training a semantic analysis model for semantic analysis of video data, and using the trained semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the compressed video, when compressing the video to be compressed according to the semantic information of the video to be compressed, different compression ratios are used for the regions with a specified target / behavior and the regions without a specified target / behavior, improving the flexibility of video compression. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic flowchart of a video compression method shown in an exemplary embodiment of the present application;

[0017] Figure 2 is a schematic flowchart of a video compression method shown in an exemplary embodiment of the present application;

[0018] Figure 3 is a schematic structural diagram of a video compression device shown in an exemplary embodiment of the present application;

[0019] Figure 4 is a schematic structural diagram of a video compression device shown in an exemplary embodiment of the present application;

[0020] Figure 5 is a schematic structural diagram of a video compression device shown in an exemplary embodiment of the present application;

[0021] Figure 6 is a schematic hardware structure diagram of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0023] The terms used in this application are for the purpose of describing particular embodiments only and are not intended to limit the present application. The singular forms "a", "said", and "the" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0024] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0025] Please refer to Figure 1 , which is a schematic flowchart of a video compression method provided by an embodiment of the present application. As Figure 1 shown, the video compression method may include the following steps:

[0026] It should be noted that the sequence numbers of the steps in the embodiments of the present application do not imply the order of execution. The execution order of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0027] Step S100: Use the trained semantic analysis model to perform semantic analysis on the video to be compressed, and obtain the semantic information of the video to be compressed. The semantic information includes the correspondence between the timestamps and the regional location information; the timestamp is the timestamp corresponding to the first video frame, and the first video frame is the video frame in the video to be compressed that has a specified target / behavior; the regional location information is the location information of the first region, and the first region is the region of the specified target / behavior in the first video frame.

[0028] In the embodiments of the present application, for any video file to be compressed (referred to as the video to be compressed herein), the pre-trained semantic analysis model can be used to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed, so as to use the semantic information of the video to be compressed to guide the compression of different video frames or different regions of the same video frame in the video to be compressed.

[0029] Exemplarily, the above semantic analysis model can be used to perform semantic analysis on the input video data, determine whether a specified target / behavior exists in each video frame of the video data, and output the corresponding relationship between the time stamps of the video frames where the specified target / behavior exists (referred to as the first video frames in this article), and the position information of the regions where the specified target / behavior exists in the first video frames (referred to as the first regions in this article) (referred to as the region position information in this article).

[0030] Exemplarily, the video to be compressed can be input into the trained semantic analysis model in the form of an image sequence (i.e., a video frame sequence), and the semantic analysis model is used to respectively determine whether a specified target / behavior exists in each video frame, and the position information of the region of the specified target / behavior in the video frame in the video frame where the specified target / behavior exists.

[0031] Exemplarily, the above specified targets may include, but are not limited to, people, animals, vehicles, etc.

[0032] Exemplarily, the above specified behaviors may include, but are not limited to, smoking, fighting, etc.

[0033] Step S110: Compress the video to be compressed according to the semantic information of the video to be compressed; wherein, the compression ratio of the first region is different from that of other regions.

[0034] In the embodiments of the present application, when performing video compression on the video to be compressed, instead of compressing the entire video to be compressed using a unified compression ratio, different compression ratios can be used for different video frames according to whether a specified target / behavior exists in the video frames, and in the same video frame, different compression ratios can also be used for different regions according to whether a specified target / behavior exists in the regions.

[0035] Exemplarily, when compressing the video to be compressed according to the semantic information of the compressed video, the compression ratio of the first region (i.e., the region where the specified target / behavior exists in the video frame) is different from that of other regions (regions where the specified target behavior does not exist, such as regions in non-first video frames, regions in the first video frame where the specified target / behavior does not exist, etc.).

[0036] It can be seen that in Figure 1 the shown method flow, by training a semantic analysis model for performing semantic analysis on video data, and using the trained semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the compressed video, when compressing the video to be compressed according to the semantic information of the video to be compressed, different compression ratios are used for the regions where the specified target / behavior exists and the regions where the specified target / behavior does not exist, improving the flexibility of video compression.

[0037] In some embodiments, a trained semantic analysis model is used to perform semantic analysis on the video to be compressed, and the semantic information of the video to be compressed can be obtained, including:

[0038] The video to be compressed and predefined text of interest are input into the semantic analysis model, and the semantic analysis model is used to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed; the text of interest includes text description information of a specified target / behavior.

[0039] Exemplarily, in order to improve the pertinence of semantic analysis by the semantic analysis model and improve the controllability of video compression, an interested target / behavior (which can be called an interested target / behavior) can also be predefined in advance, so as to use the semantic analysis model to determine whether there is an interested target / behavior in the input video data, and output the time stamp corresponding to the video frame where the interested target / behavior exists, as well as the regional position information of the interested target / behavior in the video frame.

[0040] Exemplarily, the above predefined interested target / behavior can be input into the semantic analysis model in the form of text description information as the text of interest, so that the semantic analysis model can perform semantic analysis on the input video data (such as the video to be compressed) according to the text of interest, determine whether there is an interested target / behavior in the video to be compressed, and output the time stamp corresponding to the video frame where the interested target / behavior exists, as well as the regional position information of the interested target / behavior in the video frame, that is, the above specified target / behavior can be the predefined interested target / behavior.

[0041] Correspondingly, in order to achieve the above purpose, during the training process of the semantic analysis model, the recognition ability of the semantic analysis model for the text of interest can also be trained, so that when the text of interest is input into the trained semantic analysis model, the semantic analysis model can determine the interested target / behavior according to the text of interest.

[0042] In some embodiments, the above compression of the video to be compressed according to the semantic information of the video to be compressed can include:

[0043] Compress the first region using a first compression ratio and compress other regions using a second compression ratio according to the semantic information of the video to be compressed; wherein, the first compression ratio is higher than the second compression ratio.

[0044] Exemplarily, in order to ensure the image quality of the region with the specified target / behavior in the compressed video, when compressing the video to be compressed according to the semantic information of the video to be compressed, a higher compression ratio (referred to as the first compression ratio in this article) can be used to compress the first region, and a lower compression ratio (referred to as the second compression ratio in this article) can be used to compress other regions. Thus, in the compressed video data, the image quality of the first region can still be relatively high, and at the same time, the compression ratio of the overall video data will not be too high, ensuring the compression efficiency of the overall video data.

[0045] Exemplarily, the compression ratio can be the ratio of the size of the compressed video data to the size of the video data before compression. That is, for the same video data, the higher the compression ratio, the larger the size of the compressed video data.

[0046] It should be noted that in the embodiments of the present application, in addition to using different compression ratios for the first region and other regions in the above manner, different compression ratios can also be used for video frames with the specified target / behavior and video frames without the specified target / behavior.

[0047] For example, assume that there is a specified target / behavior in video frame 1, the region with the specified target / behavior is region 1, the region without the specified target / behavior is region 2, and there is no specified target / behavior in video frame 2. Then, compression ratio 1 can be used to compress region 1 in video frame 1, compression ratio 2 can be used to compress region 2 in video frame 1, and compression ratio 3 can be used to compress video frame 2.

[0048] Exemplarily, compression ratio 1, compression ratio 2, and compression ratio 3 can decrease in sequence, that is, the compression ratio of the first region in the first video frame is the highest, the compression ratio of other regions in the first video frame is the second highest, and the compression ratio of non-first video frames is the lowest.

[0049] In some embodiments, the above semantic analysis module includes a multi-modal general model. The training data of the multi-modal general model includes picture sample data and text sample data. The picture sample data includes picture samples of various different types of targets / behaviors;

[0050] The semantic information output by the multi-modal general model includes the correspondence relationship among the timestamp, region location information, and text information. The text information is the description information of the specified target / behavior.

[0051] Exemplarily, in order to improve the semantic analysis ability of the semantic analysis model for different types of targets / behaviors in different application scenarios, the semantic analysis model for semantic analysis of the video to be compressed can include a multi-modal general model.

[0052] During the training process of the multi-modal general model, the training data used can include picture samples of various different types of targets / behaviors to ensure that the trained multi-modal general model has good semantic analysis capabilities for general visual application scenarios.

[0053] In addition, during the training process of the multi-modal general model, the training data used can also include text sample data to train the multi-modal general model's ability to convert pictures into text information, such as outputting the description information of the specified target / behavior existing in the picture.

[0054] Correspondingly, the semantic information output by the trained multi-modal general model can include text information in addition to the timestamp and regional location information. This text information is the description information of the specified target / behavior analyzed from the video to be compressed.

[0055] That is, the semantic information output by the multi-modal general model can include the corresponding relationship among the timestamp, regional location information, and text information. Through this semantic information, the timestamp corresponding to the video frame of the video to be compressed with the specified target / behavior, the region of the specified target / behavior in this video frame, and the description information of the specified target / behavior can be determined.

[0056] In an example, the multi-modal general model can include a picture encoder, a text encoder, and a similarity calculation module;

[0057] The training process of the multi-modal general model can include:

[0058] Input the picture sample data and the text sample data in the form of pairs of picture sample data and text sample data into the multi-modal general model; among them, a pair of picture sample data and text sample data includes a picture sample data and the text sample data corresponding to this picture sample data;

[0059] For any pair of picture sample data and text sample data, use the picture encoder to extract the first feature vector and the location information of the second region of this picture sample data, and use the text editor to extract the second feature vector of this text sample data; where the second region is the region of the detected specified target / behavior in this picture sample data;

[0060] Use the similarity calculation module to determine the first similarity between the first feature vector and the second feature vector, and determine the second similarity between the location information of the second region and the location information of the annotated region, and use the loss determined based on this first similarity and second similarity to optimize and update the parameters of the picture editor in the multi-modal general model;

[0061] Iteratively optimize the multi-modal general model using image sample data and text sample data until the preset training end condition is reached.

[0062] Exemplarily, in order to enable the multi-modal general model to detect the region in the image where a specified target / behavior exists and output the description information of the specified target / behavior in the image, the multi-modal general model may include an image editor, a text encoder, and a similarity calculation module. During the training process of the multi-modal general model, the image sample data and the text sample data may be input into the multi-modal general model in the form of pairs of image sample data and text sample data.

[0063] Exemplarily, a pair of image sample data and text sample data includes an image sample data and the text sample data corresponding to the image sample data.

[0064] Exemplarily, for any input pair of image sample data and sample data, on the one hand, the image editor may be used to extract the feature vector of the image sample data (the feature vector for outputting text description information, referred to as the first feature vector in this article), and the position information of the region where the specified target / behavior exists in the image sample data (referred to as the position information of the second region in this article).

[0065] On the other hand, the text editor may be used to extract the feature vector of the text sample data (referred to as the second feature vector in this article).

[0066] Exemplarily, the similarity calculation module may be used to determine the similarity between the first feature vector and the second feature vector (referred to as the first similarity in this article), and the similarity calculation module may be used to determine the similarity between the position information of the second region and the position information of the annotated region (i.e., the position information of the region where the specified target / behavior exists in the pre-annotated image sample data) (referred to as the second similarity in this article).

[0067] Furthermore, the loss may be determined based on the first similarity and the second similarity, and the parameters of the image editor may be optimized and updated based on the determined loss.

[0068] It should be noted that the text encoder and the similarity calculation module in the multi-modal general model may be pre-trained text encoder and similarity calculation module, and during the training process of the multi-modal general model, their parameters may be fixed.

[0069] Exemplarily, iteratively optimize the multi-modal general model using image sample data and text sample data until the preset training end condition is reached. For example, the number of training rounds (one round (epoch) when all training data participates in one training) reaches the preset number of rounds, and / or the multi-modal general model converges, etc.

[0070] Exemplarily, the multimodal general model may further include a text decoder, which may also be a pre-trained text decoder. During the training process of the multimodal general model, the parameters of the text decoder are fixed.

[0071] After completing the training of the multimodal general model in the above manner, for any input video frame in the input video to be compressed, it is possible to use the image encoder to determine whether there is a specified target / behavior in the video frame. If so, use the image encoder to extract the feature vector of the video frame, as well as the position information of the region where the specified target / behavior exists, and use the text decoder to decode the feature vector to obtain the description information of the specified target / behavior. Furthermore, the correspondence relationship among the timestamp (i.e., the timestamp of the video frame), the region position information, and the text information can be output.

[0072] In one example, after performing semantic analysis on the video to be compressed using the trained semantic analysis model to obtain the semantic information of the video to be compressed, it may further include:

[0073] Classify the first region according to the text information of the specified target / behavior to obtain at least two different types of first regions;

[0074] When compressing the video to be compressed, the compression ratios of different types of first regions are different.

[0075] Exemplarily, in order to further improve the flexibility of video compression, for any video to be compressed, when the semantic information of the video to be compressed is obtained using the trained semantic analysis model, the first region can be classified according to the text information of the specified target / behavior to obtain at least two different types of first regions.

[0076] For example, the first region can be divided into a foreground type (such as the specified target / behavior existing in the first region belongs to the foreground) or a background type (such as the specified target / behavior existing in the first region belongs to the background) according to the text information of the specified target / behavior.

[0077] Correspondingly, when compressing the video to be compressed according to the semantic information of the video to be compressed, the compression ratios of different types of first regions can be different.

[0078] For example, the compression ratio of the first region of the foreground type can be higher than that of the first region of the background type.

[0079] In one example, after compressing the video to be compressed according to the semantic information of the video to be compressed, it may further include:

[0080] Associate and store the semantic information of the video to be compressed, the compression ratio used for each region, and the compressed video data.

[0081] Exemplarily, in the case where the video to be compressed is semantically analyzed in the above manner and video compression is performed on the video to be compressed based on the semantic information, the semantic information of the video to be compressed, the compression ratio used for each region, and the compressed video data can also be associated and stored to facilitate subsequent secondary utilization of the semantic information.

[0082] For example, in the application of video target analysis, the target can be initially located based on the semantic information.

[0083] For example, assume that the compression ratio of the first region is higher than that of other regions. Then, based on the compression ratio used for each region, the region where the specified target / behavior exists (i.e., the above-mentioned first region) can be quickly located, and based on the text information in the semantic information, the description information of the specified target / behavior in different first regions can be determined in advance to achieve the initial location of the target.

[0084] Another example, assume that the compression ratios used for different types of first regions are different. Then, based on the compression ratio used for each region, the regions where different types of targets / behaviors exist can also be quickly located, and based on the text information in the semantic information, the description information of the specified target / behavior in different first regions can be determined in advance to achieve the initial location of the target.

[0085] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application, the technical solutions provided in the embodiments of the present application will be described below with specific examples.

[0086] In this embodiment, in order to implement semantic-based video compression in a general visual application scenario, a multi-modal general model can be introduced. During the training process of the multi-modal general model, a large amount of picture sample data and text sample data can be used, and the picture sample data can include picture samples of various different types of targets / behaviors. Thus, for general visual application scenarios, the connection between images and texts can be established using the multi-modal general model.

[0087] Exemplarily, the video to be compressed can be input into the trained multi-modal general model in the form of an image sequence (i.e., a video frame sequence), and the trained multi-modal general model can be used to perform semantic analysis on the video to be compressed and output the semantic information of the video to be compressed (which can be called video semantic information).

[0088] In Example 1, some text descriptions of target / behavior of interest can be predefined, such as "person riding a tricycle", "animal", "fighting", "smoking", etc. These descriptions are input into the multimodal general model in text form (i.e., the above-mentioned text of interest). The multimodal general model outputs the correspondence relationship among the timestamp - location area information - text information (description information of the target / behavior of interest) of the target / behavior of interest that appears in the video of the text of interest.

[0089] In Example 2, the predefined behavior / target of interest can be omitted. Through the training of the multimodal general model, the trained multimodal general model is used to convert the image sequence of the video to be compressed into text information, and the correspondence relationship among the timestamp - location area information - text information is obtained.

[0090] In this embodiment, when the semantic information of the video to be compressed is obtained in the above manner, the video to be compressed can be adaptively and dynamically compressed according to the semantic information of the video to be compressed.

[0091] Exemplarily, for the area where there is a specified target / behavior (such as the above-mentioned target / behavior of interest) (such as the above-mentioned first area), a high compression ratio can be used for compression, and for the area where there is no specified target / behavior, a low compression ratio can be used for compression. By dynamically and adaptively adjusting the compression ratio in time and space, both compression efficiency and quality can be taken into account.

[0092] In this embodiment, when the video to be compressed is semantically analyzed in the above manner and the video is compressed according to the semantic information, the semantic information of the video to be compressed, the compression ratio used for each area, and the compressed video data can also be associated and stored for subsequent secondary utilization of the semantic information. The schematic flow diagram can be as Figure 2 shown.

[0093] The method provided in this application has been described above. Next, the device provided in this application will be described:

[0094] Please refer to Figure 3 which is the structural schematic diagram of a video compression device provided in an embodiment of this application. As Figure 3 shown, the video compression device may include:

[0095] A semantic analysis unit 310 is configured to perform semantic analysis on a video to be compressed by using a trained semantic analysis model, so as to obtain semantic information of the video to be compressed, where the semantic information includes a correspondence relationship between a timestamp and regional location information; the timestamp is a timestamp corresponding to a first video frame, and the first video frame is a video frame in the video to be compressed where a specified target / behavior exists; the regional location information is location information of a first region, and the first region is a region of the specified target / behavior in the first video frame.

[0096] A compression unit 320 is configured to compress the video to be compressed according to the semantic information of the video to be compressed; wherein, a compression ratio of the first region is different from that of other regions.

[0097] In some embodiments, the semantic analysis unit 310 performs semantic analysis on a video to be compressed by using a trained semantic analysis model, so as to obtain semantic information of the video to be compressed, including:

[0098] Inputting the video to be compressed and a predefined text of interest into the semantic analysis model, and using the semantic analysis model to perform semantic analysis on the video to be compressed, so as to obtain semantic information of the video to be compressed; the text of interest includes text description information of the specified target / behavior.

[0099] In some embodiments, the compression unit 320 compresses the video to be compressed according to the semantic information of the video to be compressed, including:

[0100] Compressing the first region by using a first compression ratio and compressing other regions by using a second compression ratio according to the semantic information of the video to be compressed; wherein, the first compression ratio is higher than the second compression ratio.

[0101] In some embodiments, the semantic analysis model includes a multimodal general model, and training data of the multimodal general model includes picture sample data and text sample data, and the picture sample data includes picture samples of various different types of targets / behaviors;

[0102] The semantic information output by the multimodal general model includes a correspondence relationship among a timestamp, regional location information, and text information, and the text information is description information of the specified target / behavior.

[0103] In some embodiments, the multimodal general model includes a picture encoder, a text encoder, and a similarity calculation module;

[0104] The training process of the multimodal general model includes:

[0105] Input the picture sample data and the text sample data in the form of pairs of picture sample data and text sample data into the multi-modal general model; wherein, a pair of picture sample data and text sample data includes one picture sample data and the text sample data corresponding to this picture sample data.

[0106] For any pair of picture sample data and text sample data, use the picture encoder to extract the first feature vector of this picture sample data and the position information of the second region, and use the text editor to extract the second feature vector of this text sample data; wherein, the second region is the region of the detected specified target / behavior in this picture sample data.

[0107] Use the similarity calculation module to determine the first similarity between the first feature vector and the second feature vector, and determine the second similarity between the position information of the second region and the position information of the marked region, and use the loss determined based on this first similarity and second similarity to update the parameters of the multi-modal general model.

[0108] Use the picture sample data and the text sample data to perform iterative optimization on the multi-modal general model until a preset training end condition is reached.

[0109] In some embodiments, as Figure 4 shown, the device may further include:

[0110] A partitioning unit 330, configured to perform type partitioning on the first region according to the text information of the specified target / behavior, to obtain at least two different types of first regions; wherein, when compressing the video to be compressed, the compression ratios of different types of first regions are different.

[0111] In some embodiments, as Figure 5 shown, based on Figure 3 or Figure 4 shown device, the device (taking the optimization of the Figure 3 shown device as an example) may further include:

[0112] A storage unit 340, configured to associatively store the semantic information of the video to be compressed, the compression ratios used by each region, and the compressed video data.

[0113] An embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the video compression method described above.

[0114] Please refer to Figure 6, which is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 may communicate via a system bus 603. And by reading and executing the machine-executable instructions corresponding to the video compression logic in the memory 602, the processor 601 may execute the video compression method described above.

[0115] The memory 602 mentioned herein may be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.

[0116] In some embodiments, a machine-readable storage medium is also provided, such as Figure 6 the memory 602 in, which stores machine-executable instructions. When the machine-executable instructions are executed by a processor, the video compression method described above is implemented. For example, the storage medium may be ROM, RAM, CD-ROM, magnetic tapes, floppy disks, and optical data storage devices, etc.

[0117] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0118] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A video compression method, characterized in that, it includes: Using a trained semantic analysis model, perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed, where the semantic information includes the correspondence between timestamps and regional location information; The timestamp is the timestamp corresponding to the first video frame, and the first video frame is the video frame in the video to be compressed where there is a specified target / behavior; the regional location information is the location information of the first region, and the first region is the region of the specified target / behavior in the first video frame; Compress the video to be compressed according to the semantic information of the video to be compressed; wherein, the compression ratio of the first region is different from that of other regions; Among them, the step of compressing the video to be compressed according to the semantic information of the video to be compressed includes: Compress the first region using a first compression ratio and compress other regions using a second compression ratio according to the semantic information of the video to be compressed; wherein, the first compression ratio is higher than the second compression ratio.

2. The method according to claim 1, characterized in that, The step of using a trained semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed includes: Input the video to be compressed and a predefined text of interest into the semantic analysis model, and use the semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed; the text of interest includes the text description information of the specified target / behavior.

3. The method according to claim 1 or 2, characterized in that, The semantic analysis model includes a multi-modal general model, and the training data of the multi-modal general model includes picture sample data and text sample data, and the picture sample data includes picture samples of various different types of targets / behaviors; The semantic information output by the multi-modal general model includes the correspondence between timestamps, regional location information, and text information, and the text information is the description information of the specified target / behavior.

4. The method according to claim 3, characterized in that, The multi-modal general model includes a picture encoder, a text encoder, and a similarity calculation module; The training process of the multi-modal general model includes: Input the picture sample data and text sample data into the multi-modal general model in the form of pairs of picture sample data and text sample data; wherein, a pair of picture sample data and text sample data includes a picture sample data and the text sample data corresponding to the picture sample data; For any pair of picture sample data and text sample data, use the picture encoder to extract the first feature vector and the location information of the second region of the picture sample data, and use the text encoder to extract the second feature vector of the text sample data; wherein, the second region is the region of the detected specified target / behavior in the picture sample data; The similarity calculation module is used to determine the first similarity between the first feature vector and the second feature vector, and determine the second similarity between the position information of the second region and the position information of the labeled region, and optimize and update the parameters of the image encoder by using the loss determined according to the first similarity and the second similarity; Iteratively optimize the multi-modal general model by using the image sample data and the text sample data until a preset training end condition is reached.

5. The method according to claim 3, wherein, after using the trained semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed, it further includes: According to the text information of the specified target / behavior, classify the first region to obtain at least two different types of first regions; wherein, when compressing the video to be compressed, the compression ratios of different types of first regions are different.

6. The method according to claim 3, wherein, after compressing the video to be compressed according to the semantic information of the video to be compressed, it further includes: Associatively store the semantic information of the video to be compressed, the compression ratio used for each region, and the compressed video data.

7. A video compression device, wherein, comprising: A semantic analysis unit for using a trained semantic analysis model to perform semantic analysis on a video to be compressed to obtain the semantic information of the video to be compressed, where the semantic information includes the correspondence between the time stamp and the region position information; The time stamp is the time stamp corresponding to the first video frame, and the first video frame is a video frame in the video to be compressed where there is a specified target / behavior; the region position information is the position information of the first region, and the first region is the region of the specified target / behavior in the first video frame; A compression unit for compressing the video to be compressed according to the semantic information of the video to be compressed; wherein, the compression ratio of the first region is different from the compression ratios of other regions; wherein, the compression unit compresses the video to be compressed according to the semantic information of the video to be compressed, including: Compressing the first region with a first compression ratio and compressing other regions with a second compression ratio according to the semantic information of the video to be compressed; wherein, the first compression ratio is higher than the second compression ratio.

8. The device according to claim 7, wherein, the semantic analysis unit uses a trained semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed, including: Inputting the video to be compressed and a predefined text of interest into the semantic analysis model, and using the semantic analysis model to perform semantic analysis on the video to be compressed to obtain the semantic information of the video to be compressed; the text of interest includes the text description information of the specified target / behavior; and / or, The semantic analysis model includes a multimodal general model. The training data of the multimodal general model includes picture sample data and text sample data. The picture sample data includes picture samples of multiple different types of targets / behaviors; The semantic information output by the multimodal general model includes the correspondence relationship among the timestamp, regional location information, and text information. The text information is the description information of the specified target / behavior; Among them, the multimodal general model includes a picture encoder, a text encoder, and a similarity calculation module; The training process of the multimodal general model includes: Inputting the picture sample data and the text sample data into the multimodal general model in the form of pairs of picture sample data and text sample data; among them, a pair of picture sample data and text sample data includes a picture sample data and the text sample data corresponding to this picture sample data; For any pair of picture sample data and text sample data, use the picture encoder to extract the first feature vector of this picture sample data and the location information of the second region, and use the text encoder to extract the second feature vector of this text sample data; among them, the second region is the region of the detected specified target / behavior in this picture sample data; Use the similarity calculation module to determine the first similarity between the first feature vector and the second feature vector, and determine the second similarity between the location information of the second region and the location information of the labeled region, and use the loss determined based on this first similarity and second similarity to update the parameters of the multimodal general model; Use the picture sample data and the text sample data to perform iterative optimization on the multimodal general model until a preset training end condition is reached; Among them, the device further includes: A partitioning unit, configured to perform type partitioning on the first region according to the text information of the specified target / behavior, to obtain at least two different types of first regions; among them, when compressing the video to be compressed, the compression ratios of different types of first regions are different; Among them, the device further includes: A storage unit, configured to associate and store the semantic information of the video to be compressed, the compression ratios used by each region, and the compressed video data.

9. An electronic device, characterized in that, it includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the method according to any one of claims 1-6.

10. A machine-readable storage medium, characterized in that, the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Deep learning image compression method based on semantic analysis

    CN110517329A

  • Semantic analysis model training method and device, electronic equipment and storage medium

    CN112560496A