Video Information Generation Method, Apparatus, Electronic Device, and Computer-Readable Medium

By annotating and generating rate distortion curves of the video area of interest, the problem of not being able to evaluate the performance of the encoder and compressed encoder and not evaluating the area of interest in traditional video encoding and codec quality evaluation is solved, and a more accurate video quality evaluation is achieved.

CN114697638BActive Publication Date: 2025-07-25CHONGQING ZHONGXING MICRO ARTIFICIAL INTELLIGENCE CHIP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011604293.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-29
Publication Date
2025-07-25
Estimated Expiration
2040-12-29

AI Technical Summary

Technical Problem

The performance of the encoder and compression encoder cannot be evaluated simultaneously in the traditional video encoding and codec quality evaluation, and the video quality cannot be evaluated through the region of interest, resulting in low evaluation accuracy.

Method used

By annotating the regions of interest in the target video, a set of regions of interest information is generated, and based on this generation rate distortion curve, video quality evaluation information is finally generated.

Benefits of technology

It improves the accuracy of video quality evaluation, solves the problem of failure to highlight the importance of the region of interest in traditional methods, and provides a more accurate video quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114697638B_ABST
    Figure CN114697638B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a video information generation method, apparatus, electronic device, and computer-readable medium. A specific implementation of the method includes: obtaining a target video, where the target video includes at least one region of interest; annotating at least one region of interest included in the target video to generate a region of interest information set; generating at least one rate-distortion curve based on the target video and the region of interest information set; and generating video quality evaluation information based on the at least one rate-distortion curve. This implementation solves the problem of low accuracy in evaluating video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to methods, apparatuses, electronic devices, and computer-readable media for generating video information. Background Art

[0002] In traditional video codec quality evaluation systems, the rate-distortion curve corresponding to the restoration degree of the decoded and reconstructed images and the bit rate of the corresponding compressed bitstream is usually used as the evaluation criterion. Currently, the overall detection method is usually adopted to detect all-frame images in a video.

[0003] However, when using the above detection method, the following technical problems usually exist:

[0004] First, it is impossible to evaluate the performance of the detection encoder and the compression encoder simultaneously in the rate-distortion curve.

[0005] Second, the video quality is not evaluated through the regions of interest included in the video, resulting in a low accuracy of evaluating the video quality. Summary of the Invention

[0006] This part of the content of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. This part of the content of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution. Some embodiments of the present disclosure propose methods, apparatuses, electronic devices, and computer-readable media for generating video information to solve one or more of the technical problems mentioned in the above background art part.

[0007] In a first aspect, some embodiments of the present disclosure provide a method for generating video information, the method including: obtaining a target video, where the target video includes at least one region of interest; annotating at least one region of interest included in the target video to generate a region of interest information set; generating at least one rate-distortion curve based on the target video and the region of interest information set; and generating video quality evaluation information based on the at least one rate-distortion curve.

[0008] In a second aspect, some embodiments of the present disclosure provide a device for generating video information, the device including: an obtaining unit configured to obtain a target video, where the target video includes at least one region of interest; an annotating unit configured to annotate at least one region of interest included in the target video to generate a region of interest information set; a generating unit configured to generate at least one rate-distortion curve based on the target video and the region of interest information set; and an evaluating unit configured to generate video quality evaluation information based on the at least one rate-distortion curve.

[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.

[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.

[0011] The above various embodiments of the present disclosure have the following beneficial effects: Through the video information generation method of some embodiments of the present disclosure, the accuracy of video quality evaluation is effectively improved. Specifically, the reason for the inaccurate relevant video quality evaluation results is that the traditional method evaluates the quality of the full-frame images of the entire video without highlighting the importance of the regions of interest in the video. First, at least one region of interest included in the acquired target video is annotated to generate a region of interest information set. Thus, a reference region of interest set can be provided for generating the rate-distortion curve. Then, based on the target video and the region of interest information set, at least one rate-distortion curve is generated. Thus, data support can be provided for the quality evaluation of the compressed video. Finally, based on at least one rate-distortion curve, video quality evaluation information is generated. Thus, the problem of low accuracy in evaluating video quality is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In combination with the accompanying drawings and with reference to the following specific implementation manners, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0013] Figure 1 is a schematic diagram of an application scenario of a video information generation method according to some embodiments of the present disclosure;

[0014] Figure 2 is a flowchart of some embodiments of the video information generation method according to the present disclosure;

[0015] Figure 3 is a flowchart of some embodiments of the video information generation device according to the present disclosure;

[0016] Figure 4 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0018] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0023] Figure 1 is a schematic diagram of an application scenario of a video information generation method according to some embodiments of the present disclosure.

[0024] In Figure 1 In the application scenario, first, the computing device 101 can obtain the target video 102. Among them, the target video 102 includes at least one region of interest. Then, the computing device 101 can annotate at least one region of interest included in the target video 102 to generate a region of interest information set 103. Then, the computing device 101 can generate at least one rate-distortion curve 104 based on the target video 102 and the region of interest information set 103. Finally, the computing device 101 can generate video quality evaluation information 105 based on at least one rate-distortion curve 104.

[0025] It should be noted that the above computing device 101 can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.

[0026] It should be understood that Figure 1 the number of computing devices in

[0027] Continuing to refer to Figure 2 , a flowchart 200 of some embodiments of the video information generation method according to the present disclosure is shown. The method includes the following steps:

[0028] Step 201, obtaining a target video.

[0029] In some embodiments, the execution subject of the video information generation method (such as Figure 1 the computing device 101 shown) can obtain the target video through a wired connection method or a wireless connection method. Among them, the above target video includes at least one region of interest. In practice, the region of interest can be a region to be processed outlined in a rectangular manner from the image to be processed. Among them, the above target video can be a video containing at least one region of interest. It should be noted that the above wireless connection method can include, but is not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future-developed wireless connection methods.

[0030] Step 202, annotating at least one region of interest included in the target video to generate a region of interest information set.

[0031] In some embodiments, the above execution subject can annotate at least one region of interest included in the target video to generate a region of interest information set. Here, the annotation method can be carried out using a video calibration encoder or in a manual annotation manner. In practice, the above target video can be frame-extracted to obtain a sequence of frame images corresponding to the target video; then, each frame image included in the frame image sequence is manually annotated to obtain a region of interest information set.

[0032] Step 203, generating at least one rate-distortion curve based on the target video and the region of interest information set.

[0033] In some embodiments, the above-mentioned execution entity can analyze and process the target video and the region of interest information set through various methods to generate at least one rate-distortion curve.

[0034] In some alternative implementation manners of some embodiments, the above-mentioned execution entity can generate at least one rate-distortion curve through the following steps:

[0035] First step, use a preset detection encoder to detect at least one region of interest included in the above-mentioned target video to generate a set of detected regions. Among them, the preset detection encoder can be an encoder with a target detection function. In practice, the above-mentioned execution entity can detect at least one region of interest included in the target video through the preset detection encoder, and obtain the set of detected regions of interest as the set of detected regions.

[0036] Second step, label each detected region in the above-mentioned set of detected regions to generate a set of detected region information. In practice, by labeling the above-mentioned set of detected regions, the frame number in the frame sequence corresponding to the target video where each detected region is located, as well as the position information and size information are obtained.

[0037] Third step, input the above-mentioned target video into a preset compression encoder group to generate a set of encoded video groups. Among them, the above-mentioned preset compression encoder can be an encoder that varies values for regions of interest. Among them, the preset compression encoder is a coding method that uses printable ASCII characters to represent characters in various coding formats. In practice, the above-mentioned preset encoder inputs the target video into the preset compression encoder to generate at least 4 encoded videos at different bitrates, that is, a group of encoded videos. For the convenience of comparing the compression effects, the target video can be used to generate the same group of encoded videos under another preset compression encoder to obtain a set of encoded video groups.

[0038] Fourth step, based on the above-mentioned target video and each encoded video group in the above-mentioned set of encoded video groups, generate a group of rate-distortion information to obtain a set of rate-distortion information groups. Through the following formula, the peak signal-to-noise ratio included in the rate-distortion information corresponding to each encoded video in the above-mentioned group of encoded videos is generated:

[0039]

[0040] Among them, for each frame image in the target video, with the upper left corner of the image as the origin, in pixels, the number of columns of pixels in the image array is the abscissa, and the number of rows of pixels is the ordinate to establish an image coordinate system. (i, j) represents the coordinates in the image coordinate system. i represents the abscissa of the above coordinates. j represents the ordinate of the above coordinates. MSE represents the mean square error. PSNR represents the peak signal-to-noise ratio. n represents the number of bits of the pixel value. A represents the above at least one region of interest. card(A) represents the number of regions of interest included in the above at least one region of interest. k represents the serial number of the region of interest in the above at least one region of interest. p k represents the abscissa of the pixel at the upper left corner of the region of interest with serial number k. q k represents the ordinate of the above pixel at the upper left corner. r k represents the abscissa of the pixel at the lower right corner of the region of interest with serial number k. S k represents the ordinate of the above pixel at the lower right corner. I k (i, j) represents the value of the sub-pixel at the coordinate (i, j) on the frame image of the target video corresponding to the region of interest with serial number k in the above set A. J k (i, j) represents the value of the sub-pixel at the coordinate (i, j) on the frame image of the encoded video corresponding to the region of interest with serial number k in the above set A.

[0041] The above image can be represented by different color models, such as representing the image with the three primary colors of red, green, and blue (RGB). Each color on each pixel in the image is called a sub-pixel, and each sub-pixel processes one color channel. In practice, the above sub-pixel can be a sub-pixel that processes the color channel representing the brightness information of the image. Thus, with the bit rate of the encoded video as the abscissa and the corresponding peak signal-to-noise ratio as the ordinate, a rate-distortion information can be obtained. Under the same preset compression encoder, for different bit rates, a set of rate-distortion information can be obtained, that is, a rate-distortion information set.

[0042] Optionally, if the application requirement pays more attention to the compression coding quality within a certain period of time (t, t + Δt) in the target video, the corresponding set of regions of interest at this time is a subset A t,t+Δt of the above set A, replace the set A with its subset A t,t+Δt , and re-perform the above generation process to obtain the rate-distortion information set under this condition.

[0043] Optionally, if the application requirement pays more attention to whether there are regions of interest at certain moments in the target video or the environment where the regions of interest are located (i.e., frame image information), and does not pay attention to the positions of these regions of interest, the set of frame images with regions of interest within a certain period of time before and after encoding can be used as the input to generate the corresponding rate-distortion information set.

[0044] Step 5: Based on the above rate-distortion information set, generate at least one rate-distortion curve. In practice, a set of rate-distortion information in the above rate-distortion information set can be used as a set of coordinate points, and the above set of coordinate points can be connected and interpolated to obtain a rate-distortion curve.

[0045] The above formula and related content are an inventive point of the embodiments of the present disclosure, which solves the second technical problem mentioned in the background art, "The video quality is not evaluated through the region of interest included in the video, resulting in low accuracy of video quality evaluation." The factors that often lead to low accuracy are as follows: The video quality is not evaluated through the region of interest included in the video. If the above factors are solved, the effect of improving the accuracy of video quality evaluation can be achieved. To achieve this effect, the present disclosure introduces the region of interest and the time period of interest to improve the accuracy of video quality evaluation. When the application pays more attention to the information of the region of interest within a certain time period in the target video, the pixel values before and after the coding of the region of interest set within the time period can be used to obtain the peak signal-to-noise ratio, and then generate the rate-distortion curve and quality evaluation information. When the application pays more attention to whether there is a region of interest at a certain moment in the target video or the environment where the region of interest is located (i.e., the frame information), the pixel values before and after the coding of the frame set with a region of interest within the determined time period can be used to obtain the peak signal-to-noise ratio, and then generate the rate-distortion curve and quality evaluation information. The above two methods respectively filter the unconcerned regions or frames, thereby improving the accuracy of video quality evaluation.

[0046] Step 204: Generate video quality evaluation information based on at least one rate-distortion curve.

[0047] In some embodiments, the above execution entity can obtain video quality evaluation information through the rate-distortion curve, such as the change of the peak signal-to-noise ratio at different video bitrates. If two or more rate-distortion curves are obtained, the evaluation information can be obtained by comparison, such as the magnitude of the peak signal-to-noise ratio of the two curves at the same video bitrate, etc.

[0048] In some optional implementation manners of some embodiments, the above execution entity can generate video quality evaluation information through the following steps:

[0049] First step: Generate intersection-over-union information based on at least one region of interest and the detection region set.

[0050] In some embodiments, the above first step may include the following sub-steps:

[0051] The first sub-step: Determine the frame numbers of the regions of interest corresponding to each region of interest in at least one region of interest to obtain at least one region-of-interest frame number.

[0052] The second sub-step: Determine the detection area frame numbers included in the detection area information corresponding to each detection area in the detection area set, and obtain a detection area frame number group.

[0053] The third sub-step: For each region of interest in at least one region of interest, determine whether there is a detection area frame number in the detection area frame number group that matches the region of interest frame number corresponding to the region of interest. Matching can mean being the same.

[0054] The fourth sub-step: In response to the existence of a matching detection area frame number, perform the following processing steps:

[0055] Detect whether there is a region in the region of interest that matches the detection area corresponding to the detection area frame number. Matching can mean that the two regions have an overlapping part.

[0056] In response to the existence of a matching region, determine the area of the matching region, the area of the region of interest, and the area of the detection area. The matching region can be regarded as the intersection of the two regions.

[0057] The fifth sub-step: Determine the area of each matching region in the determined area of the matching region as the matching region area, and obtain a matching region area group.

[0058] The sixth sub-step: Determine the area of each region of interest in the determined area of the region of interest as the region of interest area, and obtain a region of interest area group.

[0059] The seventh sub-step: Determine the area of each detection area in the determined area of the detection area as the detection area area, and obtain a detection area area group.

[0060] The eighth sub-step: Based on the matching region area group, the region of interest area group, and the detection area area group, generate an intersection over union group. The intersection of a region of interest and the corresponding detection area is the matching region, which is measured by the matching region area. The combined part of the region of interest and the corresponding detection area is measured by the sum of the region of interest area and the corresponding detection area area minus the corresponding matching region area. Divide the area of the above intersection part by the area of the combined part as the intersection over union of the above region of interest, and obtain an intersection over union group.

[0061] The ninth sub-step: Based on the intersection over union group and the number of regions of interest included in at least one region of interest, generate intersection over union information. Sum the above intersection over union group, and divide the obtained result by the number of regions of interest included in at least one region of interest to obtain intersection over union information.

[0062] In the second step, the intersection over union (IoU) information is processed for scoring generation to obtain the target detection score value of the detection encoder. The IoU value can be directly used as the target detection score value of the detection encoder.

[0063] In the third step, based on at least one rate-distortion curve and the target detection score value, video quality evaluation information is generated. The target detection performance of the detection encoder is characterized by the above-mentioned score value, and the compression encoder performs encoding based on the detection region. Therefore, the detection performance of the detection encoder will affect the quality of the compressed video. The above-mentioned score value and the rate-distortion curve can be used as video quality evaluation information.

[0064] Optionally, the above-mentioned execution entity can generate video quality evaluation information through the following steps:

[0065] In the first step, at least one region of interest and a set of detection regions are used to generate intersection over union (IoU) information.

[0066] Within the interesting time period (t, t + Δt) of the target video, the set of frame numbers S corresponding to the regions of interest can be found, and the set of frame numbers T corresponding to the detection regions can also be found. All the largest consecutive subsets S1, S2,..., Sn of frame numbers are selected from S. Here, n represents the number of the largest consecutive subsets of frame numbers in S. The largest subset here refers to the subset with the largest number of elements under the same constraints. For each frame number in S1, if it also exists in T, it is put into the empty set I1. The same comparison process is also done for S2,..., Sn. Finally, the sets I1, I2,..., In are obtained.

[0067] The sets U1, U2,..., Un are respectively initialized as the sets I1, I2,..., In. Find the smallest frame number I1min and the largest frame number I1max in I1. Starting from I1min, if the frame number I1min - 1 is in T, add this frame number to U1. Then, successively judge I1min - 2, I1min - 3,... until a certain frame number is not in T, and stop this addition. Then, starting from I1max, if the frame number I1max + 1 is in T, add this frame number to U1. Then, successively judge I1max + 2, I1max + 3,... until a certain frame number is not in T, and stop the addition. The same process is also done for I2,..., In. Finally, the updated set sequence U1, U2,..., Un is obtained. Then, calculate Card(I1) / Card(U1) + Card(I2) / Card(U2) +... + Card(In) / Card(Un), and divide the above result by Card(S). The result is used as the intersection over union (IoU) information. Here, Card(S) represents the number of elements in the set S.

[0068] Second, perform scoring generation processing on the above-mentioned intersection over union information to obtain the target detection score value of the preset detection encoder. The above-mentioned intersection over union information can be used as the target detection score value of the above-mentioned detection encoder.

[0069] Third, based on the above-mentioned at least one rate-distortion curve and the above-mentioned target detection score value, generate video quality evaluation information. The target detection performance of the detection encoder is characterized by the above-mentioned score value, and the compression encoder performs encoding based on the detection region. Therefore, the detection performance of the detection encoder will affect the quality of the compressed video. The above-mentioned score value and rate-distortion curve can be used as video quality evaluation information.

[0070] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the video information generation method of some embodiments of the present disclosure, the accuracy of video quality evaluation is effectively improved. Specifically, the reason for the inaccurate relevant video quality evaluation results is that the traditional method evaluates the quality of the full-frame images of the entire video without highlighting the importance of the regions of interest in the video. First, annotate at least one region of interest included in the obtained target video to generate a region of interest information set. Thus, a reference region of interest set can be provided for generating the rate-distortion curve. Then, based on the target video and the region of interest information set, generate at least one rate-distortion curve. Thus, data support can be provided for the quality evaluation of the compressed video. Finally, based on the at least one rate-distortion curve, generate video quality evaluation information. Thus, the problem of low accuracy in evaluating the video quality is solved.

[0071] Further reference Figure 3 , as an implementation of the above-mentioned method, the present disclosure provides some embodiments of a video information generation device. These device embodiments correspond to Figure 2 the above-mentioned method embodiments shown, and the device can be specifically applied to various electronic devices.

[0072] As Figure 3 shown, a video quality evaluation device 300 of some embodiments includes: an acquisition unit 301, an annotation unit 302, a generation unit 303, and an evaluation unit 304. Among them, the acquisition unit 301 is configured to acquire a target video, where the above-mentioned target video includes at least one region of interest; the annotation unit 302 is configured to annotate at least one region of interest included in the above-mentioned target video to generate a region of interest information set; the generation unit 303 is configured to generate at least one rate-distortion curve based on the above-mentioned target video and the above-mentioned region of interest information set; and the evaluation unit 304 is configured to generate video quality evaluation information based on the above-mentioned at least one rate-distortion curve.

[0073] It can be understood that the various units recorded in the device 300 correspond to the reference Figure 2corresponds to each step in the described method. Thus, the operations, features, and beneficial effects described above for the method also apply to the apparatus 300 and the units included therein, and will not be repeated here.

[0074] Reference is now made to Figure 4 , which shows a schematic structural diagram of an electronic device (such as the computing device 101 in Figure 1 ) 400 suitable for use in implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0075] As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0076] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 4 shows an electronic device 400 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 4 Each block shown in

[0077] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from a network through a communication device 409, or installed from a storage device 408, or installed from a ROM 402. When the computer program is executed by a processing device 401, the above functions defined in the methods of some embodiments of the present disclosure are performed.

[0078] It should be noted that the computer-readable medium described in some embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0079] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0080] The computer-readable medium described above may be included in the above-described electronic device; or may exist separately without being assembled into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a target video, where the target video includes at least one region of interest; annotate at least one region of interest included in the target video to generate a set of region-of-interest information; generate at least one rate-distortion curve based on the target video and the set of region-of-interest information; and generate video quality evaluation information based on the at least one rate-distortion curve.

[0081] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0082] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0083] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a labeling unit, a generation unit, and an evaluation unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "acquiring a target video, where the above target video includes at least one region of interest".

[0084] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0085] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.

Claims

1. A method for generating video information, comprising: Obtaining a target video, wherein the target video includes at least one region of interest; Annotating at least one region of interest included in the target video to generate a region of interest information set; Detecting at least one region of interest included in the target video by using a preset detection encoder to generate a detection region set; Annotating each detection region in the detection region set to generate a detection region information set; Inputting the target video into a preset compression encoder group to generate a set of encoded video groups; Generating a rate-distortion information group based on the target video and each encoded video group in the set of encoded video groups to obtain a set of rate-distortion information groups; Generating at least one rate-distortion curve based on the set of rate-distortion information groups; Generating intersection over union information based on the at least one region of interest and the detection region set; Performing a scoring generation process on the intersection over union information to obtain a target detection score value of the detection encoder; Generating video quality evaluation information based on the at least one rate-distortion curve and the target detection score value; Wherein the rate-distortion information includes peak signal-to-noise ratio, and the generating a rate-distortion information group based on the target video and each encoded video group in the set of encoded video groups includes: In response to determining that the application pays more attention to the information of the region of interest in a certain time period of the target video, generating a peak signal-to-noise ratio by using the pixel values before and after encoding of the region of interest information set in the time period; In response to determining that the application pays more attention to whether there is a region of interest at a certain moment in the target video or pays attention to the frame map information where the region of interest is located, generating a peak signal-to-noise ratio by using the pixel values before and after encoding of the frame map set with a region of interest in a determined time period.

2. The method according to claim 1, wherein, The generating a rate-distortion information group based on the target video and each encoded video group in the set of encoded video groups includes: Generating the peak signal-to-noise ratio included in the rate-distortion information corresponding to each encoded video in the encoded video group through the following formula: For each frame image in the target video, taking the upper left corner of the image as the origin, with pixels as the unit, the number of columns of pixels in the image array as the abscissa, and the number of rows of pixels as the ordinate, an image coordinate system is established. (i, j) represents the coordinates in the image coordinate system, i represents the abscissa of the coordinates, j represents the ordinate of the coordinates, MSE represents the mean square error, PSNR represents the peak signal-to-noise ratio, n represents the number of bits of the pixel value, A represents the at least one region of interest, card(A) represents the number of regions of interest included in the at least one region of interest, k represents the serial number of the region of interest in the at least one region of interest, p k represents the abscissa of the pixel at the upper left corner of the region of interest with the serial number k, q k represents the ordinate of the pixel at the upper left corner, r k represents the abscissa of the pixel at the lower right corner of the region of interest with the serial number k, s k represents the ordinate of the pixel at the lower right corner, I k (i, j) represents the value of the sub-pixel at the coordinates (i, j) on the frame image of the target video corresponding to the region of interest with the serial number k in the set A, J k (i, j) represents the value of the sub-pixel at the coordinates (i, j) on the frame image of the encoded video corresponding to the region of interest with the serial number k in the set A.

3. The method according to claim 2, wherein The region of interest information includes: region of interest frame number, and the detection region information includes: detection region frame number; and The generating intersection over union information based on the at least one region of interest and the detection region set includes: Determining the region of interest frame number included in the region of interest information corresponding to each region of interest in the at least one region of interest to obtain at least one region of interest frame number; Determining the detection region frame number included in the detection region information corresponding to each detection region in the detection region set to obtain a detection region frame number group; For each region of interest in the at least one region of interest, determining whether there is a detection region frame number in the detection region frame number group that matches the region of interest frame number corresponding to the region of interest; In response to the existence of a matching detection region frame number, performing the following processing steps: Detecting whether there is a region in the region of interest that matches the detection region corresponding to the detection region frame number; In response to the existence of a matching region, determining the area of the matching region, the area of the region of interest, and the area of the detection region.

4. The method according to claim 3, wherein Generating intersection over union (IoU) information based on the at least one region of interest and the set of detected regions further includes: Determining the area of each of the determined matching regions in the area of the determined matching regions as the matching region area to obtain a set of matching region areas; Determining the area of each of the determined regions of interest in the area of the determined regions of interest as the region of interest area to obtain a set of region of interest areas; Determining the area of each of the determined detected regions in the area of the determined detected regions as the detected region area to obtain a set of detected region areas; Generating a set of IoUs based on the set of matching region areas, the set of region of interest areas, and the set of detected region areas; Generating IoU information based on the set of IoUs and the number of regions of interest included in the at least one region of interest.

5. A video information generation device, comprising: An acquisition unit configured to acquire a target video, where the target video includes at least one region of interest; A labeling unit that labels the at least one region of interest included in the target video to generate a set of region of interest information; A generation unit configured to use a preset detection encoder to detect the at least one region of interest included in the target video to generate a set of detected regions; label each detected region in the set of detected regions to generate a set of detected region information; input the target video into a preset compression encoder group to generate a set of encoded video groups; generate a set of rate-distortion information groups based on the target video and each encoded video group in the set of encoded video groups to obtain a set of rate-distortion information groups; and generate at least one rate-distortion curve based on the set of rate-distortion information groups; An evaluation unit configured to generate IoU information based on the at least one region of interest and the set of detected regions; perform a scoring generation process on the IoU information to obtain a target detection score value of the detection encoder; and generate video quality evaluation information based on the at least one rate-distortion curve and the target detection score value; where the rate-distortion information includes peak signal-to-noise ratio (PSNR), and generating a set of rate-distortion information groups based on the target video and each encoded video group in the set of encoded video groups includes: in response to determining that the application is more concerned about the information of the region of interest in a certain time period in the target video, generating PSNR using the pixel values before and after encoding of the set of region of interest information in the time period; and in response to determining that the application is more concerned about whether there is a region of interest at a certain moment in the target video or about the frame information where the region of interest is located, generating PSNR using the pixel values before and after encoding of the set of frame maps with regions of interest in the determined time period.

6. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-4.

7. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Method and system for objectively evaluating video coding performance

    CN101895788A

  • Image quality evaluation method and device, electronic equipment and computer readable storage medium

    CN110858394A