Video detection method and device
By obtaining the video frame to be detected, downsampling and upsampling, and determining its actual resolution, the problems of data processing complexity and high computing cost in the prior art are solved, and efficient video detection is achieved.
Patent Information
- Application Number
- CN202510297301.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-03
AI Technical Summary
Existing video detection methods require the collection and processing of large amounts of data, resulting in high complexity and computational costs, making it difficult to effectively determine the actual resolution of video frames.
By obtaining the video frame to be detected, multiple downsampled video frames of different resolutions are obtained, and upsampled them to reconstruct the original resolution. The actual resolution is determined based on the amount of information of the reconstructed video frame and the amount of information of the video frame to be detected.
It realizes that only determines the actual resolution of the video frame to be detected, avoids a large amount of data collection and processing, reduces the complexity and data requirements of video detection, and reduces the computational cost.
Smart Images

Figure CN120091160A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a video detection method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of digital video technology, the production and dissemination of video content have become increasingly convenient. However, in terms of the authenticity and integrity of video content, the production and dissemination of pseudo-high-definition videos are becoming increasingly serious. A pseudo-high-definition video is a video that disguises a non-high-definition video as a high-definition video through technical means such as increasing the resolution, adjusting the bit rate, and frame rate. Such disguises not only mislead viewers but may also be used for unfair competition.
[0003] To address this situation, it is necessary to determine the actual resolution of video frames. However, existing video detection methods require collecting and processing a large amount of data, resulting in high complexity and data requirements, and increasing the computational cost.
[0004] It should be noted that the above content is not necessarily prior art and does not limit the scope of patent protection of the present application. Summary of the Invention
[0005] Embodiments of the present application provide a video detection method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above technical problems.
[0006] One aspect of the embodiments of the present application provides a video detection method, the method including: Obtaining a video frame to be detected, where the video frame to be detected carries information about the original resolution; Based on the video frame to be detected, obtaining multiple downsampled video frames with different resolutions; Upsampling the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution; and Determining the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected.
[0007] Optionally, obtaining multiple downsampled video frames based on the video frame to be detected includes: Dividing a preset resolution interval to obtain multiple resolutions; where each of the multiple resolutions is different; Downsampling the video frame to be detected to obtain multiple downsampled video frames corresponding to the multiple resolutions.
[0008] Optionally, determining the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected includes: Obtaining a plurality of information amount difference values based on the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected; Obtaining an information amount difference change curve according to the plurality of information amount difference values; Determining the actual resolution of the video frame to be detected according to the inflection point of the information amount difference change curve.
[0009] Optionally, the information amount includes pixel information; obtaining a plurality of information amount difference values based on the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected includes: Obtaining a plurality of pixel difference values by comparing each pixel in the target reconstructed video frame with the corresponding pixel in the video frame to be detected; wherein the target reconstructed video frame is any one of the plurality of reconstructed video frames; and Obtaining the information amount difference value between the target reconstructed video frame and the video frame to be detected based on the plurality of pixel difference values.
[0010] Optionally, determining the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected further includes: Obtaining the similarity between each of the reconstructed video frames and the video frame to be detected; Obtaining a similarity change curve according to the similarity between each of the reconstructed video frames and the video frame to be detected; and Determining the actual resolution of the video frame to be detected according to the inflection point of the similarity change curve.
[0011] Optionally, upsampling the plurality of downsampled video frames to obtain a plurality of reconstructed video frames corresponding to the original resolution, including: Obtaining a plurality of new pixel values of each of the downsampled video frames based on the original resolution and the plurality of pixel values of each of the downsampled video frames; Obtaining the plurality of reconstructed video frames based on the plurality of pixel values of each of the downsampled video frames and the corresponding plurality of new pixel values.
[0012] Another aspect of the embodiments of the present application provides a video detection device, the device includes: An acquisition model, configured to acquire a video frame to be detected, where the video frame to be detected carries information of the original resolution; A downsampling module, configured to acquire a plurality of downsampled video frames with different resolutions based on the video frame to be detected; An upsampling module that upsamples the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution; and A determination module configured to determine the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected.
[0013] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0014] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0015] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0016] The embodiments of the present application adopting the above technical solutions may include the following advantages: When the resolution of the downsampled video frame is different from the actual resolution of the video frame to be detected, there is a lack of information amount in the downsampled video frame compared with the video frame to be detected, and the corresponding reconstructed video frame will amplify the information amount difference from the video frame to be detected. Thus, in the video detection process, according to the information amount of each reconstructed video frame and the information amount of the video frame to be detected, the actual resolution of the video frame to be detected is determined, realizing the determination of the actual resolution of the video frame to be detected only by the video frame to be detected, avoiding a large amount of data collection and processing, thereby reducing the complexity and data requirements of video detection, and at the same time reducing the computing cost. Description of the Drawings
[0017] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0018] Figure 1 Schematically shows an operating environment diagram of the video detection method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the video detection method according to Embodiment 1 of the present application; Figure 3 Schematically shows according to Figure 2 The sub-step flowchart of step S202 in Figure 4 Schematically shows according to Figure 2 The sub-step flowchart of step S204 in Figure 5 Schematically shows according to Figure 2 The sub-step flowchart of step S206 in Figure 6 Schematically shows according to Figure 5 The sub-step flowchart of step S500 in Figure 7 Schematically shows according to Figure 2 The sub-step flowchart of step S206 in Figure 8 Schematically shows the block diagram of the video detection device according to the second embodiment of the present application; and Figure 9 Schematically shows the schematic diagram of the hardware architecture of the computer device according to the third embodiment of the present application. Detailed implementation manners
[0019] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0020] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0021] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.
[0022] To facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: The applicant has learned that the actual resolution of video frames can be determined through machine learning and deep learning methods. However, a large amount of video data needs to be collected and processed during the training of the model, which leads to high complexity and computational cost of this method.
[0023] For this reason, the embodiments of the present application provide a video detection technical solution. In this technical solution, a large amount of data collection and processing are avoided, thereby reducing the complexity and data requirements of video detection, and at the same time reducing the computational cost. See the following for details.
[0024] Finally, for the convenience of understanding, an exemplary operating environment is provided below.
[0025] As Figure 1 shown, the operating environment diagram includes: a service platform 2, and clients (4A, 4B,..., 4N).
[0026] The service platform 2 can be connected to the clients (4A, 4B,..., 4N) through a network.
[0027] The service platform 2 can be a single server, a server cluster, or a cloud computing service center.
[0028] The service platform 2 can determine the actual resolution of the videos uploaded by the clients, etc.
[0029] The service platform 2 can be located in a data center such as a single location, or distributed in different geographical locations (for example, in multiple locations). The service platform 2 can provide services via a network. The network includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0030] Clients (4A, 4B, …, 4N) can be configured to upload video manuscripts to the service platform 2. The clients (4A, 4B, …, 4N) can include electronic devices with or external to a display panel, such as mobile devices, tablet devices, laptop computers, workstations, virtual reality devices, gaming devices, digital streaming devices, vehicle terminals, smart TVs, set-top boxes, etc., and can also include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load the virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices.
[0031] The clients (4A, 4B, …, 4N) can be associated with one or more users. A single user can also use one or more of the clients (4A, 4B, …, 4N) to access the service platform 2. The clients (4A, 4B, …, 4N) can travel to various locations and use different networks to access the service platform 2.
[0032] The following takes the service platform 2 as the execution entity and introduces the technical solutions of this application through multiple embodiments. It should be noted that these embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments described herein.
[0033] Embodiment 1 Figure 2 A flowchart of a video detection method according to Embodiment 1 of this application is schematically shown.
[0034] As Figure 2 shown, the video detection method can include steps S200~S206, where: Step S200, obtain a video frame to be detected, and the video frame to be detected carries information of the original resolution.
[0035] Step S202, based on the video frame to be detected, obtain multiple downsampled video frames with different resolutions.
[0036] Step S204, upsample the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution.
[0037] Step S206, determine the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected.
[0038] In the video detection method provided in this embodiment, when the resolution of the downsampled video frame is different from the actual resolution of the video frame to be detected, there is a lack of information in the downsampled video frame compared to the video frame to be detected, and the corresponding reconstructed video frame will amplify the information difference from the video frame to be detected. Thus, during the video detection process, based on the information content of each reconstructed video frame and the information content of the video frame to be detected, the actual resolution of the video frame to be detected is determined, achieving the determination of the actual resolution of the video frame to be detected only through the video frame to be detected, avoiding the collection and processing of a large amount of data, thereby reducing the complexity and data requirements of video detection, and at the same time reducing the computational cost.
[0039] The following will elaborate in detail on each step in steps S200 to S206 and optional other steps in combination with Figure 2 , and elaborate on each step in steps S200 to S206 and optional other steps in detail.
[0040] Step S200 , obtain the video frame to be detected, and the video frame to be detected carries information about the original resolution.
[0041] The video frame to be detected can be one or more video frames extracted from the video to be detected. The video to be detected can be a video uploaded by a user to the service platform 2 through a client, or a video stored in the service platform 2. Since the resolution of a video file can be modified by changing the resolution parameters, there are videos obtained by magnifying a low-resolution video. The original resolution of such videos is not the actual resolution. Specifically, the image quality has not been substantially improved. Thus, it is necessary to detect the video to determine its actual resolution to identify such videos.
[0042] Step S202 , based on the video frame to be detected, obtain multiple downsampled video frames with different resolutions.
[0043] The downsampled video is obtained by downsampling the video frame to be detected. Downsampling is the process of reducing the resolution of an image, and its purpose is to reduce the number of pixels in the image. During the downsampling process, the position of each pixel will correspond to a smaller image area. Since the downsampling process necessarily involves the loss of pixels, there will be a loss of information during the downsampling process. The greater the difference between the actual resolution of the video frame to be detected and the resolution of the downsampled video frame, the more information is lost.
[0044] The following will exemplarily introduce the specific process of obtaining multiple downsampled video frames with different resolutions.
[0045] In an optional embodiment, as Figure 3 shown, step S202 may include: Step S300, divide a resolution range based on the original resolution to obtain multiple resolutions; wherein, each of the multiple resolutions is different.
[0046] Step S302, downsample the video frame to be detected to obtain multiple downsampled video frames corresponding to the multiple resolutions.
[0047] A resolution range can be determined according to the original resolution, and then the resolution range is divided, and multiple resolutions are selected from this range. The image interpolation algorithm can be used to downsample the video frame to be detected. The image interpolation algorithm is used to estimate the pixel values in the new image according to the existing pixel values when changing the image resolution. In some embodiments, taking the video frame frame_1080p with an original resolution of 1080P as an example, a resolution can be taken every 40P from 240P to 1040P, for a total of 21 resolutions, and these resolutions are represented by res_i, where 1 <= i <= 21. Then, frame_1080p is downsampled at different resolutions to obtain multiple downsampled video frames, that is, frame_1080p_down_i = F(frame_1080p, res_i). Wherein, frame_1080p_down_i represents the downsampled video frame with a resolution of res_i, and F(x) represents the image interpolation algorithm.
[0048] In this embodiment, since there are differences in the amount of information in the downsampled video frames with different resolutions, the change in the amount of information when the video frame to be detected is downsampled to different resolutions can be inferred through multiple downsampled video frames with different resolutions.
[0049] Step S204 , upsample the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution.
[0050] Upsampling refers to the process of converting a low-resolution image into a high-resolution image. Since downsampling will lose image information, when upsampling back to the original resolution, the restored image quality will not be the same as the original image. In addition, because upsampling involves the generation of new pixels, the difference in the amount of information between the generated reconstructed video frame and the video frame to be detected will be further amplified. In some embodiments, frame_1080p_down_i (the downsampled video frame with a resolution of res_i) is upsampled back to the original resolution (1080P), that is, frame_1080p_down_i_up = F(frame_1080p_down_i, 1080p). Wherein, frame_1080p_down_i_up represents the reconstructed video frame corresponding to frame_1080p_down_i.
[0051] Regarding the generation of reconstructed video frames, in an alternative embodiment, as Figure 4 shown, step S204 may include: Step S400, obtaining a plurality of new pixel values for each of the downsampled video frames based on the original resolution and the plurality of pixel values of each of the downsampled video frames; Step S402, obtaining the plurality of reconstructed video frames based on the plurality of pixel values and the corresponding plurality of new pixel values of each of the downsampled video frames.
[0052] The downsampled video frames can be upsampled by bilinear interpolation and bicubic interpolation. Bilinear interpolation and bicubic interpolation are interpolation methods. Bilinear interpolation is used to estimate the value of a pixel by using the weighted average of the four neighboring pixels around the position of the target pixel. Bicubic interpolation uses 16 neighboring pixels (4x4 grid) around the target pixel for interpolation. For example, the resolution of one of the downsampled video frames is 240P, and after upsampling, it is restored to the original video frame of 1080P. The interpolation algorithm will estimate the missing pixel values in the 1080P image based on the pixel information in the existing 240P image. For the target pixel (x,y), assuming it is located within a 2x2 grid, the four neighboring pixels are (x1,y1), (x2,y1), (x1,y2), (x2,y2) respectively, then the value of the target pixel (x,y) is the weighted average of the values of these four neighboring pixels. First, interpolate in the horizontal direction using the bilinear interpolation formula: (x,y1)=(x2 - x) / (x2 - x1)·(x1,y1)+(x - x1) / (x2 - x1)·(x2,y1); (x,y2)=(x2 - x) / (x2 - x1)·(x1,y2)+(x - x1) / (x2 - x1)·(x2,y2). Then interpolate in the vertical direction: (x,y)=(y2 - y) / (y2 - y1)·(x,y1)+(y - y1) / (y2 - y1)·(x,y2).
[0053] In this embodiment, by generating a plurality of reconstructed video frames, the information quantity difference between the plurality of downsampled video frames and the video frame to be detected is amplified, so as to facilitate comparison with the information quantity of the video frame to be detected.
[0054] Step S206 , determining the actual resolution of the video frame to be detected according to the information quantity of each of the reconstructed video frames and the information quantity of the video frame to be detected.
[0055] Since multiple reconstructed video frames are obtained by downsampling the video frame to be detected to different resolutions and then upsampling back to the original resolution, there are different differences in the amount of information between each reconstructed video frame and the video frame to be detected. For example, the meta-information of a video frame A to be detected indicates that its original resolution is 1080P, but its actual resolution is 540P, that is, it is a pseudo-high-definition (pseudo-1080P) video. To detect whether the video frame A to be detected is a pseudo-1080P video, the video frame to be detected can be downsampled to any resolution in the resolution range [540P, 1080P), and then upsampled back to 1080P. The amount of information of the obtained reconstructed video frame is basically the same as that of the video frame to be detected. When the video frame to be detected is downsampled to any resolution in the range (0P, 540P) and then upsampled back to 1080P, the difference in the amount of information between the reconstructed video frame and the video frame to be detected will increase as the downsampling resolution decreases. Therefore, the actual resolution of the video frame to be detected can be determined by the change in the amount of information, so as to identify whether the video frame A to be detected is a pseudo-1080P video.
[0056] The following will exemplarily introduce the specific process of determining the actual resolution of the video frame to be detected.
[0057] Determining the actual resolution by calculating the information amount difference value. In an optional embodiment, as Figure 5 shown, step S206 may include: Step S500, based on the amount of information of each of the reconstructed video frames and the amount of information of the video frame to be detected, obtain a plurality of information amount difference values.
[0058] Step S502, according to the plurality of information amount difference values, obtain an information amount difference change curve.
[0059] Step S504, according to the inflection point of the information amount difference change curve, determine the actual resolution of the video frame to be detected.
[0060] The information quantity difference value is used to represent the difference between the information quantity of the reconstructed video frame and that of the video frame to be detected. The information quantity difference change curve can be a curve graph with the resolution of the downsampled video frame corresponding to the reconstructed video frame as the horizontal axis and the information quantity difference value as the vertical axis. The inflection point of this curve is the position where the change in information quantity is the most obvious, corresponding to the actual resolution of the video frame. For example, for a video frame to be detected with an original resolution of 1080P and an actual resolution of 540P, the information quantity difference value will continuously decrease as the downsampling resolution increases, and an obvious inflection point will start to appear near the 540P resolution and then tend to be stable. Similarly, for a video frame to be detected with an original resolution of 1080P and an actual resolution of 360P, its inflection point appears near 360P. For a video frame to be detected with an original resolution of 1080P and an actual resolution of 1080P, no inflection point will appear.
[0061] In this embodiment, the trend of information quantity loss changing with the downsampling resolution is described by the information quantity difference change curve, and then the actual resolution of the video frame to be detected is visually confirmed through the inflection point of this curve.
[0062] Regarding the calculation of the information quantity difference value, in an alternative embodiment, as Figure 6 shown, step S500 includes: Step S600, by comparing each pixel in the target reconstructed video frame with the corresponding pixel in the video frame to be detected, obtaining a plurality of pixel difference values; wherein, the target reconstructed video frame is any one of the plurality of reconstructed video frames.
[0063] Step S602, based on the plurality of pixel difference values, obtaining the information quantity difference value between the target reconstructed video frame and the video frame to be detected.
[0064] The information quantity difference value can be obtained through pixel difference averaging or mean square error. Pixel difference averaging is to directly calculate the absolute difference of each pixel between the reconstructed video frame and the video frame to be detected, and then take the mean value. The mean square error measures the overall information loss degree by calculating the sum of the squares of the pixel errors. For example, when calculating the information quantity difference value between frame_1080p_down_i_up (reconstructed video frame) and frame_1080p (video frame to be detected), the information quantity difference value diff_i = G(frame_1080p_down_i_up, frame_1080p), where G(x) is the function for calculating the information quantity difference value, that is, the function corresponding to pixel difference averaging or mean square error.
[0065] In this embodiment, the difference between the reconstructed video frame and the video frame to be detected is quantified by the information quantity difference value, so as to facilitate determining the actual resolution of the video frame to be detected subsequently.
[0066] Determine the actual resolution by calculating the similarity. In an alternative embodiment, as Figure 7 shown, step S206 may further include: Step S700, obtain the similarity between each of the reconstructed video frames and the video frame to be detected.
[0067] Step S702, obtain a similarity change curve according to the similarity between each of the reconstructed video frames and the video frame to be detected.
[0068] Step S704, determine the actual resolution of the video frame to be detected according to the inflection point of the similarity change curve.
[0069] In some embodiments, the similarity between the reconstructed video frame and the video frame to be detected may be calculated by negative PSNR and negative SSIM. PSNR is an image quality evaluation metric that measures the signal-to-noise ratio between two images and is used to evaluate the quality difference between images. The smaller the negative PSNR value (i.e., the larger the PSNR), the more similar the images and the smaller the difference. SSIM is a metric for measuring the structural similarity between two images and can better reflect the human perception of image quality than simple pixel differences. The calculation of SSIM consists of three main parts: luminance, contrast, and structure. The smaller the negative SSIM value, the more similar the two images. The similarity change curve may be a curve graph with the resolution of the downsampled video frame corresponding to the reconstructed video frame as the horizontal axis and the similarity as the vertical axis. The inflection point of this curve is the position where the similarity changes most significantly and corresponds to the actual resolution of the video frame.
[0070] In this embodiment, the trend of the information loss changing with the downsampling resolution is described by the similarity change curve, and then the actual resolution of the video frame to be detected is visually confirmed through the inflection point of this curve.
[0071] To make the present application easier to understand, the following provides an exemplary application.
[0072] In this exemplary application, take the video frame frame_1080p with an original resolution of 1080P as an example.
[0073] S1. Divide the resolution range, take a resolution every 40P from 240P to 1040P, a total of 21 resolutions, and represent this resolution by res_i, where res_i, 1 <= i <= 21.
[0074] S2. Downsample frame_1080p to obtain multiple downsampled video frames corresponding to res_i, that is, the downsampled video frame frame_1080p_down_i = F(frame_1080p, res_i), where F(x) represents an image interpolation algorithm, such as bilinear interpolation, bicubic interpolation, etc.
[0075] S3. Upsample frame_1080p_down_i back to the original resolution (1080P) to obtain multiple reconstructed video frames, that is, the reconstructed video frame frame_1080p_down_i_up = F(frame_1080p_down_i, 1080p).
[0076] S4. Calculate the information quantity difference value or similarity between frame_1080p_down_i_up and frame_1080p, that is, the information quantity difference value or similarity diff_i = G(frame_1080p_down_i_up, frame_1080p), where G(x) is a function for calculating the information quantity difference value or similarity, such as a function of taking the average of the pixel differences for calculating the information quantity difference value, or functions such as negative PSNR and negative SSIM for calculating similarity.
[0077] S5. According to the information quantity difference value or similarity diff_i, draw a curve of the change in the information quantity difference or a curve of the change in similarity, and determine the actual resolution of the video frame to be detected according to the inflection point of the curve.
[0078] In this exemplary application, only by using the video frame to be detected as the basic data to determine its actual resolution, it is possible to avoid collecting and processing a large amount of data, thereby reducing the complexity and data requirements of video detection, and at the same time reducing the computing cost.
[0079] Embodiment 2 Figure 8 Schematically shows a block diagram of a video detection device according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 8 shown, the device 800 may include: an acquisition module 810, a downsampling module 820, an upsampling module 830, and a determination module 840, where: The acquisition module 810 is configured to acquire a video frame to be detected, and the video frame to be detected carries information about the original resolution; The downsampling module 820 is configured to acquire a plurality of downsampled video frames based on the video frame to be detected; The upsampling module 830 upsamples the plurality of downsampled video frames to obtain a plurality of reconstructed video frames corresponding to the original resolution; and A determination module 840, configured to determine an actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected.
[0080] As an optional embodiment, the downsampling module 820 is further configured to: Divide a resolution range based on the original resolution to obtain a plurality of resolutions; wherein, each of the plurality of resolutions is different; Downsample the video frame to be detected to obtain a plurality of downsampled video frames corresponding to the plurality of resolutions.
[0081] As an optional embodiment, the determination module 840 is further configured to: Obtain a plurality of information amount difference values based on the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected; Obtain an information amount difference change curve according to the plurality of information amount difference values; Determine the actual resolution of the video frame to be detected according to an inflection point of the information amount difference change curve.
[0082] As an optional embodiment, the determination module 840 is further configured to: Obtain a plurality of pixel difference values by comparing each pixel in a target reconstructed video frame with a corresponding pixel in the video frame to be detected; wherein, the target reconstructed video frame is any one of the plurality of reconstructed video frames; and Obtain an information amount difference value between the target reconstructed video frame and the video frame to be detected based on the plurality of pixel difference values.
[0083] As an optional embodiment, the determination module 840 is further configured to: Obtain a similarity between each of the reconstructed video frames and the video frame to be detected; Obtain a similarity change curve according to the similarity between each of the reconstructed video frames and the video frame to be detected; and Determine the actual resolution of the video frame to be detected according to an inflection point of the similarity change curve.
[0084] As an optional embodiment, the upsampling module 830 is further configured to: Obtain a plurality of new pixel values of each of the downsampled video frames based on the original resolution and the plurality of pixel values of each of the downsampled video frames; Obtain the plurality of reconstructed video frames based on the plurality of pixel values and the corresponding plurality of new pixel values of each of the downsampled video frames.
[0085] Embodiment III Figure 9Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing the video detection method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 9 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as the program code of the video detection method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.
[0086] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0087] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0088] It should be noted that Figure 9 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0089] In this embodiment, the video detection method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0090] Embodiment 4 The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the video detection method in the embodiments are implemented.
[0091] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the video detection method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0092] Embodiment 5 The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0093] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0094] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of the present application.
Claims
1. A video detection method, characterized in that: The method comprises: Acquire a video frame to be detected, where the video frame to be detected carries information of original resolution; Based on the video frame to be detected, obtaining a plurality of downsampled video frames with different resolutions; Upsampling the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution; and The actual resolution of the video frame to be detected is determined according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected.
2. The method according to claim 1, characterized in that Based on the video frame to be detected, a plurality of down-sampled video frames are obtained, including: Dividing the resolution interval based on the original resolution to obtain multiple resolutions; wherein each resolution in the multiple resolutions is different; The to-be-detected video frame is down-sampled to obtain a plurality of down-sampled video frames corresponding to the plurality of resolutions.
3. The method according to claim 1, characterized in that Determining the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected includes: Based on the information amount of each of the reconstructed video frames and the information amount of the to-be-detected video frame, a plurality of information amount difference values are obtained; According to the multiple information amount difference values, obtaining an information amount difference change curve; The actual resolution of the to-be-detected video frame is determined according to the inflection point of the information amount difference variation curve.
4. The method according to claim 3, characterized in that The amount of information includes pixel information; based on the amount of information of each of the reconstructed video frames and the amount of information of the to-be-detected video frame, a plurality of information amount difference values are obtained, including: By comparing each pixel in the target reconstructed video frame with the corresponding pixel in the video frame to be detected, a plurality of pixel difference values are obtained; wherein the target reconstructed video frame is any one of the plurality of reconstructed video frames; and Based on the multiple pixel difference values, an information amount difference value between the target reconstructed video frame and the to-be-detected video frame is obtained.
5. The method according to any one of claims 1 to 4, characterized in that: Determining the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected, further comprising: Obtaining the similarity between each of the reconstructed video frames and the video frame to be detected; According to the similarity between each of the reconstructed video frames and the video frame to be detected, obtaining a similarity change curve; and The actual resolution of the to-be-detected video frame is determined according to the inflection point of the similarity variation curve.
6. The method according to claim 1, characterized in that Upsampling the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution includes: Based on the original resolution and the multiple pixel values of each of the downsampled video frames, obtaining multiple new pixel values of each of the downsampled video frames; The multiple reconstructed video frames are obtained based on the multiple pixel values of each of the down-sampled video frames and the corresponding multiple new pixel values.
7. A video detection device, characterized in that: The device comprises: An acquisition model is used to acquire a video frame to be detected, where the video frame to be detected carries information of original resolution; A down-sampling module, used for acquiring a plurality of down-sampled video frames based on the video frame to be detected; an upsampling module, upsampling the multiple downsampled video frames to obtain multiple reconstructed video frames corresponding to the original resolution; and The determination module is used to determine the actual resolution of the video frame to be detected according to the information amount of each of the reconstructed video frames and the information amount of the video frame to be detected.
8. A computer device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 6 are implemented.