Multimedia data quality detection method and device, equipment and medium
Through the automated multimedia data quality detection method, multi-dimensional detection of multimedia files is performed using preset quality detection algorithms, which solves the problems of low manual monitoring efficiency and low accuracy in the prior art, and achieves efficient and accurate video quality monitoring.
Patent Information
- Application Number
- CN202510153116.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art relies on manual monitoring in video quality monitoring, resulting in high labor costs, low detection efficiency and low accuracy, making it difficult to timely discover and solve video access delay and lag problems.
By obtaining multimedia files, determining the task to be detected, and performing multidimensional quality detection based on preset quality detection algorithms (including picture, sound and synchronization detection algorithms), we realize automated multimedia data quality detection.
The automation and multi-dimensional quality inspection of multimedia data are realized, the detection accuracy and efficiency are improved, potential quality problems are discovered in a timely manner, and the pressure of manual intervention and operation and maintenance is reduced.
Smart Images

Figure CN120067089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimedia data detection, and in particular, to a method, device, equipment and medium for quality detection of multimedia data. Background Art
[0002] Due to regulatory requirements, banks need to record videos of the customer business handling process. With the increase in the volume of enterprise-related video services, problems such as access latency and video stuttering frequently occur due to physical resources or network performance. Since enterprises currently lack effective video quality monitoring means, the technical department cannot proactively discover problems in a timely manner until complaints are received and feedback is provided, which seriously affects the customer experience and the healthy development of related businesses.
[0003] Existing technologies usually use an integrated business platform system to provide audio and video monitoring support for enterprises. In the financial industry with the most complex enterprise architecture, the corresponding business platform architecture is also more complex. Through the business platform system, global resources can be uniformly monitored and managed, so that IT operation and maintenance personnel can perform quality detection and monitoring on audio and video resources in global resources.
[0004] However, this method of detecting video quality requires continuous monitoring by operation and maintenance personnel, resulting in high labor costs. Due to the complex architecture of the business platform system, the dependence on IT support is increasing, and it is difficult for insufficient operation and maintenance personnel to cope with the heavy operation and maintenance requirements. At the same time, the detection requires manual judgment, which is subjective, inefficient and has low detection accuracy. In this case, an accident will directly affect the business and the responsibility is significant. Summary of the Invention
[0005] The present invention provides a method, device, equipment and medium for quality detection of multimedia data to automatically perform multi-dimensional quality detection on multimedia data and improve detection accuracy.
[0006] According to the first aspect of the present invention, there is provided a method for quality detection of multimedia data, including: Obtaining a multimedia file; Determining a task to be detected according to the multimedia file; Performing quality detection on the task to be detected based on a preset quality detection algorithm to obtain a detection result, where the preset quality detection algorithm includes a picture class detection algorithm, a sound class detection algorithm, and a synchronization class detection algorithm.
[0007] According to the second aspect of the present invention, there is provided a device for quality detection of multimedia data, including: A file acquisition module for obtaining a multimedia file; A task determination module for determining a task to be detected according to the multimedia file; A result determination module, configured to perform quality detection on the task to be detected based on a preset quality detection algorithm, and obtain a detection result, where the preset quality detection algorithm includes a picture type detection algorithm, a sound type detection algorithm, and a synchronization type detection algorithm. According to a third aspect of the present invention, there is provided an electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the quality detection method for multimedia data according to any embodiment of the present invention.
[0008] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the quality detection method for multimedia data according to any embodiment of the present invention when executed.
[0009] According to a fifth aspect of the present invention, an embodiment of the present invention further provides a computer program product, which includes a computer program that implements the quality detection method for multimedia data according to any embodiment of the present invention when executed by a processor.
[0010] The technical solution of the embodiment of the present invention obtains a multimedia file; determines a task to be detected according to the multimedia file; performs quality detection on the task to be detected based on a preset quality detection algorithm, where the preset quality detection algorithm includes a picture type detection algorithm, a sound type detection algorithm, and a synchronization type detection algorithm, and obtains a detection result. By performing quality detection on the multimedia file under different types, the detection results under different detection types are determined. The types of quality detection for the multimedia file are enriched, and through multi-dimensional quality detection, potential quality problems in the multimedia file can be discovered in a timely manner, providing a basis for subsequent quality optimization of the multimedia file.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0013] Figure 1 is a flowchart of a method for detecting the quality of multimedia data provided in Embodiment 1 of the present invention; Figure 2 is a flowchart of a method for detecting the quality of multimedia data provided in Embodiment 2 of the present invention; Figure 3 is a schematic structural diagram of a device for detecting the quality of multimedia data provided in Embodiment 3 of the present invention; Figure 4 is a schematic structural diagram of an electronic device for implementing the embodiments of the present invention. Detailed implementation manners
[0014] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0015] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0016] Embodiment 1 Figure 1 This is a flowchart of a method for detecting the quality of multimedia data provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of quality inspection of multimedia data. This method can be executed by a device for detecting the quality of multimedia data, which can be implemented in the form of hardware and / or software, and the device for detecting the quality of multimedia data can be configured in an electronic device. As Figure 1 shown, this method includes: S110. Obtain a multimedia file.
[0017] In this embodiment, a multimedia file can be understood as a file containing audio and video.
[0018] Specifically, the processor can provide a C interface call method with the upper-layer application through the API layer, and obtain the multimedia file transmitted by the upper-layer application through the API interface. For example, the API layer provides interfaces for creating, starting, stopping, and destroying quality inspection tasks and registering event callbacks, etc. Then, the multimedia task can be transmitted by the upper-layer application creating a quality inspection task.
[0019] S120. Determine the task to be detected according to the multimedia file.
[0020] In this embodiment, the task to be detected can be understood as the task that needs to be quality inspected.
[0021] Specifically, due to the excessive number of transmitted multimedia files, the processor can pre-generate, destroy, and control and monitor the process of the quality inspection task in the control layer. The processor can first decode the multimedia task and conduct a preliminary inspection on the decoding result. For example, it can include a preliminary inspection of the parameters, timestamps, and decoding status of the multimedia file to obtain a preliminary inspection result, and based on the preliminary inspection result, allocate CPU and memory resources for the multimedia task to generate the task to be detected.
[0022] S130. Perform quality inspection on the task to be detected based on a preset quality inspection algorithm to obtain a detection result, where the preset quality inspection algorithm includes a picture class detection algorithm, a sound class detection algorithm, and a synchronization class detection algorithm.
[0023] In this embodiment, the preset quality inspection algorithm can be understood as an algorithm preset for performing multi-dimensional quality inspection. It can include a picture class detection algorithm, a sound class detection algorithm, and a synchronization class detection algorithm. The detection result can be understood as the result used to characterize whether there are quality problems, specific quality problems, and positioning.
[0024] Specifically, the processor can perform quality inspection on the task to be detected through the preset quality inspection algorithm. The preset quality inspection algorithm can perform picture class detection, sound class detection, and synchronization class detection on the task to be detected in parallel to obtain multi-dimensional detection results.
[0025] The technical solution of the embodiment of the present invention obtains a multimedia file; determines the task to be detected according to the multimedia file; performs quality inspection on the task to be detected based on a preset quality inspection algorithm, where the preset quality inspection algorithm includes a picture class detection algorithm, a sound class detection algorithm, and a synchronization class detection algorithm, to obtain a detection result. By performing quality inspection on the multimedia file under different types, the detection results under different detection types are determined. It enriches the types of quality inspection of the multimedia file. Through multi-dimensional quality inspection, potential quality problems in the multimedia file can be discovered in a timely manner, providing a basis for the subsequent quality optimization of the multimedia file.
[0026] As a first alternative embodiment of the first embodiment, after performing quality inspection on the task to be inspected based on a preset quality inspection algorithm and obtaining the inspection result, it further includes: Generating a quality inspection report for the multimedia file based on the inspection result.
[0027] In this embodiment, the quality inspection report can be understood as a report for presenting the inspection results of various quality inspection items in the multimedia file.
[0028] Specifically, for the quality inspection items in different dimensions in the inspection result, the processor can fill in the inspection result according to the template of the quality inspection report to generate the quality inspection report.
[0029] Embodiment Two Figure 2 The flowchart of a quality inspection method for multimedia data provided by the second embodiment of the present invention. This embodiment is a further refinement between the above embodiments. As Figure 2 shown, the method includes: S201. Obtain a multimedia file.
[0030] S202. Decode the multimedia file to obtain file basic information and multimedia frame information. The file basic information includes file parameters and data timestamps.
[0031] In this embodiment, the file basic information can be understood as information characterizing the file attributes. For example, it may include file parameters and data timestamps. The file parameters may include parameters such as video width and height, video bit rate, audio bit rate, video and audio length, video frame rate, audio sampling rate, and audio and video coding types; the data timestamp can be understood as the download time of each multimedia data. The multimedia frame information can be understood as including the multimedia data under each frame.
[0032] Specifically, the processor can decode the multimedia file through relevant function interfaces to obtain the file basic information and the decoded multimedia frame information. For example, taking an MP4 file as an example, information such as video width and height, video bit rate, audio bit rate, video and audio length, video frame rate, audio sampling rate, and audio and video coding types can be obtained by parsing the MP4 header format.
[0033] S203. Verify the file parameters to determine the first verification result.
[0034] In this embodiment, the first verification result can be understood as a result for characterizing whether the file parameters are normal.
[0035] Specifically, the processor can verify the file parameters, and can determine whether the lengths of the audio and video in the file parameters are the same, and whether the difference is within the controllable range of the threshold time (such as 3 seconds); it can also determine whether the quality inspection ability can support the quality inspection according to the parameters of the audio and video, and can also determine whether the video bit rate is insufficient through a preset formula based on the video width, height and bit rate, so as to obtain the first verification result.
[0036] S204. Verify the interval between multimedia frame information according to each data timestamp and the preset interval threshold, and determine the second verification result.
[0037] In this embodiment, the preset interval threshold can be understood as the threshold for judging whether the interval between two videos is normal. The second verification result can be understood as the verification result for characterizing whether the time is abnormal.
[0038] Specifically, the processor can read each packet of multimedia frame information, and each packet of multimedia frame information will have a data timestamp. The processor can judge whether the interval time between two data timestamps is legal through the preset interval threshold. For example, according to a frame rate of 15 frames, the video frame interval should be about 66 milliseconds. The processor will take 66 milliseconds plus the set preset interval threshold (for example, 40 milliseconds), so the maximum interval is about 106 milliseconds. If the interval between two adjacent multimedia frames calculated from the data timestamp is greater than this value, the time point will be recorded to generate an abnormal time point to obtain the second verification result.
[0039] S205. Perform decoding verification on the multimedia frames to determine the third verification result.
[0040] In this embodiment, the third verification result is used to characterize the result of whether the decoding is successful.
[0041] Specifically, the processor can decode the multimedia frame information to determine whether the decoding is successful. If the decoding is not successful, the abnormal time point corresponding to the corresponding multimedia frame information will be recorded as the third verification result. If the audio and video coding is not supported by the quality inspection ability and the decoding is abnormal, subsequent other types of quality inspection functions cannot be performed. However, if only a certain frame in the middle fails to decode, only that frame will not be subject to subsequent quality inspection functions, and the subsequent normal data will continue to be detected.
[0042] S206. Determine the task to be detected according to the first verification result, the second verification result, the third verification result and the multimedia frame information.
[0043] Specifically, the processor can determine the multimedia data frames that cannot be subject to subsequent quality inspections from the multimedia data frame information based on the first verification result, the second verification result, and the third verification result, and generate a task to be detected. Since the quality detection algorithm includes multiple dimensions, multi-threaded detection can be performed simultaneously, and the task to be detected can be distributed to the global thread pool composed of different detection dimensions.
[0044] S207. When the preset quality detection algorithm is a picture class detection algorithm, perform color gamut conversion on each picture frame in the multimedia frame information of the task to be detected to obtain a first target picture frame.
[0045] In this embodiment, the picture frame can be understood as the picture in the multimedia frame information in units of frames. The first target picture frame can be understood as the picture frame after color gamut conversion.
[0046] Specifically, the processor can first perform color gamut conversion on each picture frame in the multimedia frame information of the task to be detected based on the picture class detection algorithm, and convert each picture frame to a black and white color gamut. For example, it can be converted from the YUV color gamut to the HSV color gamut.
[0047] Exemplarily, the model of the hue saturation value (HSV) color space corresponds to a conical subset in the cylindrical coordinate system. The top surface of the cone corresponds to V = 1, which contains the three surfaces of R = 1, G = 1, and B = 1 in the RGB model, and the colors represented are brighter. The color H is given by the rotation angle around the V axis. Red corresponds to the angle 0°, green corresponds to the angle 120°, and blue corresponds to the angle 240°. In the HSV color model, each color and its complementary color differ by 180°. The saturation S ranges from 0 to 1, so the radius of the top surface of the cone is 1. The color domain represented by the HSV color model is a subset of the CIE chromaticity diagram. For the colors with 100% saturation in this model, their purity is generally less than 100%. At the vertex (i.e., the origin) of the cone, V = 0, H and S are undefined, representing black. At the center of the top surface of the cone, S = 0, V = 1, and H is undefined, representing white. From this point to the origin represents the gray with gradually decreasing brightness, that is, the gray with different grayscales. For these points, S = 0 and the value of H is undefined. It can be said that the V axis in the HSV model corresponds to the main diagonal in the RGB color space. For the colors on the circumference of the top surface of the cone, V = 1 and S = 1, and this kind of color is a pure color. The HSV model corresponds to the method of painters' color matching. Painters obtain different hues of colors from a certain pure color by changing the color concentration and color depth. Adding white to a pure color changes the color concentration, adding black changes the color depth, and adding different proportions of white and black can obtain various different hues.
[0048] S208. Detect each first target frame according to the black-and-white ratio threshold information to obtain a sub-result of frame detection.
[0049] In this embodiment, the black-and-white ratio threshold information can be understood as the threshold for detecting whether the ratio of black and white colors is normal, including a black screen threshold and a white screen threshold. For example, the hue in the black screen threshold is 0 - 180, the saturation is 0 - 255, and the brightness is 0 - BlackVal; the hue in the white screen threshold is 0 - 180, the saturation is 0 - 30, and the brightness is WhiteVal - 255. The sub-result of frame detection can be understood as the result for characterizing whether the black and white of the frame are normal.
[0050] Specifically, the processor can compare the pixel points of each first target frame according to the black-and-white ratio threshold information to determine whether the black ratio threshold is reached and whether the white ratio threshold is reached, so as to obtain the sub-result of frame detection.
[0051] S209. Perform blur detection on each frame based on the preset blur detection threshold information to obtain a sub-result of blur detection.
[0052] In this embodiment, the preset blur detection threshold information can be understood as the threshold for determining whether each frame is blurred. For example, it can include a Hamming distance threshold and a high-pass filter threshold. The sub-result of blur detection can be understood as the detection result for characterizing whether it is blurred.
[0053] Specifically, after the processor obtains a frame of the video frame, it can calculate the pHash value of this frame by the perceptual hashing algorithm and combine it with the pHash value of the previous frame to determine the Hamming distance between the two frames. Among them, the greater the Hamming distance, the greater the difference in the change of the frames. Then, compare the Hamming distance between the two frames with the Hamming distance threshold (for example, it can be set to 5) in the preset blur detection threshold information. If it is less, continue to detect the next frame of the video frame. If it is greater than or equal to, count the high-frequency information of the video frame, and compare the high-frequency information with the high-pass filter threshold in the preset blur detection threshold information to determine the sub-result of blur detection.
[0054] It can be understood that the high-frequency information of a picture frame is the edges that appear in the picture structure. The clearer and more numerous the edges are, the higher the high-frequency information of the picture represents. To determine whether a picture is blurred or clear, it is generally by judging whether there are clear edges in the image picture. Usually, it can be understood that the clearer the picture, the higher the edge information will be. Therefore, the first step is to obtain the edge information of the picture frame, and the acquisition of edge information is based on an algorithm in the spatial domain (x, y). However, the algorithm in the spatial domain will have inconsistent measurement of the blur reference value at different resolutions. For example, a value calculated for Video A indicates a clear judgment, but the judgment may be unclear for videos of different sizes. The fundamental reason is that the resolutions are different. Therefore, by performing a Fourier transform to convert the spatial-domain image into a frequency-domain (amplitude, phase angle) image and detecting the edge information, the problem of inconsistent measurement of the blur reference value due to differences in picture size can be eliminated. When using the frequency-domain algorithm, the edge information belongs to the high-frequency information. Therefore, by using a high-pass filter to filter out the low-frequency information in the frequency domain, the remaining is the clear information of this image, which can be used to compare whether there is blur. The corresponding threshold is the range of the high-pass filter threshold: 0 to 100, with a default value of 30. Setting it to 0 means no filtering, that is, always clear, and setting it to 100 means all filtering, that is, always blurred.
[0055] S210. Cache the multimedia frame information according to a set duration, and perform stuttering picture detection based on the cache result to obtain a stuttering detection sub-result.
[0056] In this embodiment, the set duration can be understood as the caching duration. For example, each time 1 second of video frames is cached. The cache result can be understood as the result obtained after caching the multimedia frame information. The stuttering detection sub-result can include the detection results of the stuttering time point and the duration of the video picture.
[0057] Specifically, the processor can cache the multimedia frame information according to a set duration to obtain a cache result at the set duration, and perform stuttering picture detection on the cache result to obtain a stuttering detection sub-result.
[0058] Further, on the basis of the above embodiment, the step of caching the multimedia frame information according to a set duration and performing stuttering picture detection based on the cache result to obtain a stuttering detection sub-result can be refined as: Cache the multimedia frame information according to the set duration to obtain a cache result, and use the earliest frame in the cache result as the current frame; compare the similarity between the current frame and the next frame adjacent to the current frame to obtain the current frame similarity; if the current frame similarity exceeds the preset similarity threshold, determine whether there is a stuttering frame according to the number of frames between the current frame and the earliest frame and the preset stuttering threshold to obtain a judgment result, use the next frame as the current frame, and continue to execute the step of determining the current frame similarity; otherwise, use the next frame as the current frame, and continue to execute the step of determining the current frame similarity; when the similarity of all frames in the cache result has been compared, determine the stuttering detection sub-result according to the judgment result.
[0059] In this embodiment, the earliest frame can be understood as the first frame in a cache. The current frame can be understood as the frame currently used for judgment. The next frame can be understood as the frame adjacent to the current frame in the cache and in the next time sequence. The current frame similarity can be understood as representing the similarity between the current frame and the next frame. The preset similarity threshold can be understood as the threshold for determining whether two frames are similar. The number of frames between can be understood as the number of frames between the current frame and the earliest frame. The preset stuttering threshold can be understood as the threshold for determining whether there is a frame stuttering situation. The judgment result can be understood as the result representing whether there is stuttering.
[0060] Specifically, the processor can cache the multimedia frame information according to the set duration to obtain a cache result, and use the earliest frame in the cache result as the current frame. The processor can compare the similarity between the current frame and the next frame adjacent to the current frame. The range of the preset similarity threshold can be set from 0 to 100. For example, the preset similarity threshold P1 is P1 = 90. After calculating the similarity between the current frame and the next frame, a current frame similarity S1 can be obtained. Determine whether S1 exceeds P1. If the current frame similarity exceeds the preset similarity threshold, determine whether there is a stuttering frame according to the number of frames between the current frame and the earliest frame and the preset stuttering threshold to obtain a judgment result. For example, if there are 15 frames in 1 second, the default stuttering threshold number of frames is the lower threshold of 0.3, 15 * 0.3 is 5 frames, the upper threshold is 0.6, and 15 * 0.6 is 9 frames. Then frames in the range above 5 frames and below 9 frames are recorded as stuttering frames, and the others are non-stuttering frames. If it is not a stuttering frame, use the next frame as the current frame, and continue to execute the step of determining the current frame similarity; if the current frame similarity does not exceed the preset similarity threshold, use the next frame as the current frame, and continue to execute the step of determining the current frame similarity; when the similarity of all frames in the cache result has been compared, determine the stuttering detection sub-result according to the judgment result.
[0061] S211. Use the screen detection sub-result, blur detection sub-result, and stutter detection sub-result as the detection result.
[0062] Specifically, the processor can use the screen detection sub-result, blur detection sub-result, and stutter detection sub-result together as the detection result.
[0063] S212. When the preset quality detection algorithm is a sound detection algorithm, perform volume detection on each audio frame of the multimedia frame information in the task to be detected, and determine whether the volume of the audio frame is greater than the preset volume threshold.
[0064] In this embodiment, the audio frame can be understood as the audio of each frame. The preset volume threshold can be understood as the threshold for determining whether it is muted.
[0065] Specifically, when the preset quality detection algorithm is a sound detection algorithm, the processor can perform volume detection on each audio frame of the multimedia frame information in the task to be detected, and determine whether the volume of the audio frame is greater than the preset volume threshold.
[0066] S213. If not, regard the audio frame as a muted frame, accumulate the number of muted frames, and obtain the number of muted frames.
[0067] In this embodiment, the number of muted frames can be understood as the number of frames continuously determined to be muted frames.
[0068] Specifically, if the volume of the audio frame is less than or equal to the preset volume threshold, it is determined that the frame is a muted frame, and the number of muted frames is incremented by one to obtain the current number of muted frames.
[0069] S214. If the number of muted frames reaches the preset frame number threshold, perform a mute mark on the audio frames corresponding to the number of muted frames.
[0070] In this embodiment, the preset frame number threshold can be understood as the threshold for determining that a period of time is a muted frame. The mute mark can be understood as a mark used to indicate that an audio frame is a muted frame.
[0071] Specifically, the processor can compare the obtained number of muted frames with the preset frame number threshold. If the number of muted frames reaches the preset frame number threshold, perform a mute mark on the audio frames corresponding to the number of muted frames.
[0072] S215. Otherwise, continue to detect the next audio frame.
[0073] In this embodiment, the next audio frame can be understood as the audio frame of the next frame of the audio frame currently being judged.
[0074] Specifically, if the number of muted frames does not reach the preset frame number threshold, continue to perform volume detection on the next audio frame.
[0075] S216. If so, clear the number of silent frames and continue to detect the next audio frame.
[0076] Specifically, if the volume of the audio frame is greater than the preset volume threshold, clear the number of silent frames and continue to detect the next audio frame.
[0077] S217. When all audio frames in the multimedia frame information are detected, use the target audio segment with a silent mark as the detection result.
[0078] In this embodiment, the target audio segment can be understood as the set of all audio determined to be silent segments.
[0079] Specifically, when all audio frames in the multimedia frame information are detected, the processor can use the target audio segment composed of all audio frames with a silent mark as the detection result.
[0080] When the preset quality detection algorithm is a synchronization detection algorithm, perform voice activity detection on the multimedia frame information in the task to be detected and filter out voice activity segments.
[0081] In this embodiment, the voice activity segment can be understood as the set of all frames containing speech.
[0082] Specifically, when the preset quality detection algorithm is a synchronization detection algorithm, the processor can perform voice activity detection (Voice Activity Detection, VAD) on the multimedia frame information in the task to be detected, and filter out the voice activity segments containing voice activity from the multimedia frame information to reduce redundant calculations for non-voice activity segments.
[0083] S219. Perform face detection and tracking on each voice activity segment to determine the segment sequence of each speaking target.
[0084] In this embodiment, the speaking target can be understood as the target object with voice activity. The segment sequence can be understood as the segment sequence with a time sequence composed of all voice activity segments of the speaking target.
[0085] Specifically, the processor can perform face detection on each voice activity segment to determine each speaking target, and perform tracking for each speaking target to determine the segment sequence of each speaking target in the voice activity segment.
[0086] S220. Perform audio-visual synchronization detection on the audio in the segment sequence and the target action changes of the speaking target to obtain the detection result.
[0087] In this embodiment, the target action change can be understood as the change representing the voice action of the speaking target.
[0088] Specifically, the processor can perform audio-visual synchronization detection on the audio in the segment sequence and the target action changes of the speaking target, determine whether there is a problem of audio-visual out-of-synchronization, and obtain the detection result.
[0089] The technical solution of the embodiment of the present invention is as follows: By obtaining a multimedia file, decoding the multimedia file, determining the file basic information and multimedia frame information, and through pre-checking the file basic information to determine the check result, first screen out the parts that do not need to be checked to form a task to be detected. Based on the screen class detection algorithm, detect the black and white screen of the screen to obtain the screen detection sub-result; perform blurring detection on the screen frames to obtain the blurring detection sub-result; perform stuck screen detection on the cache result to obtain the stuck detection sub-result, and then obtain a comprehensive detection result in the screen dimension, enriching the types of quality detection, performing all-round screen detection, and ensuring the comprehensiveness of quality inspection. By using the sound class detection algorithm to perform mute detection on the task to be detected, the detection result of the sound class is obtained, enriching the types of quality inspection. By using the synchronization class detection algorithm to perform audio-visual synchronization detection on the multimedia frame information, the detection result of the audio-visual synchronization class is obtained. By performing quality detection on the multimedia file under different types, the detection results under different detection types are determined. The types of quality detection of the multimedia file are enriched, and through multi-dimensional quality detection, potential quality problems in the multimedia file are timely discovered, providing a basis for the subsequent quality optimization of the multimedia file.
[0090] Embodiment III Figure 3 It is a schematic structural diagram of a quality detection device for multimedia data provided by Embodiment III of the present invention. As Figure 3 shown, the device includes: A file acquisition module 31, configured to acquire a multimedia file; A task determination module 32, configured to determine a task to be detected according to the multimedia file; A result determination module 33, configured to perform quality detection on the task to be detected based on a preset quality detection algorithm to obtain a detection result, where the preset quality detection algorithm includes a screen class detection algorithm, a sound class detection algorithm, and a synchronization class detection algorithm. The technical solution of the embodiment of the present invention is as follows: By obtaining a multimedia file; determining a task to be detected according to the multimedia file; performing quality detection on the task to be detected based on a preset quality detection algorithm, where the preset quality detection algorithm includes a screen class detection algorithm, a sound class detection algorithm, and a synchronization class detection algorithm, to obtain a detection result. By performing quality detection on the multimedia file under different types, the detection results under different detection types are determined. The types of quality detection of the multimedia file are enriched, and through multi-dimensional quality detection, potential quality problems in the multimedia file are timely discovered, providing a basis for the subsequent quality optimization of the multimedia file.
[0091] Further, the task determination module 32 is specifically configured to: Decode the multimedia file to obtain file basic information and multimedia frame information, where the file basic information includes file parameters and data timestamps; Verify the file parameters to determine a first verification result; Verify the intervals between the multimedia frame information according to each of the data timestamps and a preset interval threshold to determine a second verification result; Perform decoding verification on the multimedia frame information to determine a third verification result; Determine a task to be detected according to the first verification result, the second verification result, the third verification result, and the multimedia frame information.
[0092] Further, when the preset quality detection algorithm is the picture class detection algorithm, correspondingly, the result determination module 33 includes: A first determination unit, configured to perform color gamut conversion on each picture frame of the multimedia frame information in the task to be detected to obtain a first target picture frame; A second determination unit, configured to detect each of the first target picture frames according to black-and-white ratio threshold information to obtain a picture detection sub-result; A third determination unit, configured to perform blur detection on each of the picture frames based on preset blur detection threshold information to obtain a blur detection sub-result; A fourth determination unit, configured to cache the multimedia frame information according to a set duration, and perform stuttering picture detection based on the cache result to obtain a stuttering detection sub-result; A fifth determination unit, configured to use the picture detection sub-result, the blur detection sub-result, and the stuttering detection sub-result as detection results.
[0093] Among them, the fourth determination unit is specifically configured to: Cache the multimedia frame information according to a set duration to obtain a cache result, and use the earliest picture frame in the cache result as the current frame picture; Compare the similarity between the current frame picture and the next frame picture adjacent to the current frame picture to obtain a current frame similarity; If the current frame similarity exceeds a preset similarity threshold, then perform stuttering frame determination according to the number of frames difference between the current frame picture and the earliest picture frame and a preset stuttering threshold to determine a determination result, use the next picture frame as the current frame picture, and continue to execute the step of determining the current frame similarity; Otherwise, use the next picture frame as the current frame picture, and continue to execute the step of determining the current frame similarity; When similarity comparison is performed on all frame images in the cache result, a stuttering detection sub-result is determined according to the judgment result.
[0094] Further, when the preset quality detection algorithm is the sound class detection algorithm, correspondingly, the result determination module 33 is specifically configured to: Perform volume detection on each audio frame of the multimedia frame information in the task to be detected, and determine whether the volume of the audio frame is greater than a preset volume threshold; If not, the audio frame is used as a silent frame, and the number of silent frames is accumulated to obtain the number of silent frames; If the number of silent frames reaches a preset frame number threshold, the audio frames corresponding to the number of silent frames are marked as silent; otherwise, the next audio frame is continuously detected; If so, the number of silent frames is cleared, and the next audio frame is continuously detected; When all the audio frames in the multimedia frame information are detected, the target audio segment with the silent mark is used as the detection result.
[0095] Further, when the preset quality detection algorithm is the synchronization class detection algorithm, correspondingly, the result determination module 33 is specifically configured to: Perform voice activity detection on the multimedia frame information in the task to be detected, and screen out voice activity segments; Perform face detection and tracking on each voice activity segment to determine the segment sequence of each speaking target; Perform audio-visual synchronization detection on the audio in the segment sequence and the target action change of the speaking target to obtain a detection result.
[0096] Optionally, the device further includes: A report generation module, configured to generate a quality inspection report of the multimedia file based on the detection result after performing quality detection on the task to be detected based on a preset quality detection algorithm to obtain a detection result.
[0097] The quality detection device for multimedia data provided by the embodiments of the present invention can execute the quality detection method for multimedia data provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0098] Embodiment 4 Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 40 that can be used to implement embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0099] As Figure 4 shown, the electronic device 40 includes at least one processor 41 and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0100] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0101] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the method for detecting the quality of multimedia data.
[0102] In some embodiments, the method for quality detection of multimedia data may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the method for quality detection of multimedia data described above may be performed. Alternatively, in other embodiments, the processor 41 may be configured to perform the method for quality detection of multimedia data by any other suitable means (e.g., by means of firmware).
[0103] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0104] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0105] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0106] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide an interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0107] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0108] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0109] In one embodiment, the embodiment of the present invention further includes a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the quality detection method for multimedia data according to any embodiment of the present invention.
[0110] In the process of implementation, the computer program product can be written in one or more programming languages or combinations thereof to write computer program code for performing the operations of the present invention. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0111] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitations are imposed herein.
[0112] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting the quality of multimedia data, characterized in that: include: Get multimedia files; Determining a task to be detected according to the multimedia file; The quality of the task to be detected is detected based on a preset quality detection algorithm to obtain a detection result, wherein the preset quality detection algorithm includes a picture detection algorithm, a sound detection algorithm and a synchronization detection algorithm.
2. The method according to claim 1, characterized in that The step of determining the task to be detected according to the multimedia file includes: Decoding the multimedia file to obtain basic file information and multimedia frame information, wherein the basic file information includes file parameters and data timestamp; Verifying the file parameters to determine a first verification result; Verify the interval between the multimedia frame information according to each of the data timestamps and a preset interval threshold, and determine a second verification result; Decoding and verifying the multimedia frame information to determine a third verification result; A task to be detected is determined according to the first verification result, the second verification result, the third verification result and the multimedia frame information.
3. The method according to claim 1, characterized in that When the preset quality detection algorithm is the picture type detection algorithm, correspondingly, performing quality detection on the task to be detected based on the preset quality detection algorithm to obtain a detection result includes: Performing color gamut conversion on each picture frame of the multimedia frame information in the task to be detected to obtain a first target picture frame; Detect each of the first target picture frames according to the black-white ratio threshold information to obtain a picture detection sub-result; Performing blur detection on each of the picture frames based on preset blur detection threshold information to obtain blur detection sub-results; Cache the multimedia frame information according to a set duration, and perform freeze detection based on the cached result to obtain a freeze detection sub-result; The picture detection sub-result, the blur detection sub-result and the freeze detection sub-result are used as detection results.
4. The method according to claim 3, characterized in that The caching of the multimedia frame information according to the set duration, and performing freeze screen detection based on the cached result to obtain a freeze detection sub-result, include: Cache the multimedia frame information according to a set time length to obtain a cache result, and use the earliest picture frame in the cache result as the current frame; Comparing the similarity between the current frame and the next frame adjacent to the current frame to obtain the current frame similarity; If the current frame similarity exceeds a preset similarity threshold, a stuck frame judgment is performed according to the frame difference between the current frame and the earliest frame and a preset stuck threshold, a judgment result is determined, the next frame is used as the current frame, and the step of determining the current frame similarity is continued; Otherwise, the next picture frame is used as the current frame, and the step of determining the similarity of the current frame is continued; When all frames in the cached result are compared for similarity, a freeze detection sub-result is determined according to the judgment result.
5. The method according to claim 1, characterized in that When the preset quality detection algorithm is the sound detection algorithm, correspondingly, performing quality detection on the task to be detected based on the preset quality detection algorithm to obtain a detection result includes: Performing volume detection on each audio frame of the multimedia frame information in the task to be detected, and determining whether the volume of the audio frame is greater than a preset volume threshold; If not, the audio frame is regarded as a silent frame, and the number of silent frames is accumulated to obtain the number of silent frames; If the number of silent frames reaches a preset frame number threshold, the audio frame corresponding to the number of silent frames is marked as silent; otherwise, the next audio frame is detected; If yes, clear the number of silent frames and continue to detect the next audio frame; When all the audio frames in the multimedia frame information are detected, the target audio segment with the silence mark is taken as the detection result.
6. The method according to claim 1, characterized in that When the preset quality detection algorithm is the synchronization detection algorithm, correspondingly, performing quality detection on the task to be detected based on the preset quality detection algorithm to obtain a detection result includes: Performing voice activity detection on the multimedia frame information in the task to be detected, and screening voice activity segments; Performing face detection and tracking on each of the speech activity segments to determine a segment sequence of each speaking target; Performing audio and video synchronization detection on the audio in the segment sequence and the target action change of the speaking target to obtain a detection result.
7. The method according to claim 1, characterized in that After performing quality inspection on the task to be inspected based on the preset quality inspection algorithm and obtaining the inspection result, the method further includes: A quality inspection report of the multimedia file is generated based on the inspection result.
8. A multimedia data quality detection device, characterized in that: include: A file acquisition module, used to acquire multimedia files; A task determination module, used to determine a task to be detected according to the multimedia file; The result determination module is used to perform quality detection on the task to be detected based on a preset quality detection algorithm to obtain a detection result, wherein the preset quality detection algorithm includes a picture detection algorithm, a sound detection algorithm and a synchronization detection algorithm.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multimedia data quality detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimedia data quality detection method according to any one of claims 1 to 7 when executed.
Citation Information
Patent Citations
Multi-mode parallel video quality fault detection method and device
CN103873852A
Video quality detection method based on image processing, storage device and mobile terminal
CN112291551A
Video quality detection method and device, equipment, storage medium and program product
CN114928740A
Double-recording video quality detection method and device, electronic equipment and storage medium
CN116744054A
Audio and video non-silent segment detection method and device, equipment and storage medium
CN117953925A