Video optimization method, system and device and storage medium

By analyzing and editing the audio and picture stagnant parts in low-quality short videos, the problems of low efficiency and lack of automation in the existing technology are solved, efficient and automated video optimization is achieved, and video quality and user experience are improved.

CN120223952APending Publication Date: 2025-06-27SO-YOUNG INT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311810339.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is inefficient when processing low-quality short videos, requiring manual editing, with a high threshold, and lack of online process-based solutions, resulting in poor user experience and low completion rate.

Method used

By obtaining the cloud video address of the original video, analyzing the duration of audio and picture stagnation, automatically performing video editing, and optimizing video quality, the entire process is carried out on the server side to avoid client traffic consumption.

Benefits of technology

It realizes efficient and automated video optimization, saves bandwidth costs and manual operation costs, improves video quality, increases playback rate and user stickiness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223952A_ABST
    Figure CN120223952A_ABST
Patent Text Reader

Abstract

The invention discloses a video optimization method, which is applied to a server and comprises the following steps: acquiring video information of an original video according to a cloud video address of the original video; obtaining audio loudness data according to the video information, and analyzing the audio stagnation duration of the original video; frame screenshot is carried out on the original video, and the picture stagnation duration of the original video is analyzed; and determining a part needing to be clipped in the original video according to the audio stagnation duration and the picture stagnation duration, and clipping the original video to obtain an optimized video. The invention further discloses a video optimization system, an electronic device and a computer readable storage medium. Therefore, the optimization of the low-quality short video can be completed more efficiently, and the flow of the client side is not consumed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and in particular, to a video optimization method, system, electronic device, and computer-readable storage medium. Background Art

[0002] With the development of network technologies and mobile terminals, short videos, as a new media form, have gradually been accepted by the public. Many self-media and individuals choose short videos as a way to express their views and showcase their lives. However, for low-quality short videos, necessary editing is required before they can be published on the Internet. Otherwise, it will affect the user viewing experience, resulting in a low completion rate and easy user loss.

[0003] Currently, the optimization of low-quality short videos still remains at the stage of manual editing and processing of local files, with low efficiency and time-consuming. And there is a certain threshold for short video editing. Not everyone has a good foundation in operating video processing software, which will bring insurmountable obstacles to users without a foundation. Moreover, existing video or image processing technologies are relatively scattered, and there is no publicly available and mature online process-based solution. In addition, in aspects such as image comparison and sound detection, there are still problems such as inaccuracy and high traffic consumption. Summary of the Invention

[0004] The main purpose of this application is to propose a video optimization method, system, electronic device, and computer-readable storage medium, aiming to solve at least one of the above technical problems.

[0005] In a first aspect, an embodiment of this application provides a video optimization method, which is applied to a server. The method includes:

[0006] Obtain video information of the original video according to the cloud video address of the original video;

[0007] Obtain audio loudness data according to the video information, and analyze the audio stagnation duration of the original video;

[0008] Analyze the picture stagnation duration of the original video by taking frame screenshots of the original video;

[0009] Determine the part to be cropped in the original video according to the audio stagnation duration and the picture stagnation duration, and edit the original video to obtain an optimized video.

[0010] In a second aspect, an embodiment of this application provides a video optimization system, which is applied to a server. The system includes:

[0011] A receiving module, configured to obtain video information of the original video according to the cloud video address of the original video;

[0012] An analysis module, configured to obtain audio loudness data according to the video information and analyze the audio stagnation duration of the original video;

[0013] The analysis module is further configured to analyze the picture stagnation duration of the original video by taking frame screenshots of the original video;

[0014] A clip module, configured to determine the part of the original video that needs to be cropped according to the audio stagnation duration and the picture stagnation duration, and clip the original video to obtain an optimized video;

[0015] An upload module, configured to upload the optimized video to a cloud server to obtain a new cloud video address.

[0016] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor, and a video optimization program stored on the memory and executable on the processor. When the video optimization program is executed by the processor, the video optimization method as described above is implemented.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a video optimization program is stored. When the video optimization program is executed by a processor, the video optimization method as described above is implemented.

[0018] The video optimization method, system, electronic device, and computer-readable storage medium proposed in the embodiments of the present application only need to pass the cloud video address to the service interface of the server to obtain the video information of the original video, and analyze the audio stagnation duration and the picture stagnation duration, so as to clip the original video. At the same time, the local traffic of the client is not consumed during the whole process, and all processing is completed on the server side, which greatly saves the bandwidth cost brought by downloading videos. In addition, no manual participation is required during the whole process, which saves costs and more efficiently completes the optimization of low-quality short videos. The overall quality of the video optimized by this solution is significantly improved, which can effectively increase the completion rate and browsing step length, thereby enhancing user stickiness. Description of the Drawings

[0019] The drawings here are used to provide a further understanding of the present application and constitute a part of the present application. It should be understood that these drawings only depict some embodiments disclosed according to the present application and should not be regarded as a limitation on the scope of the present application.

[0020] Figure 1 An application environment architecture diagram for implementing various embodiments of the present application;

[0021] Figure 2 A flowchart of a video optimization method proposed in the first embodiment of the present application;

[0022] Figure 3 is Figure 2 a detailed flowchart of step S202 in

[0023] Figure 4 is Figure 2 a detailed flowchart of step S204 in

[0024] Figure 5 is Figure 2 the first detailed flowchart of step S206 in

[0025] Figure 6 is Figure 2 the second detailed flowchart of step S206 in

[0026] Figure 7 is Figure 2 the third detailed flowchart of step S206 in

[0027] Figure 8 is a flowchart of a video optimization method proposed in the second embodiment of the present application;

[0028] Figure 9 is a flowchart of another form of the video optimization method in the second embodiment of the present application;

[0029] Figure 10 is a schematic diagram of the hardware architecture of an electronic device proposed in the third embodiment of the present application;

[0030] Figure 11 is a schematic diagram of the modules of a video optimization system proposed in the fourth embodiment of the present application;

[0031] Figure 12 is a schematic diagram of the modules of a video optimization system proposed in the fifth embodiment of the present application. Detailed implementation manners

[0032] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0033] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes, and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0034] The following provides the term explanations involved in the present application:

[0035] Short video: That is, a short film video, which is a way of spreading Internet content. Generally, it is a video with a duration of less than 5 minutes spread on Internet new media. With the popularization of mobile terminals and the acceleration of network speed, short, flat, and fast large-traffic dissemination content has gradually gained the favor of major platforms, fans, and capital.

[0036] Cloud server: It is a simple, efficient, secure, and reliable computing service with elastic scalability of processing power. Its management method is simpler and more efficient than that of physical servers. Users can quickly create or release any number of cloud servers without the need to purchase hardware in advance.

[0037] MySQL: It is a relational database management system. Relational databases store data in different tables instead of putting all data in a large warehouse, which increases speed and improves flexibility.

[0038] FFmpeg: It is a set of open-source computer programs that can be used to record, convert digital audio and video, and convert them into streams. As a multimedia video processing tool, FFmpeg has very powerful functions, including video capture, video format conversion, video capture, adding watermarks to videos, etc.

[0039] RPC (Remote Procedure Call): It is a computer communication protocol that allows calls between programs in different process spaces. The implementation of RPC usually includes two parts: a client and a server. The client initiates a remote call request, and the server receives the request and responds. Its client and server can be on the same machine or on different machines, and when used, it is just like calling a local program without the need to pay attention to implementation details. Therefore, it has higher flexibility and scalability compared to other communication methods. In RPC, the parameters and return values of remote calls can be of any data type, such as basic data types, custom objects, etc.

[0040] Go (also known as Golang): A statically typed, compiled, concurrent programming language with garbage collection developed by Google. The syntax of the Go language is similar to that of the C language, but the variable declarations are different. Go embeds associative arrays (also known as hash tables or dictionaries), just like the string type.

[0041] Loudness: The subjective perceptual quantity of the strength of sound, also known as volume. The unit of loudness is not decibel. There is an international standard for loudness measurement - EBU R128. EBU R128 is a recommendation on audio signal loudness standardization and maximum level, which is mainly used in the audio mixing of TV and radio programs and is used to measure and control the loudness of programs. EBU R128 adopts an international standard to measure audio loudness and specifically creates units for loudness measurement: LU (loudness units, relative loudness measurement unit) and LUFS (loudness units referenced to fullscale, absolute loudness measurement unit). Additionally, a signal below the absolute threshold of -70 LUFS can be regarded as the silent part of the audio.

[0042] pHash algorithm: An algorithm used for image retrieval and similarity calculation. It is based on the perceptual hash algorithm and represents the features of an image by converting the image into a fixed-length hash value. The pHash algorithm can be applied to various fields, such as image search, copyright protection, and image recognition.

[0043] Please refer to Figure 1 , Figure 1 For an application environment architecture diagram for implementing various embodiments of this application. This application can be applied to an application environment including, but not limited to, client 2, server 4, and cloud server 6.

[0044] Among them, the client 2 is used to upload the original video to the cloud server 6, obtain the online cloud video address, and send the cloud video address to the server 4 so that the server 4 can perform video editing processing.

[0045] The server 4 is used to obtain the original video information according to the cloud video address, analyze the audio stagnation duration and the video frame stagnation duration, compare the two durations and take the maximum value to obtain the duration that needs to be trimmed, perform editing processing on the original video, and upload the optimized video to the cloud server 6 to generate a new cloud video address.

[0046] The cloud server 6 is used to save video files and cloud video addresses.

[0047] The client 2 can be a terminal device such as a PC (Personal Computer), mobile phone, tablet computer, portable computer, wearable device, etc. The server 4 can be a computing device such as a rack server, blade server, tower server or cabinet server, and can be an independent server or a server cluster composed of multiple servers. The cloud server 6 can be a cloud server.

[0048] The client 2, the server 4, and the cloud server 6 are communicatively connected through a wireless or wired network for data transmission and interaction.

[0049] Embodiment 1

[0050] As Figure 2 shown, it is a flowchart of a video optimization method proposed in the first embodiment of this application. It can be understood that the flowchart in the embodiment of this method is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. The following takes the server as the execution subject to illustrate this method.

[0051] This method includes the following steps:

[0052] S200, obtain the video information of the original video according to the cloud video address of the original video.

[0053] In this embodiment, the client uploads the local video (original video file) to the cloud server, and the server obtains the video information of the original video through the online cloud video address URL (Uniform Resource Locator).

[0054] The cloud server creates a data table, such as a MySQL data table tb_post_video_info, and stores the cloud video address URL in the backup_url field for backup. Therefore, it is ensured that the original video can be conveniently searched, and the original video can be continuously optimized multiple times, and the video file can also be restored to the original state.

[0055] Optionally, the video information includes the audio loudness information of the original video file, and obtaining the video information of the original video includes: calculating the audio loudness information of the original video file corresponding to the cloud video address through a filter. Format and output the audio loudness information. In the solution of this embodiment, through filtering and formatting processing of the original video file, the audio loudness information of the original video file can be accurately obtained.

[0056] The server builds an RPC service based on the Go language. In the RPC service, the external service method GetTheMuteTime is used to receive the cloud video address URL parameter. In this method, the FFmpeg command is executed at the system layer by calling the exec.Command() function to obtain video information. Among them, the exec.Command() function is used to execute external commands in the Go language. Specifically, through the FFmpeg command, the video information of the original video file corresponding to the cloud video address can be formatted and output to obtain a corresponding string.

[0057] For example, by adding the -filter_complex ebur128 -c:v copy -f null / dev / null parameter to the FFmpeg command, the corresponding audio information of the original video can be obtained. Among them, the -filter_complex ebur128 command is used to calculate the audio loudness information of the video file using the EBU R128 filter; the -c:v copy command means to keep the video stream unchanged without re-encoding; the -f null / dev / null command means to redirect the output to the null device because only the audio loudness information is concerned at this time and there is no need to actually output to a file.

[0058] S202. Obtain the audio loudness data according to the video information, and analyze the audio stagnation duration of the original video.

[0059] Specifically, further refer to Figure 3 , which is a refined flowchart of the above step S202. It can be understood that this flowchart is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. In this embodiment, the step S202 specifically includes:

[0060] S2020. Match the video information through a preset regular expression to obtain the audio loudness values at multiple time points in the original video.

[0061] In this embodiment, the audio loudness data can be obtained by matching the string of the formatted output of the video information through a corresponding regular expression. Specifically, first, the video information is matched using a first regular expression to obtain multiple time points at preset time intervals in the original video. Then, based on the multiple time points, a second regular expression is used to match the video information to obtain the audio loudness value corresponding to each time point. For example, by using the regular expression t:[\s]*(\d+[\.\d+]*) to match the video information, multiple time points at a preset time interval (e.g., 100 milliseconds) in the video can be obtained; by using the regular expression LRA:\s+(\d+\.\d)\sLU to match the video information, the audio loudness value corresponding to each time point can be obtained.

[0062] S2022. Compare the audio loudness value of each time point with the absolute threshold to obtain the audio loudness state of each time point.

[0063] Loop through and judge the audio loudness state of each of the above time points, that is, compare the audio loudness value of each time point with the absolute threshold of -70 LUFS. If the audio loudness value of a certain time point is lower than -70 LUFS, it is considered that the current time point is in a silent state; if the audio loudness value of a certain time point is higher than -70 LUFS, it is considered that the current time point is in a non-silent state.

[0064] S2024. Determine the audio stagnation duration of the original video according to the audio loudness state of each time point.

[0065] Based on the above judgment, after obtaining the audio loudness state of each time point, that is, the silent state or the non-silent state, the start and end time points of silence in the original video can be obtained, so as to obtain the audio stagnation duration of the original video.

[0066] For example, assume that the first time point to the ninth time point are all in the silent state, and the tenth time point is in the non-silent state. This means that the first time point is the start time point of silence, the tenth time point is the end time point of silence, and the duration between the first time point and the tenth time point is the audio stagnation duration.

[0067] Back to Figure 2 , S204. Analyze the video frame freeze duration of the original video by taking frame screenshots of the original video.

[0068] Specifically, further refer to Figure 4, which is a detailed process schematic diagram of the above step S204. It can be understood that this flowchart is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. In this embodiment, the step S204 specifically includes:

[0069] S2040, perform frame screenshots on the original video at a preset interval.

[0070] In this embodiment, the original video on the cloud server can be frame-sampled at an interval of 200 milliseconds to obtain a plurality of screenshots arranged in time sequence. Specifically, the screenshot information of the original video on the cloud server can be directly obtained through corresponding method parameters without downloading the original video file to the local.

[0071] S2042, perform similarity comparison on the screenshots to determine the screen state.

[0072] After obtaining a plurality of screenshots arranged in time sequence, perform similarity comparison between the front and rear screenshots to determine whether the video screen is stagnant. Among them, when the similarity between two consecutive screenshots is equal to zero, it means that the screen has not changed, that is, the screen is in a static state; when the similarity between two consecutive screenshots is greater than zero, it means that the screen has changed, that is, the non-static state. In this embodiment, the pHash algorithm is used for similarity comparison of screenshots. The pHash algorithm is a recognition algorithm with extremely high accuracy in current image similarity comparison algorithms.

[0073] S2044, determine the screen stagnation duration of the original video according to the screen state.

[0074] Based on the above judgment, the screen state between every two screenshots can be obtained, that is, the static state or the non-static state. Then, the screen stagnation duration in the original video can be obtained according to the screen state.

[0075] For example, assume that the comparison result between the first screenshot and the second screenshot is a static state, the comparison result between the second screenshot and the third screenshot is also a static state, and the comparison result between the third screenshot and the fourth screenshot is a non-static state, that is, the screen has changed. Then, the duration between the first screenshot and the third screenshot is the screen stagnation duration.

[0076] Return to Figure 2 , S206, determine the part of the original video that needs to be cropped according to the audio stagnation duration and the screen stagnation duration, and edit the original video to obtain an optimized video.

[0077] In this embodiment, determining the portion to be cropped in the original video according to the audio stagnation duration and the video stagnation duration may include various processing methods. For example, taking the maximum value of the two stagnation durations, using the video portion corresponding to the maximum stagnation duration as the cropped portion, taking the overlapping portion of the two stagnation durations as the cropped portion, or taking the video portions corresponding to the two stagnation durations as the cropped portions, etc.

[0078] Specifically, as Figure 5 shown, it is the first refined flow diagram of the above step S206. In Figure 5 , the step S206 includes:

[0079] S2060, compare the audio stagnation duration and the video stagnation duration to obtain the maximum stagnation duration among them.

[0080] According to the above steps, one or more of the audio stagnation durations and one or more of the video stagnation durations may be obtained. Compare all these durations and take the maximum value among them, that is, the maximum stagnation duration.

[0081] S2061, use the video portion corresponding to the maximum stagnation duration as the cropped portion, and edit the original video to obtain an optimized video.

[0082] That is to say, the first processing method is to crop the portion with the longest stagnation time of the audio or video in the original video, and the edited video is the optimized video.

[0083] As Figure 6 shown, it is the second refined flow diagram of the above step S206. In Figure 6 , the step S206 includes:

[0084] S2062, determine whether there is an overlapping portion between the audio stagnation duration and the video stagnation duration in the original video.

[0085] For the one or more audio stagnation durations and the one or more video stagnation durations, respectively find the corresponding portions in the original video and determine whether these portions overlap. For example, assume that the 0 - 2400 milliseconds of the original video is the audio stagnation portion, and the 0 - 2800 milliseconds is the video stagnation portion, then the 0 - 2400 milliseconds is the overlapping portion.

[0086] S2063, use the overlapping portion as the cropped portion, and edit the original video to obtain an optimized video.

[0087] That is to say, the second processing method is to crop the portion where both the audio and video are stagnant in the original video, and the edited video is the optimized video.

[0088] As shown Figure 7 in the figure, it is a schematic diagram of the third refined process of the above step S206. In Figure 7 , the step S206 includes:

[0089] S2064, determining whether the total duration of the audio stagnation duration and the video stagnation duration exceeds a preset duration.

[0090] For the one or more audio stagnation durations and one or more video stagnation durations, summarize them, calculate the total duration of the audio stagnation or video stagnation part in the original video. Then compare whether the total duration exceeds the preset duration, so as to determine whether it is necessary to cut off the video parts corresponding to the audio stagnation duration and the video stagnation duration. Among them, the preset duration can be set according to the total duration of the short video. For example, assuming that the original video is 5 minutes in total, the preset duration can be set to 10%, that is, 30 seconds.

[0091] S2065, in the case where the total duration does not exceed the preset duration, take the video parts corresponding to the audio stagnation duration and the video stagnation duration as the cut parts, and edit the original video to obtain an optimized video.

[0092] If the total duration of the audio stagnation or video stagnation part in the original video does not exceed the preset duration, it means that the stagnation part is not too long and can be all cut off. Therefore, the third processing method is to cut off the audio or video stagnation parts in the original video, and the edited video is the optimized video.

[0093] After editing through the above processing method, the video quality can be improved to obtain an optimized video. Then the server uploads the optimized video to the cloud server again, and a new cloud video address URL can be generated. The cloud server can also save the new cloud video address to the MySQL data table.

[0094] In addition, according to actual application needs, some subsequent processing can also be performed on the optimized video. For example, obtain the first frame of the optimized video to generate a video cover image. Or perform watermarking processing on the optimized video, generate gif pictures, etc.

[0095] The video optimization method proposed in this embodiment is developed based on cloud video and RPC services, which onlineizes and processes the optimization process of low-quality short videos. Through the online service of the FFmpeg tool, only by passing the cloud video address to the service interface on the server side can the video information of the original video be obtained, and the audio stagnation duration and the picture stagnation duration can be analyzed, so as to clip the original video. At the same time, the local traffic of the client will not be consumed during the whole process, and all processing is done on the server side, greatly saving the bandwidth cost brought by downloading videos. In addition, no manual participation is required during the whole process, which saves costs and more efficiently completes the optimization of low-quality short videos. The overall quality of the video optimized by this method is significantly improved, which can effectively increase the completion rate and the browsing step length, thus enhancing user stickiness. And this method can be dynamically extended to support adding more optimization parameters and processing logics.

[0096] Embodiment 2

[0097] As Figure 8 shown, it is a flowchart of a video optimization method proposed in the second embodiment of this application. In the second embodiment, on the basis of the above first embodiment, the video optimization method further includes steps S304 and S310. It can be understood that the flowchart in the embodiment of this method is not used to limit the execution order of steps. According to needs, some steps in this flowchart can also be added or deleted.

[0098] This method includes the following steps:

[0099] S300, obtaining the video information of the original video according to the cloud video address of the original video.

[0100] S302, obtaining the audio loudness data according to the video information and analyzing the audio stagnation duration of the original video.

[0101] S304, obtaining the average audio loudness value of the original video according to the video information.

[0102] In this embodiment, the average audio loudness value can also be obtained by matching the string of the formatted output of the video information through the corresponding regular expression. For example, by using the regular expression \s+I:\s+([-+\s]\d+\.\d)LUFS to match the video information, the average audio loudness value of the original video can be obtained.

[0103] S306, analyzing the picture stagnation duration of the original video by taking frame screenshots of the original video.

[0104] S308. Determine the part of the original video that needs to be cropped according to the audio stagnation duration and the video stagnation duration, and clip the original video to obtain an optimized video.

[0105] The implementation of the above steps S300 - S302 and S306 - S308 is the same as the implementation principle of steps S200 - S206 in the foregoing first embodiment. The specific implementation process can refer to the description in the first embodiment, and will not be elaborated herein in the embodiments of the present application.

[0106] S310. Adjust the overall audio loudness of the optimized video according to the average audio loudness value.

[0107] In this embodiment, according to the average audio loudness value, it can be determined whether the overall audio loudness of the original video is too large or too small. Especially when multiple videos are played continuously after the subsequent release of the video, if the audio loudness difference between the two consecutive videos is too large, it is very likely to affect the user's auditory experience and cause a very bad user experience.

[0108] Therefore, in this embodiment, the average audio loudness value can be compared with a preset threshold. If the average audio loudness value is lower than the first threshold, it means that the audio loudness of the original video is too small, and the overall audio loudness of the original video can be increased by a preset ratio. If the average audio loudness value is higher than the second threshold, it means that the audio loudness of the original video is too large, and the overall audio loudness of the original video can be decreased by a preset ratio. Wherein, the second threshold is greater than the first threshold. Additionally, the overall audio loudness of the original video can also be adjusted to an intermediate value.

[0109] It should be noted that for the same original video, both the video clipping and the audio loudness adjustment can be performed, or only one of the optimization processing methods can be performed. If the original video has undergone the video clipping process, the average audio loudness value after clipping can be obtained according to the video information after clipping, and then compared and adjusted.

[0110] In addition, in an optional embodiment, at this time, only the average audio loudness value can be recorded. For example, the average audio loudness value is uploaded and saved in the MySQL data table of the cloud server. When multiple videos are played continuously for users after the subsequent video distribution, the audio loudness difference between the two consecutive videos is calculated according to the recorded average audio loudness value, so as to perform an overall audio loudness adjustment on the released video.

[0111] After clipping and adjusting the audio loudness through the above processing method, the video quality can be improved to obtain an optimized video. Then the server uploads the optimized video to the cloud server again, and a new cloud video address URL can be generated. The cloud server can also save the new cloud video address to the MySQL data table.

[0112] In addition, according to actual application needs, some subsequent processing can also be performed on the optimized video. For example, obtaining the first frame of the optimized video to generate a video cover image. Another example is to perform watermarking on the optimized video, generate gif pictures, etc.

[0113] As Figure 9 shown, it is a schematic flowchart of another form of the video optimization method of this embodiment. Figure 9 The specific implementation processes of each step in

[0114] The video optimization method proposed in this embodiment only needs to pass the cloud video address to the service interface of the server to obtain the video information of the original video, analyze the audio stagnation duration and the video frame stagnation duration, and then clip the original video. At the same time, the local traffic of the client will not be consumed during the whole process, and all processing is done on the server, which greatly saves the bandwidth cost brought by downloading videos. In addition, no manual participation is required during the whole process, which saves costs and more efficiently completes the optimization of low-quality short videos. Through the optimization process of this solution, the overall quality of the video is significantly improved, which can effectively increase the completion rate and the browsing step length, thereby enhancing user stickiness. And by obtaining the average audio loudness value of the original video, it is possible to adjust the situation where the overall audio loudness of the video is too large or too small, so that the volume will not be too high or too low when multiple videos are played continuously, improving the user's auditory experience and further enhancing the user experience and user stickiness.

[0115] Embodiment III

[0116] As Figure 10 shown, it is a schematic diagram of the hardware architecture of an electronic device 20 proposed in the third embodiment of the present application. In this embodiment, the electronic device 20 may include, but is not limited to, a memory 21, a processor 22, and a network interface 23 that are communicatively connected to each other through a system bus. It should be noted that Figure 10 only the electronic device 20 with components 21 - 23 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. In this embodiment, the electronic device 20 may be the server of the server side.

[0117] The memory 21 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 21 may be an internal storage unit of the electronic device 20, such as the hard disk or memory of the electronic device 20. In other embodiments, the memory 21 may also be an external storage device of the electronic device 20, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 20. Of course, the memory 21 may also include both the internal storage unit and the external storage device of the electronic device 20. In this embodiment, the memory 21 is generally used to store the operating system and various application software installed in the electronic device 20, such as the program code of the video optimization system 60, etc. In addition, the memory 21 may also be used to temporarily store various data that have been output or will be output.

[0118] In some embodiments, the processor 22 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 22 is generally used to control the overall operation of the electronic device 20. In this embodiment, the processor 22 is used to run the program code stored in the memory 21 or process data, such as running the video optimization system 60, etc.

[0119] The network interface 23 may include a wireless network interface or a wired network interface, and the network interface 23 is generally used to establish a communication connection between the electronic device 20 and other electronic devices.

[0120] Embodiment Four

[0121] As Figure 11 shown, a schematic diagram of modules of a video optimization system 60 is proposed in the fourth embodiment of this application. The video optimization system 60 may be divided into one or more program modules, and one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of this application. The program modules referred to in the embodiments of this application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment.

[0122] In this embodiment, the video optimization system 60 includes:

[0123] A receiving module 600, configured to obtain video information of the original video according to the cloud video address of the original video.

[0124] In this embodiment, the client uploads a local video (original video file) to the cloud server to obtain an online cloud video address URL.

[0125] The cloud server creates a data table, such as a MySQL data table tb_post_video_info, and stores the cloud video address URL in the backup_url field for backup. Therefore, it is ensured that the original video can be conveniently searched, and the original video can be continuously optimized multiple times, and the video file can also be restored to its original state.

[0126] The receiving module 600 receives the cloud video address URL parameter through the external service method GetTheMuteTime in the RPC service. In this method, the FFmpeg command is executed at the system layer by calling the exec.Command() function to obtain video information. Specifically, through the FFmpeg command, the video information of the original video file corresponding to the cloud video address can be formatted and output to obtain a corresponding string. In this embodiment, the video information mainly includes the audio loudness information of the original video file.

[0127] An analysis module 602, configured to obtain audio loudness data according to the video information and analyze the audio stagnation duration of the original video.

[0128] First, the video information is matched through a preset regular expression to obtain the audio loudness values at multiple time points in the original video. Then, the audio loudness value at each time point is compared with an absolute threshold to obtain the audio loudness state at each time point, including the mute state or the non-mute state. Then, the audio stagnation duration of the original video is determined according to the audio loudness state at each time point.

[0129] The analysis module 602 is further configured to analyze the picture stagnation duration of the original video by taking frame screenshots of the original video.

[0130] First, frame screenshots of the original video are taken at a preset interval. Then, the similarity of the screenshots is compared to determine the picture state, including the static state or the non-static state. Then, the picture stagnation duration of the original video is determined according to the picture state.

[0131] A clip module 604, configured to determine the part to be cropped in the original video according to the audio stagnation duration and the picture stagnation duration, and clip the original video to obtain an optimized video.

[0132] In this embodiment, determining the part to be cropped in the original video according to the audio stagnation duration and the video stagnation duration may include various processing methods. For example, taking the maximum value of the two stagnation durations and using the video part corresponding to the maximum stagnation duration as the cropped part, taking the overlapping part of the two stagnation durations as the cropped part, or taking the video parts corresponding to the two stagnation durations as the cropped parts, etc.

[0133] After editing through the above processing methods, the video quality can be improved to obtain an optimized video. Re-uploading the optimized video to the cloud server can generate a new cloud video address URL. The cloud server can also save the new cloud video address to the MySQL data table.

[0134] In addition, according to actual application needs, some subsequent processing can also be performed on the optimized video. For example, obtaining the first frame of the optimized video to generate a video cover image. Another example is to perform watermarking processing on the optimized video, generate a gif image, etc.

[0135] For the specific implementation process of the functions of each of the above modules, reference can be made to the description in the first embodiment above, and details will not be elaborated here.

[0136] Embodiment Five

[0137] As Figure 12 shown, a schematic diagram of modules of a video optimization system 60 proposed in the fifth embodiment of the present application is shown. In this embodiment, in addition to including the receiving module 600, the analysis module 602, and the editing module 604 in the fourth embodiment, the video optimization system 60 further includes an acquisition module 606 and an adjustment module 608.

[0138] The acquisition module 606 is configured to obtain the average audio loudness value of the original video according to the video information.

[0139] In this embodiment, the average audio loudness value can be obtained by matching the string of the formatted output of the video information through a corresponding regular expression.

[0140] The adjustment module 608 is configured to adjust the overall audio loudness of the optimized video according to the average audio loudness value.

[0141] In this embodiment, the average audio loudness value can be compared with a preset threshold. If the average audio loudness value is lower than the first threshold, it indicates that the audio loudness of the original video is relatively small, and the overall audio loudness of the original video can be increased by a preset proportion. If the average audio loudness value is higher than the second threshold, it indicates that the audio loudness of the original video is relatively large, and the overall audio loudness of the original video can be decreased by a preset proportion. Wherein, the second threshold is greater than the first threshold. Additionally, the overall audio loudness of the original video can also be adjusted to an intermediate value.

[0142] For the specific implementation process of the functions of each of the above modules, reference can be made to the description in the second embodiment above, which will not be elaborated here.

[0143] Embodiment Six

[0144] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium, which stores a video optimization program, and the video optimization program can be executed by at least one processor, so that the at least one processor executes the steps of the video optimization method as described above.

[0145] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device. Of course, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the video optimization method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.

[0146] It should be noted that, in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element.

[0147] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0148] Obviously, those skilled in the art should understand that the various modules or steps of the embodiments of the present application described above can be implemented by a general-purpose computing device. They can be centralized on a single computing device or distributed across a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0149] The above are only the preferred embodiments of the embodiments of the present application, and do not thus limit the patent scope of the embodiments of the present application. Any equivalent structural or equivalent process transformation made by using the description and drawings of the embodiments of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the embodiments of the present application.

Claims

1. A video optimization method, applied to a server, characterized in that, The method includes: Obtaining video information of the original video according to the cloud video address of the original video; Obtaining audio loudness data according to the video information and analyzing the audio stagnation duration of the original video; Analyzing the picture stagnation duration of the original video by taking frame screenshots of the original video; Determining the part to be cropped in the original video according to the audio stagnation duration and the picture stagnation duration, and editing the original video to obtain an optimized video.

2. The video optimization method according to claim 1, wherein The method further includes: Obtaining the average audio loudness value of the original video according to the video information; Adjusting the overall audio loudness of the optimized video according to the average audio loudness value.

3. The video optimization method according to claim 1, wherein The obtaining video information of the original video according to the cloud video address of the original video includes: Calculating the audio loudness information of the original video corresponding to the cloud video address through a filter; Formatting and outputting the audio loudness information.

4. The video optimization method according to claim 1 or 3, characterized in that, The obtaining audio loudness data according to the video information and analyzing the audio stagnation duration of the original video includes: Matching the video information through a preset regular expression to obtain the audio loudness values at multiple time points in the original video; Comparing the audio loudness value at each time point with an absolute threshold to obtain the audio loudness state at each time point, including a mute state or a non-mute state; Determining the audio stagnation duration of the original video according to the audio loudness state at each time point.

5. The video optimization method according to claim 4, wherein The matching the video information through a preset regular expression to obtain the audio loudness values at multiple time points in the original video includes: Using a first regular expression to match the video information to obtain multiple time points at preset time intervals in the original video; Based on the multiple time points, using a second regular expression to match the video information to obtain the audio loudness value corresponding to each time point among the multiple time points.

6. The video optimization method according to claim 1, wherein The analyzing the picture stagnation duration of the original video by taking frame screenshots of the original video includes: Taking frame screenshots of the original video at a preset interval to obtain multiple screenshots arranged in time sequence; Comparing the similarity of the multiple screenshots to determine the picture state, including a static state or a non-static state; Determining the picture stagnation duration of the original video according to the picture state.

7. The video optimization method according to claim 1, wherein The determining the part to be cropped in the original video according to the audio stagnation duration and the picture stagnation duration includes: Comparing the audio stagnation duration and the picture stagnation duration to obtain the maximum stagnation duration therein; Taking the video part corresponding to the maximum stagnation duration as the cropped part; Or Judging whether there is an overlapping part between the audio stagnation duration and the picture stagnation duration in the original video; Taking the overlapping part as the cropped part; Or Judging whether the total duration of the audio stagnation duration and the picture stagnation duration exceeds a preset duration; In the case where the total duration does not exceed the preset duration, taking the video parts corresponding to the audio stagnation duration and the picture stagnation duration as the cropped parts.

8. The video optimization method according to claim 2, wherein The adjusting the overall audio loudness of the optimized video according to the average audio loudness value includes: In the case where the average audio loudness value is lower than the first threshold, increase the overall audio loudness of the original video by a preset ratio; In the case where the average audio loudness value is higher than the second threshold, decrease the overall audio loudness of the original video by a preset ratio, where the second threshold is greater than the first threshold.

9. A video optimization system, applied to a server, characterized in that The system includes: A receiving module, configured to obtain video information of the original video according to the cloud video address of the original video; An analysis module, configured to obtain audio loudness data according to the video information and analyze the audio stagnation duration of the original video; The analysis module is further configured to analyze the picture stagnation duration of the original video by taking frame screenshots of the original video; A clip module, configured to determine the part of the original video that needs to be cropped according to the audio stagnation duration and the picture stagnation duration, and clip the original video to obtain an optimized video.

10. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a video optimization program stored on the memory and executable on the processor. When the video optimization program is executed by the processor, it implements the video optimization method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, A video optimization program is stored on the computer-readable storage medium. When the video optimization program is executed by the processor, it implements the video optimization method according to any one of claims 1 to 8.