Loudness processing method, apparatus and terminal device
By identifying the loudness of video segments through terminal devices or utilizing cache, the problem of low efficiency in uniform loudness processing is solved, achieving more efficient and accurate loudness adjustment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-27
- Publication Date
- 2026-05-29
AI Technical Summary
The efficiency of loudness uniform processing in existing technologies is relatively low, mainly because terminal devices need to calculate the loudness of the video material corresponding to each video segment in real time, resulting in slow processing speed.
The terminal device acquires multiple video segments from the video editing draft, identifies or determines the first loudness of each video segment based on the cache, and processes them according to these loudnesses to ensure that the loudness of all video segments is the same, thus avoiding repeated calculations of the video materials.
It improves the efficiency and accuracy of loudness calculation, reduces the impact of invalid segments on processing, and enhances the efficiency and accuracy of unified loudness processing.
Smart Images

Figure CN122120544A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video editing technology, and in particular to a loudness processing method, apparatus, and terminal device. Background Technology
[0002] Loudness unification processing can adjust the loudness of video clips in a video editing draft to the same loudness. Therefore, the terminal device needs to calculate the loudness of each video clip.
[0003] Currently, terminal devices can calculate the loudness of the video clip corresponding to the video material in real time and determine the loudness of the video material as the loudness of the video clip. However, in the above method, the calculation speed of the video clip's loudness is slow, resulting in low efficiency in unified loudness processing. Summary of the Invention
[0004] This disclosure provides a loudness processing method, apparatus, and terminal device to solve the technical problem of low efficiency in the uniform loudness processing of the prior art.
[0005] In a first aspect, embodiments of this disclosure provide a loudness processing method, the loudness processing method comprising:
[0006] Obtain a video editing draft, which includes multiple video clips formed based on video footage;
[0007] A first loudness is determined for the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache;
[0008] The loudness of the video editing draft is processed based on the first loudness of the plurality of video clips so that the loudness of the plurality of video clips is the same.
[0009] Secondly, embodiments of this disclosure provide a loudness processing apparatus, which includes an acquisition module, a determination module, and a processing module, wherein:
[0010] The acquisition module is used to acquire a video editing draft, which includes multiple video clips formed based on video materials.
[0011] The determining module is used to determine the first loudness of the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache;
[0012] The processing module is used to process the loudness of the video editing draft according to the first loudness of the plurality of video segments, so that the loudness of the plurality of video segments is the same.
[0013] Thirdly, this disclosure provides a terminal device including: a processor and a memory;
[0014] The memory stores computer-executed instructions;
[0015] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the loudness processing methods described in the first aspect above and various possible aspects of the first aspect.
[0016] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the loudness processing methods described in the first aspect and various possible aspects of the first aspect.
[0017] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the loudness processing methods described in the first aspect above and various possible aspects of the first aspect.
[0018] This disclosure provides a loudness processing method, apparatus, and terminal device. The terminal device can acquire a video editing draft, which may include multiple video segments formed based on video footage. The terminal device can determine the first loudness of the multiple video segments. The first loudness of the multiple video segments can be obtained by identifying each video segment separately, or based on the loudness of the video footage corresponding to the video segments determined by a cache. The terminal device can process the loudness of the video editing draft based on the first loudness of the multiple video segments to ensure that the loudness of the multiple video segments is the same. In the above method, since the terminal device can calculate the first loudness of each video segment separately without calculating the loudness of the video footage, the calculation efficiency of the first loudness can be improved. Alternatively, the terminal device can reuse the loudness of the video footage stored in the cache, thus improving the efficiency of determining the loudness of the video segments. This improves the efficiency and accuracy of unified loudness processing. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0021] Figure 2 A flowchart illustrating a loudness processing method provided in an embodiment of this disclosure;
[0022] Figure 3 This is a schematic diagram illustrating a process for determining a valid segment, provided in an embodiment of the present disclosure.
[0023] Figure 4 A schematic diagram of the first loudness provided for an embodiment of this disclosure;
[0024] Figure 5 A schematic diagram illustrating the process of storing the loudness of video material according to an embodiment of this disclosure;
[0025] Figure 6 This is a schematic diagram of a method for determining a first loudness according to an embodiment of the present disclosure;
[0026] Figure 7 A schematic diagram illustrating another method for determining a first loudness provided in an embodiment of this disclosure;
[0027] Figure 8 This is a schematic diagram of the structure of a loudness processing device provided in an embodiment of the present disclosure;
[0028] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this disclosure. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0031] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0034] For ease of understanding, the concepts involved in the embodiments of this disclosure will be explained below.
[0035] Terminal equipment: A device with wireless transceiver capabilities. Terminal equipment can be deployed on land, including indoors or outdoors, handheld, wearable, or vehicle-mounted. The terminal equipment can be a mobile phone, tablet, computer with wireless transceiver capabilities, virtual reality (VR) terminal equipment, augmented reality (AR) terminal equipment, wireless terminals in industrial control, vehicle-mounted terminal equipment, wireless terminals in self-driving vehicles, wireless terminal equipment in remote medical care, wireless terminal equipment in smart grids, wireless terminal equipment in transportation safety, wireless terminal equipment in smart cities, wireless terminal equipment in smart homes, wearable terminal equipment, etc. The terminal equipment involved in the embodiments of this disclosure can also be referred to as a terminal, user equipment (UE), access terminal equipment, vehicle-mounted terminal, industrial control terminal, UE unit, UE station, mobile station, mobile station, remote station, remote terminal equipment, mobile device, UE terminal equipment, wireless communication equipment, UE agent, or UE device, etc. Terminal devices can be fixed or mobile.
[0036] Below, in conjunction with Figure 1The application scenarios of the embodiments of this disclosure will be described.
[0037] Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this disclosure. Please refer to [link / reference]. Figure 1 Including: terminal equipment ( Figure 1 The video editing page (not shown) may include a preview area, a video track, and a loudness unification control. The video track includes video clip 1, video clip 2, and video clip 3. The loudness of video clip 1 is loudness a, the loudness of video clip 2 is loudness b, and the loudness of video clip 3 is loudness c.
[0038] Please see Figure 1 When a user clicks the loudness unification control, the terminal device can determine the loudness d based on loudness a, loudness b, and loudness c, and adjust the loudness of video clip 1, video clip 2, and video clip 3 to loudness d. This eliminates the need for the user to manually adjust the loudness of each video clip, thus reducing the complexity of user operations.
[0039] It should be noted that, Figure 1 This is an example of an application scenario for the embodiments of this disclosure, and is not intended to limit the application scenarios of the embodiments of this disclosure.
[0040] In related technologies, since loudness unification processing can adjust the loudness of video clips in a video editing draft to the same loudness, the terminal device needs to calculate the loudness of each video clip. Currently, the terminal device can calculate the loudness of the video material corresponding to the video clip in real time to obtain the loudness of the video clip. For example, the video editing draft includes video clip 1, video clip 2, and video clip 3, where video clip 1 and video clip 2 belong to video material A, and video clip 3 belongs to video material B. When the terminal device calculates the loudness of video clip 1, it can calculate the overall loudness of video material A to obtain the loudness of video clip 1. When the terminal device calculates the loudness of video clip 2, it can again calculate the overall loudness of video material A to obtain the loudness of video clip 2. When the terminal device calculates the loudness of video clip 3, it can calculate the overall loudness of video material B to obtain the loudness of video clip 3. However, when there are a large number of video clips, the terminal device needs to calculate the loudness of the corresponding video material when determining the loudness of each video clip, which results in low efficiency of unified loudness processing. In addition, the video material also includes some video clips that do not need to be edited, which leads to low accuracy of unified loudness processing.
[0041] To address the technical problems in related technologies, this disclosure provides a loudness processing method. A terminal device can acquire a video editing draft, which includes multiple video segments formed based on video materials. The terminal device can determine the duration percentage of the multiple video segments in the video materials and the number of video segments that are identical to the original video materials. Based on the duration percentage and the number of video segments that are identical to the original video materials, the terminal device determines the first loudness of the multiple video segments. The terminal device can process the loudness of the video editing draft based on the first loudness of the multiple video segments to ensure that the loudness of the multiple video segments is the same.
[0042] In this way, since the terminal device can calculate the first loudness of video segments separately without calculating the loudness of the video material, the accuracy of loudness calculation is improved, as is the accuracy of unified loudness processing. Alternatively, the terminal device can reuse the loudness of video material stored in the cache. Therefore, the terminal device can quickly determine the first loudness of video segments, improving the efficiency of loudness calculation and the efficiency of unified loudness processing.
[0043] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.
[0044] Figure 2 This is a schematic flowchart illustrating a loudness processing method provided in an embodiment of this disclosure. Please refer to [link / reference]. Figure 2 The method may include:
[0045] S201. Obtain the video editing draft.
[0046] The execution entity of this disclosure embodiment can be a terminal device or a loudness processing device installed in the terminal device. The loudness processing device can be implemented in software, or it can be implemented using a combination of software and hardware; this disclosure embodiment does not limit this approach.
[0047] A video editing draft can include multiple video clips formed based on video footage. For example, a video editing draft can be a preliminary or incomplete version during the video editing process. A video editing draft can include multiple video clips formed based on at least one video clip (which can be any video clip such as a live stream, a recorded video, or a movie; this disclosure does not limit this). In this way, the terminal device can add a video editing draft on the video editing page and edit multiple video clips. For example, a video editing draft can be a temporary stage in the video editing process, and can be used to record the current progress of video editing, user editing operations on the video clips in the video editing draft, etc. For example, a video editing draft can include video clip 1, video clip 2, and video clip 3, where video clip 1 and video clip 2 are video clips formed based on video footage A, and video clip 3 is a video clip formed based on video footage B.
[0048] Optionally, the video editing draft may also include segments based on audio materials, image materials, etc., which are not limited in this embodiment.
[0049] The terminal device can respond to user actions and retrieve video editing drafts. For example, when a user clicks on a video editing application, the terminal device can launch the application and display its main page. This main page can include a draft area containing one or more video editing drafts. When the user clicks on any of these drafts, the terminal device can display the video editing page, showing the selected draft.
[0050] Optionally, the terminal device may also obtain the video editing draft according to any feasible implementation method (e.g., the terminal device may receive the video editing draft sent by the server), and this embodiment of the disclosure does not limit this.
[0051] The video clips can be segments created from video footage. For example, if the video footage is 10 minutes long, 10 video clips can be generated from it, each clip being 1 minute long. Alternatively, a video clip can be a valid segment from the video footage, where a valid segment can be a segment to be edited. For instance, if a user extracts at least one video clip from the video footage and edits it, that clip is considered a valid clip. Unextracted segments from the video footage are considered invalid clips. Invalid clips will affect the accuracy of loudness unification calculations.
[0052] Optionally, multiple video clips may correspond to one video clip or multiple video clips; this disclosure does not limit this. For example, a video editing draft may include five video clips, all of which come from the same video clip; or, a video editing draft may include two video clips, one of which comes from one video clip and the other from another video clip.
[0053] Below, in conjunction with Figure 3 The effective segments in the embodiments of this disclosure will be described.
[0054] Figure 3 This is a schematic diagram illustrating a process for determining a valid segment according to an embodiment of this disclosure. Please refer to... Figure 3 This includes video footage. The video footage includes video clips 1, 2, 3, 4, and 5. Terminal devices ( Figure 3 (Not shown) A new video is cut from the existing video footage, where the new video includes video clip 1, video clip 3, and video clip 5. The terminal device can determine that valid clips include video clip 1, video clip 3, and video clip 5, and invalid clips include video clip 2 and video clip 4.
[0055] It should be noted that, Figure 3 The embodiments shown are examples of determining valid segments. It is not limited to generating a new video before determining valid segments. During the video editing process, video segments displayed on the video editing page can be valid segments, while video segments not displayed on the video editing page can be invalid segments.
[0056] S202, Determine the first loudness of multiple video clips.
[0057] Optionally, the first loudness of multiple video clips can be the loudness obtained by recognizing each video clip separately. For example, if the video editing draft includes video clip 1 and video clip 2, then the first loudness of video clip 1 is the loudness obtained by the terminal device recognizing the loudness of video clip 1, and the first loudness of video clip 2 is the loudness obtained by the terminal device recognizing the loudness of video clip 2.
[0058] For example, if a video clip is 10 seconds long, its first loudness can be obtained by loudness identification of that 10-second clip; if the video clip is 20 seconds long, its first loudness can be obtained by loudness identification of that 20-second clip. In this way, the terminal device does not need to calculate the overall loudness of the video material, and invalid segments will not affect the unified loudness processing, thus significantly improving the efficiency and accuracy of the unified loudness processing.
[0059] It should be noted that the terminal device can identify the loudness of video segments according to any feasible implementation method, and the embodiments disclosed herein are not limited in this respect.
[0060] Below, in conjunction with Figure 4 The first loudness level will be explained.
[0061] Figure 4 A schematic diagram illustrating a first loudness provided for an embodiment of this disclosure. Please refer to... Figure 4 This includes video footage. The video footage includes 5-second clips, 10-second clips, and 7-second clips. If the video clip is 10 seconds long, then the terminal device ( Figure 4 (Not shown) The first loudness of a 10-second segment can be determined based on its waveform and amplitude. In this way, the terminal device does not need to determine the loudness of the 10-second segment based on the waveform and amplitude of the entire video material (22 seconds of video), thus improving the efficiency of determining the first loudness.
[0062] Optionally, the first loudness of multiple video clips can be the loudness of the video clip corresponding to the video clip determined based on the cache. For example, the first loudness of a video clip can be the loudness of the video clip corresponding to the video clip, where the loudness of the video clip can be obtained from the cache. For example, the cache can store the loudness of at least one video clip. When the terminal device determines the first loudness of a video clip, it can obtain the loudness of the video clip corresponding to that video clip from the cache to obtain the first loudness of the video clip. In this way, the terminal device can reuse the loudness of the video clip in the cache, improving the efficiency of determining the first loudness of a video clip.
[0063] Optionally, the terminal device may pre-store the loudness of at least one video clip in a cache. For example, a video editing draft may include multiple video segments formed based on video clip 1 and multiple segments formed based on video clip 2. The terminal device may pre-determine the loudness of video clip 1 and the loudness of video clip 2, and store the loudness of video clip 1 and video clip 2 in a cache.
[0064] Optionally, the terminal device can determine the loudness of the video clip corresponding to the video clip and store the loudness of the video clip in a cache. For example, when determining the first loudness of a video clip, the terminal device can determine the loudness of the video clip corresponding to the video clip, thereby obtaining the first loudness of the video clip, and storing the loudness of the video clip in the cache. In this way, the terminal device can store the loudness of each video clip in the cache, saving device resources and improving the efficiency of loudness calculation.
[0065] Optionally, the cache may also include the identifier of the video material. For example, if the cache does not include the loudness of the video material, the terminal device may store the loudness of the video material and the identifier of the video material in the cache. The identifier of the video material may be a unique identifier such as the path of the video material, which is not limited in this embodiment. For example, the path of the video material may be the storage path of the file corresponding to the video material in the terminal device, where the path of the video material may be the unique identifier of the video material.
[0066] Optionally, the terminal device may determine the path of the video material according to any feasible implementation method, and this embodiment of the disclosure does not limit this.
[0067] For example, the cache may include the loudness of video clip 1, the path of video clip 1, the loudness of video clip 2, and the path of video clip 2. If the path of the video clip corresponding to the i-th video segment in the video editing draft is path A, it means that the cache includes the loudness of the video clip, and the terminal device can obtain the loudness of the i-th video segment from the cache. If the path of the video clip corresponding to the i-th video segment is path C, it means that the cache does not include the loudness of the video clip corresponding to the i-th video segment, and the terminal device can determine the loudness of the video clip corresponding to the i-th video segment and store the loudness and path C of the video clip corresponding to the i-th video segment in the cache.
[0068] Below, in conjunction with Figure 5 This section explains the process of storing the loudness of video footage in the cache.
[0069] Figure 5 This is a schematic diagram illustrating a process for storing the loudness of video footage, as provided in an embodiment of this disclosure. Please refer to [link / reference]. Figure 5 This includes: video clip A, video clip B, and cache. Video clip A includes video segment 1, video clip B includes video segments 2 and 3, and the cache is empty. Because the cache is empty, the terminal device ( Figure 5 (Not shown) The loudness of video material A can be determined, loudness a can be obtained, and the loudness of video segment 1 can be determined as loudness a. The terminal device can store the path-loudness a of video material A in the cache.
[0070] Please see Figure 5 When the terminal device determines the loudness of video clip 2, since the cache does not include the path of video material B, the terminal device can determine the loudness of video material B as loudness b, and thus determine the loudness of video clip 2 as loudness b. The terminal device can store the path and loudness b of video material B in the cache.
[0071] Please see Figure 5When the terminal device determines the loudness of video clip 3, since the cache includes the path of video material B (the video material corresponding to video clip 3), the terminal device can determine the loudness of video clip 3 as the loudness of the cached loudness b (the loudness corresponding to the path of video material B). In this way, the terminal device can store the path of each video material in the cache sequentially, saving device resources. Furthermore, the terminal device can reuse the loudness of video materials in the cache, improving the efficiency of determining the loudness of video clips.
[0072] S203. Based on the first loudness of multiple video clips, process the loudness of the video editing draft to ensure that the loudness of multiple video clips is the same.
[0073] Specifically, the terminal device can adjust the loudness of multiple video clips in a video editing draft to the same loudness based on their initial loudness. For example, if video clip 1 has a loudness of loudness 'a', video clip 2 has a loudness of loudness 'b', and video clip 3 has a loudness of loudness 'c', where loudness 'a', loudness 'b', and loudness 'c' are different, the terminal device can unify the loudness of video clips 1, 2, and 3 based on loudness 'a', loudness 'b', and loudness 'c'. After the loudness unification process, the loudness of video clip 1, video clip 2, and video clip 3 will be the same.
[0074] Optionally, the terminal device processes the loudness of the video editing draft based on the first loudness of multiple video clips. Specifically, this can involve: determining the average of the first loudness of the multiple video clips to obtain a second loudness, and adjusting the loudness of the multiple video clips in the video editing draft to the second loudness. For example, if the loudness of video clip 1 is loudness a, the loudness of video clip 2 is loudness b, and the loudness of video clip 3 is loudness c, the terminal device can calculate the average of loudness a, loudness b, and loudness c to obtain loudness d (second loudness), and adjust the loudness of video clip 1, video clip 2, and video clip 3 to loudness d.
[0075] Optionally, the terminal device may also process multiple first loudnesses to obtain a second loudness according to any feasible implementation method (e.g., determining the loudness of the largest among multiple first loudnesses as the second loudness, or determining the loudness of the smallest among multiple first loudnesses as the second loudness, etc.). This disclosure does not limit this.
[0076] This disclosure provides a loudness processing method. A terminal device can acquire a video editing draft and determine the first loudness of multiple video segments in the draft. The terminal device can process the loudness of the video editing draft based on the first loudness of the multiple video segments to ensure that the loudness of the multiple video segments is the same. In the above method, since the first loudness of the multiple video segments is obtained by recognizing each video segment separately, or the terminal device determines the loudness of the video material corresponding to the video segment based on a cache, the terminal device can calculate the first loudness of each video segment separately without calculating the loudness of the video material. Therefore, the calculation efficiency of the first loudness can be improved. Alternatively, the terminal device can reuse the loudness of the video material stored in the cache, thereby improving the efficiency of determining the loudness of the video segments. This improves both the efficiency and accuracy of unified loudness processing.
[0077] exist Figure 2 Based on the embodiments shown, the following, in conjunction with Figure 6 The method for determining the first loudness of multiple video segments in the above loudness processing method will be explained.
[0078] Figure 6 This is a schematic diagram illustrating a method for determining a first loudness according to an embodiment of this disclosure. Please refer to [link / reference]. Figure 6 The method process includes:
[0079] S601. Determine the duration percentage of multiple video clips in the video material, and the number of video clips that belong to the same video material.
[0080] The duration percentage can refer to the proportion of multiple video clips within a single video source. For example, a video editing draft may contain 10 video clips belonging to the same source material. If the source material is 100 minutes long and each clip is 1 minute long, the terminal device can determine the duration percentage as 10%. If the source material is 100 minutes long and each clip is 5 minutes long, the terminal device can determine the duration percentage as 50%. Similarly, a video editing draft may contain 10 video clips, with 8 clips corresponding to video source 1 and 2 clips corresponding to video source 2. If video source 1 is 80 minutes long and video source 2 is 20 minutes long, with each clip being 1 minute long, the terminal device can determine the duration percentage as 10%. If video source 1 is 30 minutes long and video source 2 is 20 minutes long, with each clip being 2 minutes long, the terminal device can determine the duration percentage as 40%.
[0081] It should be noted that the terminal device can also determine the duration ratio of multiple video segments in the video material according to any feasible implementation method, and this disclosure embodiment does not limit this.
[0082] The number of video clips belonging to the same video source can be the same as the number of video clips belonging to the same video source. For example, if a video editing draft includes 10 video clips, and these 10 video clips belong to the same video source, the terminal device can determine that the number of video clips belonging to the same video source is 10. Similarly, if a video editing draft includes 30 video clips, and 10 video clips belong to one video source, while the remaining 20 video clips belong to another video source, the terminal device can determine that the number of video clips belonging to the same video source is 30.
[0083] It should be noted that the terminal device can also determine the number of video segments with the same video material according to any feasible implementation method, and this disclosure does not limit this.
[0084] S602. Determine multiple first loudness levels based on the duration ratio and the number of video clips with the same video material.
[0085] Among them, the terminal device determines multiple first loudnesses based on the duration ratio and the number of video clips with the same video material, and there are two cases:
[0086] Case 1: The duration percentage is less than or equal to the first threshold, or the number of video clips with the same video material is less than or equal to the second threshold.
[0087] Specifically, if the duration percentage is less than or equal to a first threshold, or if the number of video segments with the same video material is less than or equal to a second threshold, then multiple first loudness scores are obtained by identifying each video segment within the multiple video segments. For example, the first threshold can be 50%. If the duration percentage is less than or equal to 50%, it indicates that the video segment accounts for a small proportion of the video material. Therefore, the terminal device can identify the loudness of multiple video segments and obtain the first loudness scores for multiple video segments. In this way, invalid segments will not affect the loudness of video segments, improving the accuracy of the first loudness scores. For example, the second threshold can be 2. If the number of video segments with the same video material is less than or equal to 2, it indicates that the number of video segments with the same video material is small. Therefore, the terminal device can identify the loudness of multiple video segments and obtain the first loudness scores for multiple video segments. Since the number of video segments is small, the efficiency of calculating the first loudness scores can be improved.
[0088] In scenario 1, the experiment was conducted with 5 video clips. If each video clip is completely identified, it would take 5 video clips to be identified, each taking about 26 seconds, for a total of 130 seconds. If the loudness of each video clip is identified, it would take 5 video clips, for a total of about 115 seconds. This can improve the efficiency of determining the first loudness.
[0089] Case 2: The duration percentage is greater than the first threshold, and the number of video clips with the same video material is greater than the second threshold.
[0090] If the duration percentage is greater than the first threshold, and the number of video segments with the same video material is greater than the second threshold, then the first loudness of multiple video segments is determined based on the cache. For example, the first threshold can be 50%, and the second threshold can be 2. If the duration percentage is greater than 50%, and the number of video segments with the same video material is greater than 2, it indicates that the video segment occupies a large proportion of the duration in the video material, and the number of video segments is large. Therefore, the terminal device can determine the loudness of the video material as the first loudness of the video segment. Furthermore, the terminal device can obtain the loudness of the video material from the cache. In this way, by reusing the loudness of the video material in the cache, the first loudness of the video segment can be quickly determined, improving the efficiency of determining the first loudness.
[0091] This disclosure provides a method for unified loudness processing. A terminal device can determine the duration percentage of multiple video segments within a video clip and the number of video segments belonging to the same video clip. If the duration percentage is less than or equal to a first threshold, or the number of video segments belonging to the same video clip is less than or equal to a second threshold, then each video segment in the multiple video clips is identified to obtain multiple first loudnesses. If the duration percentage is greater than the first threshold, and the number of video segments belonging to the same video clip is greater than the second threshold, then the first loudness of the multiple video segments is determined based on a cache. In this way, the terminal device can quickly, flexibly, and accurately determine the first loudness based on the duration percentage and the number of video segments belonging to the same video clip, thereby improving the efficiency and accuracy of determining the first loudness and increasing the flexibility of the determination.
[0092] Based on any of the above embodiments, the following, in conjunction with Figure 7 The method for determining the first loudness of multiple video segments based on the cache in the above loudness processing method will be explained.
[0093] Figure 7 A schematic diagram illustrating another method for determining a first loudness provided in an embodiment of this disclosure. Please refer to... Figure 7 The method process includes:
[0094] S701. Determine the video material corresponding to the video segment.
[0095] Optionally, the terminal device may determine the video material corresponding to the video segment according to any feasible implementation method, and this embodiment of the disclosure does not limit this.
[0096] S702. Determine the first loudness of the video segment based on the video material and cache corresponding to the video segment.
[0097] Among them, Figure 7 In the illustrated embodiment, for any given video segment, the terminal device can determine the loudness of the corresponding video material as the first loudness of the video segment. Specifically, the terminal device can determine the first loudness of the video segment using the following feasible implementation: It determines whether the cache contains the loudness of the corresponding video material; if the cache contains the loudness of the corresponding video material, then the loudness of the corresponding video material is determined as the first loudness of the video segment; if the cache does not contain the loudness of the corresponding video material, then the loudness of the corresponding video material is identified to obtain the first loudness of the video segment.
[0098] Optionally, if the cache contains an identifier for the video clip corresponding to the video material, the terminal device can determine that the cache includes the loudness of the video material corresponding to the video clip; if the cache does not contain an identifier for the video material corresponding to the video clip, the terminal device can determine that the cache does not include the loudness of the video material corresponding to the video clip. For example, since the cache can include both the loudness and the identifier of the video material, the terminal device can determine whether the cache includes the loudness of the video material based on the identifier of the video material.
[0099] For example, for any given video clip, the cache may or may not store the loudness of that video clip (e.g., the loudness of the video clip corresponding to the first video segment may not be stored in the cache). Therefore, the terminal device can also determine whether the cache includes the loudness of the video clip corresponding to the video segment based on the identifier of the video clip in the cache.
[0100] If the cache includes the loudness of the video material corresponding to the video segment, it means that the terminal device can quickly obtain the loudness of the video material corresponding to the video segment from the cache, and thus obtain the first loudness of the video segment. For example, if the path of video material a corresponding to video segment 1 is path A, and the path of video material b corresponding to video segment 2 is path B, if the cache includes both path A and path B, it means that the cache includes the loudness of video material a and the loudness of video material b. Therefore, when the terminal device determines the first loudness of video segment 1, it can obtain the loudness of video material a from the cache to obtain the first loudness of video segment 1. When the terminal device determines the first loudness of video segment 2, it can obtain the loudness of video material b from the cache to obtain the first loudness of video segment 2.
[0101] If the cache does not contain the loudness of the video material corresponding to the video segment, it means that the terminal device cannot obtain the loudness of the video material corresponding to that video segment from the cache. Therefore, the terminal device can perform loudness recognition on the video material corresponding to the video segment to obtain the first loudness of the video segment. For example, if the path of video material a corresponding to video segment 1 is path A, and the path of video material b corresponding to video segment 2 is path B, if the cache includes path A but not path B, it means that the cache includes the loudness of video material a but not the loudness of video material b. Therefore, when the terminal device determines the first loudness of video segment 1, it can obtain the loudness of video material a from the cache to obtain the first loudness of video segment 1. When the terminal device determines the first loudness of video segment 2, it can perform loudness recognition on video material b to obtain the first loudness of video segment 2.
[0102] If the cache does not include the loudness of the video material corresponding to the video segment, the terminal device can store the loudness of the video material corresponding to the video segment and the identifier of the video material corresponding to the video segment in the cache after identifying the loudness of the video material corresponding to the video segment. In this way, the terminal device can reuse the loudness of the video material in the cache, thereby improving the accuracy of determining the first loudness of the video segment.
[0103] For example, a video editing draft includes video clip 1, video clip 2, video clip 3, and video clip 4. Video clips 1, 2, and 3 belong to video material A, and video clip 4 belongs to video material B. The cache is empty. When the terminal device determines the first loudness of video clip 1, because the cache is empty, the terminal device can determine the loudness of video material A corresponding to video clip 1, thus obtaining the first loudness of video clip 1. Furthermore, the terminal device can store the loudness and identifier of video material A in the cache. When the terminal device determines the first loudness of video clips 2 and 3, because the cache includes the loudness of video material A... Therefore, the terminal device can obtain the loudness of video material A from the cache to obtain the first loudness of video clips 2 and 3. When the terminal device determines the first loudness of video clip 4, since the cache does not include the loudness of video material B, the terminal device can determine the loudness of video material B to obtain the first loudness of video clip 4, and store the loudness and identifier of video material B in the cache. In this way, since the terminal device can reuse the loudness of video materials based on the cache, when there are many video clips with the same video material, the terminal device can quickly determine the first loudness of the video clips, improving the accuracy of determining the first loudness.
[0104] The experiment involved 5 video clips, each from a single video source. If each video clip's corresponding video source was fully identified, it would require identifying the video source 5 times, each time taking approximately 26 seconds, for a total time of 130 seconds. If only the loudness of the first video clip was identified, it would require identifying the video source once, taking 26 seconds. The loudness of the other video clips could be reused, meaning the total time was still 26 seconds. This improved the efficiency of determining the first loudness and the efficiency of unified loudness processing.
[0105] This disclosure provides a method for unified loudness processing. A terminal device can determine the video material corresponding to a video segment, and whether the cache contains the loudness of the corresponding video material. If the cache contains the loudness of the corresponding video material, then the loudness of the corresponding video material is determined as the first loudness of the video segment. If the cache does not contain the loudness of the corresponding video material, then the loudness of the corresponding video material is identified to obtain the first loudness of the video segment. In this way, since the terminal device can reuse the loudness of the video material based on the cache, it can quickly determine the first loudness of the video segment, improving the accuracy of determining the first loudness.
[0106] Figure 8 This is a schematic diagram of a loudness processing device provided in an embodiment of this disclosure. Please refer to [link / reference]. Figure 8The loudness processing device 800 includes an acquisition module 801, a determination module 802, and a processing module 803, wherein:
[0107] The acquisition module 801 is used to acquire a video editing draft, which includes multiple video clips formed based on video materials.
[0108] The determining module 802 is used to determine the first loudness of the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache;
[0109] The processing module 803 is used to process the loudness of the video editing draft according to the first loudness of the plurality of video segments, so that the loudness of the plurality of video segments is the same.
[0110] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0111] Determine the duration percentage of the multiple video segments in the video material, and the number of video segments that belong to the same video material;
[0112] Multiple first loudnesses are determined based on the duration percentage and the number of video segments with the same video material.
[0113] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0114] If the duration percentage is less than or equal to the first threshold, or if the number of video segments with the same video material is less than or equal to the second threshold, then each video segment in the plurality of video segments is identified to obtain a plurality of first loudnesses;
[0115] If the duration percentage is greater than the first threshold, and the number of video segments with the same video material is greater than the second threshold, then the first loudness of the multiple video segments is determined based on the cache.
[0116] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0117] Determine the video material corresponding to the video segment;
[0118] The first loudness of the video segment is determined based on the video material corresponding to the video segment and the cache.
[0119] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0120] Determine whether the cache includes the loudness of the video material corresponding to the video segment;
[0121] If the cache includes the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is determined as the first loudness of the video segment;
[0122] If the cache does not include the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is identified to obtain the first loudness of the video segment.
[0123] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0124] If the cache contains an identifier for the video material corresponding to the video segment, then the cache is determined to include the loudness of the video material corresponding to the video segment.
[0125] If the identifier of the video material corresponding to the video segment does not exist in the cache, it is determined that the loudness of the video material corresponding to the video segment is not included in the cache.
[0126] According to one or more embodiments of this disclosure, the determining module 802 is further configured to:
[0127] The cache stores the loudness of the video material corresponding to the video segment and the identifier of the video material corresponding to the video segment.
[0128] According to one or more embodiments of this disclosure, the processing module 803 is specifically used for:
[0129] The average of the first loudness of the plurality of video segments is determined to obtain the second loudness;
[0130] The loudness of multiple video clips in the video editing draft is adjusted to the second loudness.
[0131] The loudness processing device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0132] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this disclosure. Please refer to [link / reference]. Figure 9The diagram illustrates a structural schematic suitable for implementing the terminal device 900 of the embodiments of the present disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The terminal device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0133] like Figure 9 As shown, the terminal device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the terminal device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0134] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows terminal device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 A terminal device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0135] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0136] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0137] The aforementioned computer-readable medium may be included in the aforementioned terminal device; or it may exist independently and not assembled into the terminal device.
[0138] The aforementioned computer-readable medium carries one or more programs, which, when executed by the terminal device, cause the terminal device to perform the method shown in the above embodiments.
[0139] This disclosure provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements various methods that may be involved in the above embodiments.
[0140] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements various methods that may be involved in the above embodiments.
[0141] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0143] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0144] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0146] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0147] In a first aspect, embodiments of this disclosure provide a loudness processing method, the loudness processing method comprising:
[0148] Obtain a video editing draft, which includes multiple video clips formed based on video footage;
[0149] A first loudness is determined for the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache;
[0150] The loudness of the video editing draft is processed based on the first loudness of the plurality of video clips so that the loudness of the plurality of video clips is the same.
[0151] According to one or more embodiments of this disclosure, determining the first loudness of the plurality of video segments includes:
[0152] Determine the duration percentage of the multiple video segments in the video material, and the number of video segments that belong to the same video material;
[0153] Multiple first loudnesses are determined based on the duration percentage and the number of video segments with the same video material.
[0154] According to one or more embodiments of this disclosure, determining multiple first loudnesses based on the duration percentage and the number of video segments with the same video material includes:
[0155] If the duration percentage is less than or equal to the first threshold, or if the number of video segments with the same video material is less than or equal to the second threshold, then each video segment in the plurality of video segments is identified to obtain a plurality of first loudnesses;
[0156] If the duration percentage is greater than the first threshold, and the number of video segments with the same video material is greater than the second threshold, then the first loudness of the multiple video segments is determined based on the cache.
[0157] According to one or more embodiments of this disclosure, for any one video segment; determining the first loudness of the plurality of video segments based on the cache includes:
[0158] Determine the video material corresponding to the video segment;
[0159] The first loudness of the video segment is determined based on the video material corresponding to the video segment and the cache.
[0160] According to one or more embodiments of this disclosure, determining the first loudness of the video segment based on the video material corresponding to the video segment and the cache includes:
[0161] Determine whether the cache includes the loudness of the video material corresponding to the video segment;
[0162] If the cache includes the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is determined as the first loudness of the video segment;
[0163] If the cache does not include the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is identified to obtain the first loudness of the video segment.
[0164] According to one or more embodiments of this disclosure, determining whether the cache includes the loudness of the video material corresponding to the video segment includes:
[0165] If the cache contains an identifier for the video material corresponding to the video segment, then the cache is determined to include the loudness of the video material corresponding to the video segment.
[0166] If the identifier of the video material corresponding to the video segment does not exist in the cache, it is determined that the loudness of the video material corresponding to the video segment is not included in the cache.
[0167] According to one or more embodiments of this disclosure, after identifying the loudness of the video material corresponding to the video segment, the method further includes:
[0168] The cache stores the loudness of the video material corresponding to the video segment and the identifier of the video material corresponding to the video segment.
[0169] According to one or more embodiments of this disclosure, the loudness of the video editing draft is processed based on a first loudness of the plurality of video clips, including:
[0170] The average of the first loudness of the plurality of video segments is determined to obtain the second loudness;
[0171] The loudness of multiple video clips in the video editing draft is adjusted to the second loudness.
[0172] Secondly, embodiments of this disclosure provide a loudness processing apparatus, which includes an acquisition module, a determination module, and a processing module, wherein:
[0173] The acquisition module 801 is used to acquire a video editing draft, which includes multiple video clips formed based on video materials.
[0174] The determining module 802 is used to determine the first loudness of the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache;
[0175] The processing module 803 is used to process the loudness of the video editing draft according to the first loudness of the plurality of video segments, so that the loudness of the plurality of video segments is the same.
[0176] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0177] Determine the duration percentage of the multiple video segments in the video material, and the number of video segments that belong to the same video material;
[0178] Multiple first loudnesses are determined based on the duration percentage and the number of video segments with the same video material.
[0179] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0180] If the duration percentage is less than or equal to the first threshold, or if the number of video segments with the same video material is less than or equal to the second threshold, then each video segment in the plurality of video segments is identified to obtain a plurality of first loudnesses;
[0181] If the duration percentage is greater than the first threshold, and the number of video segments with the same video material is greater than the second threshold, then the first loudness of the multiple video segments is determined based on the cache.
[0182] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0183] Determine the video material corresponding to the video segment;
[0184] The first loudness of the video segment is determined based on the video material corresponding to the video segment and the cache.
[0185] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0186] Determine whether the cache includes the loudness of the video material corresponding to the video segment;
[0187] If the cache includes the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is determined as the first loudness of the video segment;
[0188] If the cache does not include the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is identified to obtain the first loudness of the video segment.
[0189] According to one or more embodiments of this disclosure, the determining module 802 is specifically used for:
[0190] If the cache contains an identifier for the video material corresponding to the video segment, then the cache is determined to include the loudness of the video material corresponding to the video segment.
[0191] If the identifier of the video material corresponding to the video segment does not exist in the cache, it is determined that the loudness of the video material corresponding to the video segment is not included in the cache.
[0192] According to one or more embodiments of this disclosure, the determining module 802 is further configured to:
[0193] The cache stores the loudness of the video material corresponding to the video segment and the identifier of the video material corresponding to the video segment.
[0194] According to one or more embodiments of this disclosure, the processing module 803 is specifically used for:
[0195] The average of the first loudness of the plurality of video segments is determined to obtain the second loudness;
[0196] The loudness of multiple video clips in the video editing draft is adjusted to the second loudness.
[0197] Thirdly, this disclosure provides a terminal device including: a processor and a memory;
[0198] The memory stores computer-executed instructions;
[0199] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the loudness processing methods described in the first aspect above and various possible aspects of the first aspect.
[0200] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the loudness processing methods described in the first aspect and various possible aspects of the first aspect.
[0201] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the loudness processing methods described in the first aspect above and various possible aspects of the first aspect.
[0202] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0203] It is understood that the data involved in this technical solution (including but not limited to the data itself, its acquisition, or its use) shall comply with the requirements of relevant laws, regulations, and provisions. Data may include information, parameters, and messages, such as flow control instructions.
[0204] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0205] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms of implementing the claims.
Claims
1. A loudness processing method, characterized in that, include: Obtain a video editing draft, which includes multiple video clips formed based on video footage; A first loudness is determined for the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache; The loudness of the video editing draft is processed based on the first loudness of the plurality of video clips so that the loudness of the plurality of video clips is the same.
2. The method according to claim 1, characterized in that, Determining the first loudness of the plurality of video segments includes: Determine the duration percentage of the multiple video segments in the video material, and the number of video segments that belong to the same video material; Multiple first loudnesses are determined based on the duration percentage and the number of video segments with the same video material.
3. The method according to claim 2, characterized in that, The determination of multiple first loudnesses based on the duration ratio and the number of video segments with the same video material includes: If the duration percentage is less than or equal to the first threshold, or if the number of video segments with the same video material is less than or equal to the second threshold, then each video segment in the plurality of video segments is identified to obtain a plurality of first loudnesses; If the duration percentage is greater than the first threshold, and the number of video segments with the same video material is greater than the second threshold, then the first loudness of the multiple video segments is determined based on the cache.
4. The method according to claim 3, characterized in that, For any given video segment; determining the first loudness of the plurality of video segments based on the cache includes: Determine the video material corresponding to the video segment; The first loudness of the video segment is determined based on the video material corresponding to the video segment and the cache.
5. The method according to claim 4, characterized in that, The step of determining the first loudness of the video segment based on the video material corresponding to the video segment and the cache includes: Determine whether the cache includes the loudness of the video material corresponding to the video segment; If the cache includes the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is determined as the first loudness of the video segment; If the cache does not include the loudness of the video material corresponding to the video segment, then the loudness of the video material corresponding to the video segment is identified to obtain the first loudness of the video segment.
6. The method according to claim 5, characterized in that, Determining whether the cache includes the loudness of the video material corresponding to the video segment includes: If the cache contains an identifier for the video material corresponding to the video segment, then the cache is determined to include the loudness of the video material corresponding to the video segment. If the identifier of the video material corresponding to the video segment does not exist in the cache, it is determined that the loudness of the video material corresponding to the video segment is not included in the cache.
7. The method according to claim 5 or 6, characterized in that, After identifying the loudness of the video material corresponding to the video segment, the method further includes: The cache stores the loudness of the video material corresponding to the video segment and the identifier of the video material corresponding to the video segment.
8. The method according to any one of claims 1-6, characterized in that, Based on the first loudness of the plurality of video clips, the loudness of the video editing draft is processed, including: The average of the first loudness of the plurality of video segments is determined to obtain the second loudness; The loudness of multiple video clips in the video editing draft is adjusted to the second loudness.
9. A loudness processing device, characterized in that, It includes an acquisition module, a determination module, and a processing module, wherein: The acquisition module is used to acquire a video editing draft, which includes multiple video clips formed based on video materials. The determining module is used to determine the first loudness of the plurality of video segments; wherein the first loudness of the plurality of video segments is the loudness obtained by recognizing each video segment in the plurality of video segments respectively, or the loudness of the video material corresponding to the video segment determined based on the cache; The processing module is used to process the loudness of the video editing draft according to the first loudness of the plurality of video segments, so that the loudness of the plurality of video segments is the same.
10. A terminal device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the loudness processing method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by the processor, implement the loudness processing method as described in any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the loudness processing method as described in any one of claims 1-8.