Audio / video processing method, device, equipment, and storage medium

The audio-video processing method addresses inefficiencies in clipping by using text-timestamp mapping for precise editing, offering one-click solutions and intelligent enhancements, thereby improving accuracy and user experience.

JP7764507B2Active Publication Date: 2025-11-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023578889
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-22
Filing Date
2022-09-02
Publication Date
2025-11-05
Estimated Expiration
2042-09-02

AI Technical Summary

Technical Problem

Current audio-video clipping methods require repetitive and cumbersome adjustments to achieve accurate start and end points, leading to inefficiencies in user operations.

Method used

An audio-video processing method that presents text data mapped to timestamps, allowing for precise audio-video clipping through operations like selection, deletion, and modification, with features like one-click editing for keywords and silent clips, audio reinforcement, background music addition, loudness balancing, and intelligent teaser controls.

Benefits of technology

Improves the accuracy of audio-video clipping and simplifies user operations by enabling intuitive and efficient editing, reducing the user's operation threshold and enhancing the overall listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764507000001
    Figure 0007764507000001
  • Figure 0007764507000002
    Figure 0007764507000002
  • Figure 0007764507000003
    Figure 0007764507000003
Patent Text Reader

Abstract

An audio-video processing method, apparatus, device and storage medium are provided, wherein the method includes presenting text data corresponding to an audio-video to be edited, the text data having a mapping relationship with audio-video timestamps of the audio-video to be edited, presenting the audio-video to be edited according to a time axis trajectory, determining an audio-video timestamp corresponding to the target text data as a target audio-video timestamp in response to a preset operation triggered on target text data in the text data, and processing an audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited based on the preset operation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This disclosure claims priority from Chinese patent application number "202111109213.4", filed on September 22, 2021, entitled "Audio-Video Processing Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to the field of data processing, and in particular to audio / video processing methods, devices, equipment and storage media. [Background technology]

[0003] With the increasing abundance of information on the Internet, watching audio and video has become a popular entertainment activity in people's daily lives. In order to improve users' viewing experience, clipping audio and video before uploading various audio and video is an important step.

[0004] Currently, during audio-video clipping, for some minor changes such as invalid term clipping, users generally need to repeatedly listen to the audio-video and fine-tune the time start and end points to clip the audio-video, which is cumbersome and requires improving the accuracy of audio-video clipping. Summary of the Invention

[0005] To solve or at least partially solve the above technical problems, an embodiment of the present disclosure provides an audio-video processing method that can improve the accuracy of audio-video clipping and simplify user operations.

[0006] According to a first aspect, the present disclosure provides an audio-video processing method, said method comprising: Presenting text data corresponding to the audio-video to be edited, the text data having a mapping relationship with an audio-video timestamp of the audio-video to be edited; presenting the ready-to-edit audio-video according to a time axis trajectory; determining an audio / video timestamp corresponding to the target text data as a target audio / video timestamp in response to a predetermined operation triggered on the target text data in the text data; processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation; Includes:

[0007] In one alternative embodiment, the method comprises: presenting a first editing entry for a preset keyword or a preset silent clip; displaying the preset keyword or the preset silent clip in the text data in accordance with a preset second display manner in response to a trigger operation on the first editing entrance; Further includes:

[0008] In one alternative embodiment, the first editing entry corresponds to a first editing card on which a one-click deletion control is set, and in response to the trigger operation on the first editing entry, the preset keyword or the preset silent clip in the text data is displayed in accordance with a preset second display manner, and then: The method further includes deleting the preset keyword or the preset silent clip from the text data in response to a trigger operation on the one-click deletion control.

[0009] In one alternative embodiment, the method comprises: presenting audio reinforcement controls on a second edit card; and The method further includes performing reinforcement processing on human voices in the audio video to be edited in response to a trigger operation on the audio reinforcement control.

[0010] In one alternative embodiment, the method comprises: determining background music corresponding to the audio video to be edited based on a music genre of the audio video to be edited and / or content in text data corresponding to the audio video to be edited; adding said background music to said ready-to-edit audio / video clip; Further includes:

[0011] In one alternative embodiment, the method comprises: presenting a loudness balance control on a third edit card; and In response to a trigger operation on the loudness balance control, normalizing the loudness of the volume of the audio video to be edited. Further includes:

[0012] In one alternative embodiment, the method comprises: presenting an intelligent teaser control in a fourth editing card; In response to a trigger operation on the intelligent teaser control, a volume of music and a volume of human voice in an audio-video clip within a previous preset time period in the audio-video to-be-edited are adjusted to obtain a volume-adjusted audio-video clip, wherein the volume of music in the volume-adjusted audio-video clip is inversely proportional to the volume of human voice; Further includes:

[0013] In one alternative embodiment, the preset operation includes a selection operation, and processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited based on the preset operation as described above includes: Displaying an audio / video clip corresponding to the target audio / video timestamp in the audio / video to-be-edited according to a first preset display manner.

[0014] In one alternative embodiment, the preset operation includes a delete operation, and processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation as described above includes: Deleting the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited based on the deletion operation.

[0015] In one alternative embodiment, the preset operation includes a modification operation, and processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation as described above includes: acquiring corrected text data corresponding to the correction operation; generating an audio / video clip based on the corrected text data and tone information in the audio / video to be edited, to create an audio / video clip to be edited; performing a replacement process on an audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited using the audio-video clip to be modified; Includes:

[0016] In one alternative embodiment, the method comprises: When an additional operation for first text data in the text data is received, generating a first audio / video clip based on the first text data and tone information in the audio / video to be edited; determining a first audio / video timestamp corresponding to the first text data based on location information of the first text data within the text data; adding the first audio-video clip to the audio-video to-be-edited based on the first audio-video timestamp; Further includes:

[0017] According to a second aspect, the present disclosure further provides an audio / video processing apparatus, the apparatus comprising: a first presentation module for presenting text data corresponding to the audio-video to be edited, the text data having a mapping relationship with an audio-video timestamp of the audio-video to be edited; a second presentation module for presenting the ready-to-edit audio-video according to a time axis trajectory; a determination module for determining an audio / video timestamp corresponding to the target text data as a target audio / video timestamp in response to a predetermined operation being triggered on target text data in the text data; an editing module for processing an audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation; Includes:

[0018] According to a third aspect, the present disclosure provides a computer-readable storage medium having instructions stored therein, the instructions causing the terminal device to implement the above-described method when executed on the terminal device.

[0019] According to a fourth aspect, the present disclosure provides an apparatus, the apparatus including a memory, a processor, and a computer program stored in the memory and runnable on the processor, the apparatus realizing the above-described method when the processor executes the computer program.

[0020] According to a fifth aspect, the present disclosure provides a computer program product, the computer program product including computer programs / instructions which, when executed by a processor, implement the above method.

[0021] The technical solution according to the embodiments of the present disclosure has the following advantages over the related art:

[0022] An embodiment of the present disclosure provides an audio-video processing method, which presents text data corresponding to an audio-video to be edited, and in response to a preset operation triggered on target text data in the text data, determines an audio-video timestamp corresponding to the target text data as a target audio-video timestamp, and processes an audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited according to the preset operation. As can be seen, the audio-video processing method according to the embodiment of the present disclosure can improve the accuracy of audio-video clipping, simplify user operation, and reduce the user operation threshold. [Brief explanation of the drawings]

[0023] The drawings herein are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure, and together with the specification serve to explain the principles of the present disclosure.

[0024] In order to more clearly explain the technical solutions in the embodiments of the present disclosure or related art, the following briefly introduces drawings that need to be used in the description of the embodiments or prior art, and it is obvious that those skilled in the art can also derive other drawings based on these drawings without exerting any creative effort. [Figure 1] 1 is a flowchart of an audio-video processing method according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram of an audio-video processing interface according to an embodiment of the present disclosure. [Figure 3]FIG. 2 is a schematic diagram of another audio-video processing interface according to an embodiment of the present disclosure. [Figure 4] 10 is a flowchart of another audio-video processing method according to an embodiment of the present disclosure. [Figure 5] FIG. 2 is a schematic diagram of another audio-video processing interface according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a schematic diagram of another audio-video processing interface according to an embodiment of the present disclosure. [Figure 7] 1 is a structural schematic diagram of an audio-video processing device according to an embodiment of the present disclosure; [Figure 8] 1 is a structural schematic diagram of an audio-video processing device according to an embodiment of the present disclosure; DETAILED DESCRIPTION OF THE INVENTION

[0025] In order to make the above-mentioned objects, features and advantages of the present disclosure more clearly understood, the present disclosure will be further described below. What is described below is that unless there is a conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0026] In the following description, numerous specific details are set forth to facilitate a thorough understanding of the present disclosure; however, the present disclosure may be implemented in other ways different from those described herein, and it is apparent that the embodiments in the specification are merely some embodiments of the present disclosure, but not all embodiments.

[0027] An embodiment of the present disclosure provides an audio-video processing method, and referring to FIG. 1, there is shown a flowchart of the audio-video processing method according to an embodiment of the present disclosure, which includes the following steps:

[0028] S101: Present text data corresponding to audio / video to be edited.

[0029] Here, the text data has a mapping relationship with the audio / video timestamps of the audio / video waiting to be edited, and the audio / video timestamps are used to indicate the playback time of the audio / video for each frame.

[0030] In the embodiments of the present disclosure, the ready-to-edit audio-video includes, but is not limited to, recorded audio-video, audio-video obtained based on a script, etc. The text data may be obtained by performing speech recognition on the ready-to-edit audio-video, or may be a script, where if the text data is a script, the text data may be matched with the ready-to-edit audio-video to obtain a mapping relationship between the text data and the audio-video timestamps of the ready-to-edit audio-video, and the speech recognition method includes, but is not limited to, ASR (Automatic Speech Recognition) technology.

[0031] In this embodiment, text data may be presented in the interface. For example, an example interface is shown in FIG. 2, where area P indicates the presented text data. If the audio video to be edited contains the voices of different users, the text data of the different users can be determined, for example, the text data of user a and user b presented in FIG. 2.

[0032] S102: Present the audio / video to be edited according to the time axis trajectory.

[0033] In this embodiment, the audio-video to be edited can be presented on the interface according to the time axis trajectory. As an example, the area Q in FIG. 2 shows the audio-video to be edited that is presented.

[0034] The description does not specifically limit the order in which steps 102 are performed.

[0035] S103: In response to a preset operation triggered on target text data in the text data, an audio / video timestamp corresponding to the target text data is determined as a target audio / video timestamp.

[0036] In this embodiment, the preset operations include, but are not limited to, a selection operation, a deletion operation, and a modification operation. Because there is a mapping relationship between the text data and the audio-video timestamps of the audio-video to be edited, for target text data in the text data, a target audio-video timestamp corresponding to the target text data can be determined based on the mapping relationship.

[0037] S104: Process the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to a preset operation.

[0038] In an embodiment of the present disclosure, the corresponding audio-video clip in the audio-video to be edited can be determined based on the audio-video timestamp, and audio-video clips corresponding to the target audio-video timestamp in the audio-video to be edited can be processed to realize audio-video clipping based on text, and the corresponding audio-video clips can be clipped in conjunction with the clipping of the text to realize relatively accurate clipping of the audio-video.

[0039] In one optional embodiment, when the preset operation includes a selection operation, processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited based on the preset operation includes displaying the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited according to a preset first display manner.

[0040] As an example, the first display manner is, for example, a highlight display, and Figure 3 shows a schematic diagram of another interface. Referring to Figure 3, based on the selection operation, the target text data can be highlighted, and based on the time axis trajectory, the audio / video clip corresponding to the target audio / video timestamp can be highlighted, and the highlighted part is the dotted part as shown in Figure 3.

[0041] In one alternative embodiment, when the preset operation includes a delete operation, processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited based on the preset operation includes deleting the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited based on the delete operation.

[0042] Here, based on the delete operation, the target text data can be deleted and the audio-video clip corresponding to the target audio-video timestamp can be deleted. For example, as shown in Figure 3, after the target text data is selected, a delete control can be presented, and in response to a trigger operation on the delete control, the target text data and the audio-video clip corresponding to the target audio-video timestamp are deleted.

[0043] In one alternative embodiment, when the preset operation includes a modification operation, processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited based on the preset operation includes obtaining modified text data corresponding to the modification operation, generating an audio-video clip based on the modified text data and tone information in the audio-video to be edited to create the audio-video clip to be modified, and using the audio-video clip to be modified to perform a replacement process on the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited.

[0044] Here, the target text data can be modified based on the modification operation. For example, as shown in Figure 3, after the target text data is selected, a modification control can be presented, and in response to a trigger operation on the modification control, modified text data is generated based on the received modification content. An audio-video clip to be modified is generated based on the modified text data and tone information, and the audio-video clip corresponding to the target audio-video timestamp is replaced based on the audio-video clip to be modified, thereby realizing the modification of the audio-video to be edited.

[0045] In an audio-video processing method according to an embodiment of the present disclosure, by presenting text data corresponding to the audio-video to be edited, in response to a preset operation triggered on target text data in the text data, an audio-video timestamp corresponding to the target text data is determined as a target audio-video timestamp, and an audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited is processed based on the preset operation. As can be seen from this, the audio-video processing method according to an embodiment of the present disclosure can clip the audio-video based on the text, and since there is a mapping relationship between the text and the audio-video timestamp, by clipping the text and clipping the corresponding audio-video clip in conjunction with it, it is possible to achieve relatively high-precision clipping of the audio-video, and by presenting text data having a mapping relationship with the audio-video timestamp, it is possible to intuitively present the audio-video content, which simplifies user operation and reduces the user's operation threshold compared to the solution of users clipping audio-video content in the related art.

[0046] Based on the above embodiment, in the audio / video processing scene, there is a demand for various functions to improve the listening experience, such as clipping invalid intonation words, background music, teaser creation, etc. According to the method of the embodiment of the present disclosure, the above functions can be easily realized and the user's usage threshold can be reduced, as described below.

[0047] In one alternative embodiment, because invalid intonation words such as "um," "yeah," and "uh," as well as silent clips, commonly appear in dialogue, there is a need to edit the audio-video to-be-edited to remove the invalid intonation words and silent clips to ensure consistency of the dialogue.

[0048] Therefore, as shown in FIG. 4, the audio-video processing method of the embodiment of the present disclosure further includes the following steps:

[0049] Step 401: present a first editing entry for a preset keyword or a preset silent clip.

[0050] In this embodiment, the text data corresponding to the audio / video to be presented for editing is detected, and a preset keyword or a preset silent clip in the text data is determined. If the preset keyword or the preset silent clip exists in the text data, a first editing entry point can be presented. For example, the control shown in area A in FIG. 3 is the first editing entry point where the information "Revision Suggestion 01: Remove invalid intonation words" is presented.

[0051] Here, the preset keywords may include terms such as invalid inflection words, and there are various ways to determine the preset keywords in the text data. For example, the preset keywords in the text data can be determined by a matching method, or the preset keywords in the text data can be determined based on natural language processing technology.

[0052] Here, the predetermined silent clip is determined based on the interval between the audio-video timestamps corresponding to two adjacent characters, for example, if the interval is greater than a predetermined threshold, it is determined that there is a predetermined silent clip between the two adjacent characters. The silent clip may be presented in the interface in the form of a space, and the presentation length of the silent clip may optionally be determined based on the value of the interval.

[0053] Step 402: In response to a trigger operation on the first editing entry, a preset keyword in the text data or a preset silent clip is displayed according to a preset second display manner.

[0054] Here, the trigger operation for the first editing entry includes, but is not limited to, a click operation, a voice command, and a touch trace, and the second display manner may be a highlight display or other display manner, and is not specifically limited thereto.

[0055] FIG. 5 shows a schematic diagram of the interface, and the pre-set keywords "eh," "um," and "um" in FIG. 5 are highlighted in the interface as shown by the dotted lines.

[0056] Step 403: Delete a preset keyword or a preset silent clip from the text data in response to a trigger operation on the one-click deletion control.

[0057] In an embodiment of the present disclosure, the first editing entry corresponds to a first editing card with a one-click delete control set, and the first editing card is presented in response to a trigger operation on the first editing entry, and the presentation manner of the first editing card includes, but is not limited to, a pull-down menu, a floating window, etc.

[0058] For example, referring to Figure 5, the first editing card can compile statistics on the number of occurrences for each pre-set keyword, as shown in area B in Figure 5, and present the pre-set keywords and the corresponding number of occurrences in the first editing card.

[0059] Optionally, in response to a trigger operation on a target keyword in the preset keywords, the target keyword is removed from the preset keywords and the occurrence count of the preset keywords presented on the first editing card is synchronously modified, so that the user can remove keywords that do not belong to the invalid inflection words by an operation such as clicking, to avoid them being deleted with one click.

[0060] In this embodiment, the deletion operation of a preset keyword or a preset silent clip can be presented in the form of an editing card, providing one-click operation, saving editing time, simplifying user operations, and reducing the user's usage threshold.

[0061] In one alternative embodiment, the audio-video processing method further includes presenting an audio reinforcement control on a second editing card and reinforcing human voices in the audio-video to be edited in response to a trigger operation on the audio reinforcement control.

[0062] In this embodiment, a second editing entrance is presented for the audio / video to be edited, and the second editing entrance corresponds to a second editing card on which an audio reinforcement control is set. For example, noise detection is performed based on the audio / video to be edited, and if noise is detected, the second editing entrance can be presented. As an example, the control shown in area C in Figure 2 is the second editing entrance on which information "Enhancement suggestion: audio reinforcement" is presented. Furthermore, in response to a trigger operation for the second editing entrance, the second editing card is presented.

[0063] Referring to Figure 6, the second editing card, as shown in area D in Figure 6, presents an audio reinforcement control "reinforcement audio", and in response to a trigger operation on this audio reinforcement control, reinforces the human voice in the audio video waiting to be edited, and the trigger operation includes, but is not limited to, a click operation, a voice command, and a touch trajectory.

[0064] In this embodiment, the voice reinforcement operation can be presented in the form of an editing card, providing one-click operation, reinforcing the user's human voice to satisfy the listening experience, simplifying user operation, and reducing the user's usage threshold.

[0065] In one alternative embodiment, the audio-video processing method further includes determining background music corresponding to the audio-video to be edited based on the music genre of the audio-video to be edited and / or content in the text data corresponding to the audio-video to be edited, and adding the background music to the audio-video clip to be edited.

[0066] In this embodiment, multiple tags can be pre-set, and each tag has a mapping relationship with one or more background music pieces. Based on the music genre of the audio video to be edited and / or the content in the text data corresponding to the audio video to be edited, the tag corresponding to the music genre and / or the content in the text data is determined, and based on the mapping relationship between the tag and the background music, the background music corresponding to the audio video to be edited is determined.

[0067] As an example, for the content in the text data corresponding to the audio / video to be edited, the theme of the content is determined to be "exercise" based on natural language processing technology, and further background music corresponding to the "exercise" tag is determined as the background music corresponding to the audio / video to be edited, and this background music is added to the audio / video clip to be edited.

[0068] As another example, a corresponding tag is determined based on the music genre of the audio / video to be edited, and background music corresponding to this tag is set as background music corresponding to the audio / video to be edited, and this background music is added to the audio / video clip to be edited.

[0069] In this embodiment, background music is intelligently recommended based on the content and genre of text data, so as to meet the scene needs for adding background music, enrich the variety of listening sensations, improve the listening experience, simplify user operations, and reduce the user's usage threshold.

[0070] In one alternative embodiment, the audio-video processing method further includes presenting a loudness balance control on a third edit card, and performing a normalization process on the loudness of the volume in the audio-video to be edited in response to a trigger operation on the loudness balance control.

[0071] In this embodiment, a third editing entry point for the audio / video to be edited is presented, and the third editing entry point corresponds to a third edit card on which loudness balance control is set. For example, if loudness detection is performed on the audio / video to be edited and it is detected that the audio / video to be edited does not satisfy a preset loudness balance condition, the third editing entry point can be presented. Furthermore, in response to a trigger operation for the third editing entry point, the third edit card is presented, and in response to a trigger operation for the loudness balance control, a normalization process is performed on the loudness of the audio / video to be edited, for example, so that the loudness of the audio / video to be edited falls within a preset range.

[0072] In this embodiment, the loudness balance operation can be presented in the form of an editing card, providing one-click operation, improving the listening experience, simplifying user operation, and reducing the user's usage threshold.

[0073] In one alternative embodiment, the audio-video processing method further includes presenting an intelligent teaser control on a fourth editing card, and adjusting the volume of music and the volume of human voices in an audio-video clip within a previous preset time period in the audio-video to be edited in response to a trigger operation on the intelligent teaser control, to obtain a volume-adjusted audio-video clip.

[0074] In this embodiment, a fourth editing entry point for the audio-video to be edited is presented, and the fourth editing entry point corresponds to a fourth editing card on which the intelligent teaser control is set. In response to a trigger operation on the fourth editing entry point, the fourth editing card is presented, and in response to a trigger operation on the intelligent teaser control, the volume of music and the volume of human voice in the audio-video clip within a previous preset time period in the audio-video to be edited are adjusted, for example, the volume of the human voice is increased by a first volume value and the volume of music is decreased by a second volume value, or in the audio-video clip where human voice is detected, the volume of music is decreased by a third volume value, thereby obtaining a volume-adjusted audio-video clip.

[0075] Here, the volume of the music in the audio video clip after volume adjustment is inversely proportional to the volume of the human voice.

[0076] Optionally, opening generation may be further realized based on the intelligent teaser control presented in the fourth editing card, for example, in response to a trigger operation on the intelligent teaser control, the currently selected second text data and the second audio-video clip corresponding to the second text data are determined, and the second text data and the second audio-video clip are copied and pasted into a preset opening area to realize the teaser effect.

[0077] In this embodiment, the intelligent teaser function is presented in the form of an editing card, providing one-click operation, realizing the effect of teaser, simplifying user operation, and reducing the user's usage threshold.

[0078] In one optional embodiment, the audio-video processing method further includes, when an add operation for first text data is received in the text data, generating a first audio-video clip based on the first text data and tone information in the audio-video to be edited, determining a first audio-video timestamp corresponding to the first text data based on position information of the first text data in the text data, and adding the first audio-video clip to the audio-video to be edited based on the first audio-video timestamp.

[0079] In this embodiment, the first text data may be obtained in response to an input operation or may be obtained by copying existing text data. The timbre information of each user can be obtained based on the audio-video to-be-edited. When adding the first text data, a corresponding first audio-video timestamp is determined based on the position information of the first text data in the text data, and the first audio-video clip is added at the position of the first audio-video timestamp.

[0080] It should be noted that the editing entry may be automatically presented based on the detection result or may be presented in the interface in response to a triggering operation.

[0081] In this embodiment, tone cloning and voice announcement technology is adopted to clone the tone based on the added text, intelligently generate audio-video clips, and add audio-video clips based on the text input, thereby reducing the time cost and editing cost of re-recording and simplifying user operations.

[0082] Based on the above method embodiment, the present disclosure further provides an audio-video processing device. Referring to FIG. 7, there is shown a structural schematic diagram of an audio-video processing device according to an embodiment of the present disclosure, the device comprising: a first presentation module 701 for presenting text data corresponding to the audio-video to be edited, the text data having a mapping relationship with an audio-video timestamp of the audio-video to be edited; a second presentation module 702 for presenting the ready-to-edit audio-video according to a time axis trajectory; a determining module 703 for determining an audio-video timestamp corresponding to the target text data as a target audio-video timestamp in response to a preset operation triggered on the target text data in the text data; an editing module 704 for processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation; Includes:

[0083] In one alternative embodiment, the audio / video processing device comprises: The system further includes a first processing module for presenting a first editing entry for a predetermined keyword or a predetermined silent clip, and displaying the predetermined keyword or the predetermined silent clip in the text data according to a predetermined second display manner in response to a trigger operation for the first editing entry.

[0084] In one alternative embodiment, the first editing entry corresponds to a first editing card on which a one-click delete control is set, and the first editing module is further used to delete the preset keyword or the preset silent clip from the text data in response to a trigger operation on the one-click delete control.

[0085] In one alternative embodiment, the audio / video processing device comprises: The system further includes a second processing module for presenting an audio reinforcement control on a second editing card and performing reinforcement processing on human voices in the audio video to be edited in response to a trigger operation on the audio reinforcement control.

[0086] In one alternative embodiment, the audio / video processing device comprises: The audiovisual device further includes a first adding module for determining background music corresponding to the audiovisual to be edited based on the music genre of the audiovisual to be edited and / or content in text data corresponding to the audiovisual to be edited, and adding the background music to the audiovisual clip to be edited.

[0087] In one alternative embodiment, the audio / video processing device comprises: The third editing card further includes a third processing module for presenting a loudness balance control on the third editing card and performing a normalization process on the loudness of the volume in the audio video to be edited in response to a trigger operation on the loudness balance control.

[0088] In one alternative embodiment, the audio / video processing device comprises: The fourth processing module further includes: a fourth processing module for presenting an intelligent teaser control on a fourth editing card; and adjusting, in response to a trigger operation on the intelligent teaser control, the volume of music and the volume of human voice in an audio-video clip within a previous preset time period in the audio-video to be edited to obtain a volume-adjusted audio-video clip, wherein the volume of music in the volume-adjusted audio-video clip is inversely proportional to the volume of the human voice.

[0089] In one alternative embodiment, the preset operation includes a selection operation, and the editing module 704 is specifically used to display the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited according to the preset first display manner.

[0090] In one alternative embodiment, the pre-set operation includes a delete operation, and the editing module 704 is specifically used to delete the audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited based on the delete operation.

[0091] In one alternative embodiment, the pre-set operation includes a modification operation, and the editing module 704 is specifically used to obtain modified text data corresponding to the modification operation, generate an audio / video clip based on the modified text data and tone information in the audio / video to be edited as the audio / video clip to be modified, and use the audio / video clip to be modified to perform a replacement process on the audio / video clip corresponding to the target audio / video timestamp in the audio / video to be edited.

[0092] In one alternative embodiment, the audio / video processing device comprises: The system further includes a second adding module for, when receiving an add operation for first text data in the text data, generating a first audio-video clip based on the first text data and tone information in the audio-video to be edited, determining a first audio-video timestamp corresponding to the first text data based on position information of the first text data in the text data, and adding the first audio-video clip to the audio-video to be edited based on the first audio-video timestamp.

[0093] The explanations for the audio / video processing method in the previous embodiment are also applicable to the audio / video processing device of this embodiment, and will not be further explained here.

[0094] In an audio-video processing device according to an embodiment of the present disclosure, by presenting text data corresponding to the audio-video to be edited, in response to a preset operation triggered on target text data in the text data, an audio-video timestamp corresponding to the target text data is determined as a target audio-video timestamp, and an audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited is processed based on the preset operation. As can be seen from this, the audio-video processing method according to an embodiment of the present disclosure can clip audio-video based on text, and since there is a mapping relationship between the text and the audio-video timestamp, by clipping the text and clipping the corresponding audio-video clip in conjunction with it, relatively high accuracy clipping of audio-video can be achieved, and by presenting text data having a mapping relationship with the audio-video timestamp, audio-video content can be presented intuitively, which simplifies user operation and reduces the user's operation threshold compared to the solutions of users clipping audio-video content in the related art.

[0095] In addition to the above methods and apparatuses, embodiments of the present disclosure further provide a computer-readable storage medium having instructions stored therein, which, when executed by a terminal device, cause the terminal device to implement the audio-video processing method described in the embodiments of the present disclosure.

[0096] An embodiment of the present disclosure further provides a computer program product, which includes a computer program / instruction, which, when executed by a processor, realizes the audio-video processing method described in the embodiment of the present disclosure.

[0097] In addition, an embodiment of the present disclosure further provides an audio-video processing device, and referring to FIG. The audio / video processing device may include a processor 801, a memory 802, an input device 803, and an output device 804. The number of processors 801 in the audio / video processing device may be one or more, and one processor is shown as an example in Figure 8. In some embodiments of the present disclosure, the processor 801, the memory 802, the input device 803, and the output device 804 may be connected by a bus or other method, and here, Figure 8 shows them connected by a bus as an example.

[0098] The memory 802 may be used to store software programs and modules, and the processor 801 executes various functional applications and data processing of the audio / video processing device by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. Here, the program storage area may store an operating system, applications required for at least one function, etc. The memory 802 may include high-speed random access memory or non-volatile memory, such as at least one magnetic disk memory device, flash memory device, or other volatile solid-state memory device. The input device 803 may be used to receive input numeric or character information and to generate signal inputs related to user configuration and function control of the audio / video processing device.

[0099] Specifically, in this embodiment, the processor 801 loads executable files corresponding to one or more application processes into the memory 802 in accordance with the following instructions, and the processor 801 runs the applications stored in the memory 802, thereby realizing various functions of the above-mentioned audio / video processing device.

[0100] It should be noted that, in this specification, related terms such as "first," "second," etc., are used only to distinguish one entity or operation from another, and do not necessarily require or imply the existence of any actual relationship or ordering between those entities or operations. Furthermore, the terms "comprise," "include," "includes," or any other variant thereof are intended to cover the non-exclusive "comprise," whereby a process, method, article, or device comprising a set of elements not only includes those elements, but also other elements not expressly listed, or further elements inherent in such process, method, article, or device. Absent further limitations, an element qualified by the phrase "comprises one of..." does not exclude the presence of other identical elements in the process, method, article, or device comprising said element.

[0101] The above are merely specific embodiments of the present disclosure to enable those skilled in the art to understand or realize the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art. The general principles defined herein may be embodied in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A computer-implemented audio / video processing method comprising: a processor of the computer, Presenting text data corresponding to the audio-video to be edited, the text data having a mapping relationship with an audio-video timestamp of the audio-video to be edited; presenting the ready-to-edit audio-video according to a time axis trajectory; In response to a predetermined operation being triggered on target text data in the text data, determining an audio / video timestamp corresponding to the target text data as a target audio / video timestamp; processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation; Including, Wherein the processor: Presenting a first editing entry for an invalid intonation word or a preset silent clip, wherein the first editing entry corresponds to a first editing card having a one-click delete control; displaying the invalid intonation word or the preset silent clip in the text data in accordance with a preset second display manner in response to a trigger operation on the first editing entry point; deleting the invalid intonation words or the preset silent clips from the text data in response to a trigger operation on the one-click deletion control; Further implementation of method.

2. The method comprises: presenting audio reinforcement controls on a second edit card; and In response to a trigger operation on the audio reinforcement control, reinforcement processing is performed on human voices in the audio video to be edited. The method of claim 1 further comprising:

3. The method comprises: determining background music corresponding to the audio video to be edited based on a music genre of the audio video to be edited and / or content in text data corresponding to the audio video to be edited; adding said background music to said ready-to-edit audio / video clip; The method of claim 1 further comprising:

4. The method comprises: presenting loudness balance control on a third edit card; and In response to a trigger operation on the loudness balance control, normalizing the loudness of the volume of the audio video to be edited. The method of claim 1 further comprising:

5. The method comprises: presenting an intelligent teaser control in a fourth editing card; In response to a trigger operation on the intelligent teaser control, a volume of music and a volume of human voice in an audio-video clip within a previous preset time period in the audio-video to-be-edited are adjusted to obtain a volume-adjusted audio-video clip, wherein the volume of music in the volume-adjusted audio-video clip is inversely proportional to the volume of human voice; The method of claim 1 further comprising:

6. The preset operation includes a selection operation, and processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited based on the preset operation includes: displaying an audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to a first preset display manner; The method of claim 1.

7. The preset operation includes a deletion operation, and processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation as described above includes: The method of claim 1 , further comprising deleting an audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited based on the deletion operation.

8. The preset operation includes a modification operation, and processing the audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation includes: acquiring corrected text data corresponding to the correction operation; generating an audio / video clip based on the corrected text data and tone information in the audio / video to be edited, to create an audio / video clip to be edited; performing a replacement process on an audio-video clip corresponding to the target audio-video timestamp in the audio-video to be edited using the audio-video clip to be modified; The method of claim 1 , comprising:

9. The method comprises: When an additional operation for first text data in the text data is received, generating a first audio / video clip based on the first text data and tone information in the audio / video to be edited; determining a first audio / video timestamp corresponding to the first text data based on location information of the first text data within the text data; adding the first audio-video clip to the audio-video to-be-edited based on the first audio-video timestamp; The method of claim 1 further comprising:

10. 1. An audio / video processing device, comprising: a first presentation module for presenting text data corresponding to the audio-video to be edited, the text data having a mapping relationship with an audio-video timestamp of the audio-video to be edited; a second presentation module for presenting the ready-to-edit audio-video according to a time axis trajectory; a determination module for determining an audio / video timestamp corresponding to the target text data as a target audio / video timestamp in response to a predetermined operation being triggered on target text data in the text data; an editing module for processing an audio-video clip corresponding to the target audio-video timestamp in the audio-video to-be-edited according to the preset operation; Including, Wherein, the audio and video processing device comprises: a first processing module that presents a first editing entry for an invalid intonation word or a preset silent clip, and displays the invalid intonation word or the preset silent clip in the text data according to a preset second display manner in response to a trigger operation on the first editing entry; Wherein, the first editing entry corresponds to a first editing card on which a one-click deletion control is set, and the editing module is further used for deleting the invalid intonation phrase or the preset silent clip from the text data in response to a trigger operation on the one-click deletion control; Audio-video processing equipment.

11. 10. A computer-readable storage medium having stored thereon instructions that, when executed on a terminal device, cause the terminal device to implement a method according to any one of claims 1 to 9.

12. 10. An apparatus comprising a memory, a processor, and a computer program stored in said memory and operable to run on said processor, said apparatus implementing the method of any one of claims 1 to 9 when said processor executes said computer program.

13. A computer program comprising computer program instructions which, when executed by a processor, implements the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Interactive information processing method and device, equipment and medium

    CN112231498A

  • Video editing method and recording medium recording procedure for the same

    JP1998191248A

  • Device and program for data creation and onboard device

    JP2004287193A

  • Device and program for generating dictionary for broadcast speech

    JP2006227363A

  • Nonlinear editing apparatus, and program therefor

    JP2007295218A