Methods, apparatus, devices, and storage media for audio editing
The audio editing method efficiently highlights and deletes unwanted characters by allowing users to select and confirm deletion, improving editing speed and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2023-05-05
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional audio editing methods require repeated listening to locate and delete unwanted characters or words, leading to inefficiency and potential errors such as omissions and incorrect deletions.
An audio editing method that highlights invalid characters in text corresponding to audio, allowing users to select and confirm deletion with a single click, automatically removing the corresponding audio portion.
Enhances editing efficiency by enabling quick and accurate removal of invalid characters, reducing accidental deletions and saving time during the audio editing process.
Smart Images

Figure 0007852078000001 
Figure 0007852078000002 
Figure 0007852078000003
Abstract
Description
Technical Field
[0001] [Cross - reference to Related Applications] This application claims priority to a Chinese invention patent application filed on May 6, 2022, with the invention title "Method, Apparatus, Device, and Storage Medium for Audio Editing" and application number 202210488246.2.
[0002] [Technical Field] Exemplary embodiments of the present invention generally relate to the field of computers, and in particular, to a method, apparatus, device, and computer - readable storage medium for audio editing.
Background Art
[0003] Audio data is a common information interaction method in all aspects of people's life, work, social communication, etc. Currently, people can produce and obtain audio data more and more conveniently and can also share recorded audio. In order to output high - quality audio, various editing operations are expected to be performed on the audio data, such as adjusting the volume, speed, timbre, etc. In some cases, it is also expected to delete words that are not expected to appear from the audio data.
Summary of the Invention
[0004] According to an exemplary embodiment of the present invention, a solution for audio editing is provided.
[0005] In a first aspect of the present invention, a method for audio editing is provided. The method includes presenting one or more invalid characters included in the text corresponding to the audio in a prominent manner in a predefined mode for the audio. The method further includes detecting a deletion confirmation instruction for at least one target invalid character among the one or more invalid characters, and in response to the detection of the deletion confirmation instruction, deleting at least one audio portion corresponding to at least one target invalid character from the audio.
[0006] A second aspect of the present invention provides an apparatus for audio editing. The apparatus comprises: a highlighting module for highlighting one or more invalid characters contained in text corresponding to the audio in a predefined mode for the audio; an instruction detection module for detecting a deletion confirmation instruction for at least one target invalid character among the one or more invalid characters; and an audio deletion module for deleting at least one audio portion corresponding to at least one target invalid character from the audio in response to the detection of a deletion confirmation instruction.
[0007] A third aspect of the present invention provides an electronic device comprising at least one processing unit and at least one memory coupled to the at least one processing unit for use in storing instructions executed by the at least one processing unit. When an instruction is executed by the at least one processing unit, the device is made to execute the method of the first aspect.
[0008] A fourth embodiment of the present invention provides a computer-readable storage medium. The medium stores a computer program, and the computer program is executed by a processor to implement the method of the first embodiment.
[0009] It should be understood that the contents described in the summary section of the present invention are not intended to limit the main or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will be readily apparent from the following description. [Brief explanation of the drawing]
[0010] The above-mentioned features and other features, advantages, and aspects of each embodiment of the present invention will become clearer when viewed in conjunction with the drawings in the following detailed description. In the drawings, the same or similar symbols indicate the same or similar elements, among which, [Figure 1] This figure shows a schematic diagram of an exemplary environment in which an embodiment of the present invention can be implemented. [Figure 2] This figure shows a flowchart of the process for audio editing according to some embodiments of the present invention. [Figure 3A] This figure shows a schematic diagram illustrating an example of the interaction of editing pages for audio editing according to some embodiments of the present invention. [Figure 3B] This figure shows a schematic diagram illustrating an example of the interaction of editing pages for audio editing according to some embodiments of the present invention. [Figure 3C] This figure shows a schematic diagram illustrating an example of the interaction of editing pages for audio editing according to some embodiments of the present invention. [Figure 3D] This figure shows a schematic diagram illustrating an example of the interaction of editing pages for audio editing according to some embodiments of the present invention. [Figure 3E] This figure shows a schematic diagram illustrating an example of the interaction of editing pages for audio editing according to some embodiments of the present invention. [Figure 3F] This figure shows a schematic diagram illustrating an example of the interaction of editing pages for audio editing according to some embodiments of the present invention. [Figure 4] This figure shows a flowchart of the process for prominently displaying invalid characters according to several embodiments of the present invention. [Figure 5] This figure shows a flowchart of the process for prominently displaying invalid characters according to some other embodiments of the present invention. [Figure 6A] This figure shows a schematic diagram of an example of user selection for invalid characters in an editing page according to some embodiments of the present invention. [Figure 6B] This figure shows a schematic diagram of an example of user selection for invalid characters in an editing page according to some embodiments of the present invention. [Figure 6C] This figure shows a schematic diagram of an example of user selection for invalid characters in an editing page according to some embodiments of the present invention. [Figure 6D]This figure shows a schematic diagram of an example of user selection for invalid characters in an editing page according to some embodiments of the present invention. [Figure 6E] This figure shows a schematic diagram of an example of user selection for invalid characters in an editing page according to some embodiments of the present invention. [Figure 6F] This figure shows a schematic diagram of an example of user selection for invalid characters in an editing page according to some embodiments of the present invention. [Figure 7A] This figure shows a schematic diagram illustrating the presentation of an exemplary page of audio editing according to some embodiments of the present invention. [Figure 7B] This figure shows a schematic diagram illustrating the presentation of an exemplary page of audio editing according to several embodiments of the present invention. [Figure 8] This figure shows a block diagram of an audio editing apparatus according to some embodiments of the present invention. [Figure 9] This figure shows a block diagram of a device that can implement multiple embodiments of the present invention. [Modes for carrying out the invention]
[0011] The embodiments of the present invention will be described in more detail below with reference to the drawings. Although specific embodiments of the present invention are shown in the drawings, the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided for a more thorough and complete understanding of the present invention. The drawings and embodiments of the present invention are for illustrative purposes only and should not be used to limit the scope of protection of the present invention.
[0012] In the description of the embodiments of the present invention, the term "including" and its similar terms are open-ended inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may be other explicit and implicit definitions hereinafter.
[0013] It is understood that the data related to the present technical solution (including but not limited to the data itself, data acquisition, or data use) should comply with the corresponding laws and regulations and related specified requirements.
[0014] Before using the technical solutions disclosed in each embodiment of the present invention, it should be understood that the types, scope of use, usage scenarios, etc. of the personal information related to the present invention should be notified to the user in an appropriate manner in accordance with relevant laws and regulations, and the user's permission should be obtained.
[0015] For example, when responding to receiving an uncommitted request from a user, by sending presentation information to the user, it is explicitly presented to the user that the requested operation requires the acquisition and use of the user's personal information. Thereby, the user can independently select whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that executes the operation of the technical solution of the present invention based on the presentation information.
[0016] As a selective but non-limiting implementation form, the method of sending presentation information to the user in response to receiving an uncommitted request from the user may be, for example, a method using a pop-up window, and the presentation information can be displayed in text form within the pop-up window. In addition, the pop-up window may further include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0017] The notification and user authorization process described above is merely a general overview and does not limit the implementation of the present invention. It is understood that other means that satisfy the relevant laws and regulations can also be applied to the implementation of the present invention.
[0018] Figure 1 shows a schematic diagram of an exemplary environment 100 in which an embodiment of the present invention can be implemented. In this exemplary environment 100, a terminal device 110 may have an audio editing application 112 installed for editing audio 114. For example, the audio editing application 112 can edit audio 114 based on the operation of user 102. In this specification, the audio 114 to be edited may be in any audio format and may have any appropriate audio length. As an example, audio 114 may be a podcast, audio corresponding to a short video, a radio drama, an audiobook, a recording of a meeting or interview, an audiobook course, an audio note, etc.
[0019] In some embodiments, audio 114 may be collected by an audio acquisition device 105 (e.g., a device equipped with a microphone) and provided to an audio editing application 112 for editing. For example, the audio acquisition device 105 may collect audio from at least a user 104. In some embodiments, the audio editing application 112 may provide an audio recording function for recording the audio 114 collected by the audio acquisition device 105. In some embodiments, the audio 114 edited by the audio editing application 112 may be from any other data source, such as audio 114 downloaded or received from another device. Embodiments of the present invention are not limited in this respect.
[0020] The description shows user 102 performing editing operations on audio 114 and user 104 outputting audio 114, but these users may be the same user and are not limited to them in this specification. Furthermore, although shown as a separate device, it will be understood that the audio acquisition device 105 can be integrated with the terminal device 110. In other implementations, the audio acquisition device 105 may be connected to the terminal device 110 by other means to collect and provide audio 114.
[0021] The terminal device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbooks, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio receivers, e-book devices, game devices, or any combination of the foregoing, and includes accessories and peripherals for these devices or any combination thereof. In some embodiments, the terminal device 110 may also support any type of interface to the user (such as a “wearable” circuit).
[0022] In some embodiments, the terminal device 110 can communicate with a remote computing device 122 to implement editing of the audio 114. For example, the computing device 122 may perform storage functions for the audio 114, perform specific analysis tasks, etc., to extend the storage and processing capabilities of the terminal device 110. The computing device 122 may be any type of computing system / server that can provide computing functions, including but not limited to mainframes, edge computing nodes, and computing devices in a cloud environment. In the example shown in Figure 1, the computing device 122 may be located in a cloud environment 120.
[0023] The configuration and functions of environment 100 are described for illustrative purposes only and should be understood as not to imply any limitation on the scope of the invention. For example, terminal device 110 does not have to communicate with remote computing device 122. Also, for example, user 104 and audio acquisition device 105 may be omitted.
[0024] In audio editing scenarios, it may be expected to remove characters or words that are meaningless or useless to the expression in the audio, or characters or words that are not expected to appear in the audio. In this specification, such characters or words may be referred to as “invalid characters,” “invalid words,” “useless words,” “words without actual meaning,” or “useless words,” and “invalid characters” may be a single character, a word, or a group of words, or any text unit of different sizes, which can vary in size across different natural languages. In some embodiments, invalid characters may include modal particles, mantras, etc., that appear in spoken expressions, such as “a,” “ya,” “on,” “e,” “kono,” and “ano,” and these meaningless words are considered invalid expressions. In some embodiments, invalid characters may additionally or alternatively include other characters or words that are not expected to appear in the audio, such as sensitive words. Sensitive words that are not expected to appear may vary across different application scenarios and are identified as needed.
[0025] In conventional solutions, removing characters or words that are not expected to appear in the audio requires the audio editor to repeatedly listen to the audio, find the character or word to be removed, pinpoint its exact location, select the corresponding audio portion, and delete it. Such an editing process is inefficient and prone to many problems, such as omissions and incorrect deletions (e.g., the audio portion to be deleted is too long or too short).
[0026] According to an embodiment of the present invention, an improved audio editing solution is proposed. In this solution, one or more invalid characters present in the audio are identified and highlighted in a text-based format corresponding to the audio, allowing the user to select a specific invalid character or a selection of invalid characters and confirm whether to delete them. After a deletion confirmation instruction for an invalid character is detected, the audio portion corresponding to the invalid character that has been confirmed to be deleted is automatically removed from the audio.
[0027] This solution supports the convenient removal of invalid characters in audio, significantly improving the efficiency of audio editing. From the user's perspective, it enables the identification and removal of invalid characters with a single click, avoiding redundant operations and saving time during audio editing. By providing users with a prominent indication of potentially removable invalid characters, accidental deletions and omissions can be effectively avoided.
[0028] In the following sections, several exemplary embodiments of the present invention will be described with reference to the drawings.
[0029] Figure 2 shows a flowchart of process 200 for audio editing according to several embodiments of the present invention. Process 200 can be implemented in terminal device 110. For ease of explanation, process 200 will be described with reference to environment 100 in Figure 1.
[0030] In block 210, the terminal device 110, in a predefined mode for audio, prominently displays invalid character sets in the text corresponding to audio 114.
[0031] In embodiments of the present invention, corresponding text is identified from audio 114 to assist in editing the audio 114. In some embodiments, automatic speech recognition (ASR) technology can be used to identify the corresponding text from audio 114. Text identification may be performed on terminal device 110. In other embodiments, text identification is performed by a remote computing device, for example, computing device 122 in environment 100. Terminal device 110 can also receive text from computing device 122.
[0032] In embodiments of the present invention, a predefined mode is provided in which a set of invalid characters in text can be positioned and highlighted, and the set of invalid characters includes one or more invalid characters. Hereinafter, for ease of explanation, this predefined mode will be referred to as the "invalid character positioning mode." In some embodiments, the invalid character positioning mode can be entered in response to user selection.
[0033] In embodiments of the present invention, the invalid characters to be prominently displayed are determined on a text basis. In some embodiments, the prominently displayed invalid characters may include one or more invalid characters automatically identified from the text. Automatically identifying invalid characters saves the user time in identifying them. In particular, compared to methods of positioning invalid characters by listening to audio, automatic identification can present the presence of invalid characters to the user more quickly. In this way, after being triggered and entering invalid character positioning mode, invalid characters identified from the text can be automatically and quickly presented prominently.
[0034] In some of the other embodiments described below, the prominently displayed invalid characters may include, additionally or alternatively, one or more invalid characters selected and determined by the user. For example, the user may be able to select one or more characters as invalid characters from the presented text. Compared to the method of locating invalid characters by listening to audio, the user can identify invalid characters more easily and accurately within the text.
[0035] In some embodiments, automatic identification of invalid characters may be performed on the terminal device 110. In other embodiments, automatic identification of invalid characters may be performed by a remote computing device, for example, a computing device 122 in environment 100. The terminal device 110 can also obtain a set of invalid characters automatically identified from the computing device 122.
[0036] Invalid characters in text can be automatically identified using various methods. In some embodiments, an invalid character list can be pre-selected, created, and maintained, recording common invalid characters such as "あ", "や", "おん", "え", "この", "あの", and / or other characters or words that are not expected to appear in the audio, such as sensitive words. By matching each character in the text corresponding to audio 114 against the invalid character list, the invalid characters contained in the text can be determined. Only non-restrictive examples of invalid characters are listed here, and it should be understood that more, fewer, or other invalid characters may be recorded in the invalid character list in various languages and application scenarios.
[0037] Alternatively or additionally, in some embodiments, invalid character recognition models can be built and trained to identify invalid characters from input text. Such invalid character recognition models can be built and trained based on various machine learning or deep learning algorithms. The input to an invalid character recognition model may include text, and the output may include identification results. The identification results indicate whether or not invalid characters are present in the text, and if so, further include instructions for the identified invalid characters.
[0038] The training data for training such an invalid character recognition model may include sample text, and may further include annotation information for invalid characters within the sample text. The invalid character recognition model can be constructed using a machine learning or deep learning model suitable for text processing, and the model can be trained using an appropriate machine learning or deep learning training algorithm. The embodiments of the present invention do not specifically limit the configuration and training process of the invalid character recognition model.
[0039] It will be understood that invalid character identification can be performed on the terminal device 110's local or remote computing device 122, whether based on an invalid character list or an invalid character identification model. In some embodiments, invalid character identification may be initiated after receiving a trigger to enter invalid character positioning mode. In some embodiments, invalid character identification can be performed asynchronously, for example, the terminal device 110 or computing device 112 can obtain audio 114, identify a set of invalid characters from the text corresponding to audio 114, and record these identified invalid characters. Then, after entering invalid character positioning mode, the previously identified invalid characters can be quickly and prominently displayed.
[0040] In some embodiments, the audio editing application 112 can perform editing on the audio 114, such as deleting audio portions corresponding to invalid characters. For example, the audio editing application 112 can provide an editing page for the audio 114. The audio editing application 112 can also provide an invalid character positioning mode. In the invalid character positioning mode, sets of invalid characters in the text are highlighted on the editing page. In some embodiments, the text can be displayed on the editing page, and sets of invalid characters can be highlighted while the text is being displayed.
[0041] The prominent presentation of invalid characters means that invalid characters are displayed differently from other characters in the text. One or more prominent presentation methods can be used to implement prominent presentation of invalid characters. Examples of prominent presentation methods may include adding a strikethrough line (i.e., a line through the middle of the character) or underline to invalid characters, changing the formatting of invalid characters (such as color, font size, typeface, and / or weight) to distinguish them from other characters, overlaying a background pattern of a specific color or shape on invalid characters, adding special shapes or annotations on invalid characters, and methods that allow any other invalid characters to be prominently presented.
[0042] In some embodiments, when other characters in the text are displayed simultaneously, invalid characters can be made more prominent by changing the display manner of the other characters. For example, the formatting of the other characters (e.g., color, font size, typeface and / or boldness) may be changed, or the other characters may be hidden or at least partially hidden.
[0043] In some embodiments, invalid characters can be highlighted in a single manner, for example, by adding only a strikethrough line to the invalid characters. In some embodiments, multiple highlighting modes can be superimposed on the invalid characters simultaneously, for example, by adding a strikethrough line and a background pattern of a specific color at the same time.
[0044] The manner in which invalid characters are prominently displayed can be selected according to the actual application. The embodiments of the present invention do not limit the prominent display manner.
[0045] To better understand some embodiments of the present invention, further explanation is provided below with reference to user interface diagrams.
[0046] Figure 3A shows a schematic diagram of an example of the interaction of an editing page 300 for audio editing according to several embodiments of the present invention. It should be understood that the pages shown in Figure 3A and the other drawings described below are illustrative only, and various page designs may actually exist. Each graphic element within a page may have a different arrangement and different visual representation, one or more elements may be omitted or replaced, and one or more other elements may be present. Embodiments of the present invention are not limited in this respect.
[0047] In editing page 300, content corresponding to audio 114 is presented in page area 310. Certain text is presented in the drawings for interpretive and illustrative purposes, but such text does not constitute any limitation to embodiments of the present invention. Further audio information related to audio 114 (also called related information for audio 114), including sound wave representation information 320 and time-length information 322 of audio 114, may be presented in editing page 300. In other embodiments, one or more of this audio information may not be presented.
[0048] The editing page 300 further provides one or more selectable editing functions. In the example in Figure 3A, function 330, labeled "Remove words with no actual meaning in one click," indicates a function to enter invalid character positioning mode. Figure 3A further shows other exemplary editing functions, including a split function 342 for dividing audio 114 into one or more audio segments, a volume adjustment function 344 for adjusting the volume of audio 114, a speed adjustment function 346 for adjusting the speed of audio 114, and a delete function 348 for deleting one or more audio segments of audio 114. The editing page 300 further presents a playback indicator 363 to indicate that audio is playing. In some implementations, the user can position one or more characters in the text or drag the progress control bar 312 to position the start position of audio playback.
[0049] Please understand that the text annotations in function 330 and the other editing functions shown are all examples. Editing page 300 may offer more, fewer, or other editing functions.
[0050] In Figure 3B, in response to the detection of a user selection for function 330, such as a user click on function 330, the terminal device 110 or the audio editing application 112 enters invalid character positioning mode. For interpretive and illustrative purposes, note that Figure 3B and several subsequent embodiments demonstrate user selection based on touch gestures. However, understand that other means of receiving user selection may exist, depending on the capabilities of the terminal device 110, such as mouse selection or voice control.
[0051] In some embodiments, when switching to invalid character positioning mode, the terminal device 110 can identify and position invalid characters in the text presented in the page area 310. As previously mentioned, invalid character identification may be performed on the terminal device 110's local or remote computing device 112, and may be performed after being triggered to enter invalid character positioning mode, or beforehand.
[0052] In some cases, as shown in Figure 3C, a positioning wait instruction 350 can be provided to indicate that an invalid character is being positioned in the page area 310. In some cases, invalid character identification may require a certain amount of time, or the positioning of the invalid character on the editing page 300 and the rendering of the prominent display of the invalid character may also require a certain amount of time. The positioning wait instruction 350 can present the user with the current operation of the terminal device 110.
[0053] After determining the invalid characters, invalid characters 360-1 "え", 360-2 "あの", and 360-3 "おん" are highlighted in the page area 310, as shown in Figure 3D. In this example, the invalid characters are highlighted by adding strikethrough lines and a colored background.
[0054] In some embodiments, in addition to prominently displaying invalid characters, additional information about the invalid characters may be displayed. This additional information may include at least the number of invalid characters that are prominently displayed. As shown in Figure 3D, on the editing page 300, a character indicator 362 for the number of invalid characters that are prominently displayed is displayed, and the number of invalid characters (e.g., "3") is also displayed on the "Delete Confirmation" option 372. Such display allows the user to quickly understand the total number of invalid characters in the text, and is particularly useful when the text is longer or there are many identified invalid characters. In some embodiments, as will be described further below, the number of displayed invalid characters can be dynamically modified as new invalid characters are successively selected by the user and / or as invalid characters are deselected.
[0055] By prominently displaying invalid characters, users can accurately identify characters that may be deleted and, depending on their editing needs, further confirm whether to delete one or more of these invalid characters. Returning to process 200 in Figure 2, in block 220, the terminal device 110 detects a delete confirmation instruction for at least one target invalid character in the invalid character set. The at least one target invalid character indicates an invalid character that has been confirmed to be deleted. In some embodiments, the delete confirmation instruction can also be detected based on user selection.
[0056] In some embodiments, a deletion confirmation option for invalid characters can be presented to the user for selection. For example, in the example in Figure 3E, a “deletion confirmation” option 372 is provided, and selecting this option triggers a deletion confirmation instruction.
[0057] In some embodiments, as described below, the user can also selectively check for automatically identified invalid characters and / or catch even more invalid characters.
[0058] If one or more prominently presented characters are determined not to need to be deleted, such as when these characters are determined to be "invalid characters" based on user selection, the remaining invalid characters are determined to be target invalid characters to be deleted.
[0059] Referring to Figure 2, in block 230, the terminal device 110 determines whether a delete confirmation instruction has been detected. In response to the detection of a delete confirmation instruction for at least one target invalid character, in block 240, the terminal device 110 obtains updated audio by deleting at least one audio portion corresponding to at least one target invalid character from audio 114. If no delete confirmation instruction for at least one target invalid character is detected, the terminal device 110 can continue to wait.
[0060] In some embodiments, the terminal device 110 can determine at least one audio portion corresponding to at least one target invalid character in audio 114 based on the temporal correspondence between audio 114 and text. The correspondence between audio 114 and text can indicate an audio portion corresponding to each text character or text string in the text, and can, for example, indicate timestamp information of the corresponding audio portion, including the start time and end time. In this way, after determining one or more target invalid characters to be deleted, these audio portions can be positioned in audio 114 by determining the timestamp information of the corresponding audio portions based on the correspondence.
[0061] After removing audio portions corresponding to one or more invalid target characters from audio 114, the updated audio may have a shorter duration. The updated audio can be constructed by concatenating the portions before and after the removed audio portion. In some embodiments, the updated audio itself may be stored locally or remotely by the terminal device 110 as a separate audio file.
[0062] In some embodiments, in addition to deleting the audio portion, updated text corresponding to the updated audio can be obtained by further deleting one or more target invalid characters identified from the text corresponding to audio 114. The updated text does not include the deleted target invalid characters. In some embodiments, the updated text can be further presented. In some embodiments, information related to the updated audio (also called related information), such as time-length information and / or sound wave representation information, can be further presented. When the audio is updated, this audio information may also be updated accordingly.
[0063] For example, if the user selects the “Delete Confirmation” option 372 in Figure 3E, the currently prominently displayed invalid characters 360-1, 360-2, and 360-3 are identified as target invalid characters. Therefore, the audio portions corresponding to these target invalid characters are deleted from audio 114, and these target invalid characters are also deleted from the text. As shown in Figure 3F, the updated text can be displayed in the text area 310 of the editing page 300, where the target invalid characters are no longer displayed.
[0064] Furthermore, on the editing page 300 in Figure 3F, updated audio-related information such as voiceprint representation information 324 and duration information 326 shown in Figure 3F is presented. This updated related information allows the user to visually see the results of deleting invalid characters in the audio. After deleting the audio portion corresponding to the invalid characters, the user can select the audio to play and hear the updated audio without the invalid characters.
[0065] As mentioned above, after entering invalid character positioning mode, automatically identified invalid characters are highlighted. Additionally or alternatively, the user can selectively check whether automatically identified invalid characters can be deleted and / or select a different invalid character to delete. Such embodiments are described in detail below.
[0066] Figure 4 shows a flowchart of process 400 for prominently displaying invalid characters according to several embodiments of the present invention. Process 400 can be implemented in terminal device 110. Process 400 in Figure 4 schematically illustrates the display of invalid characters determined based on automatic identification and manual user selection.
[0067] In block 410, the terminal device 110 presents text corresponding to the audio 114, such as the text shown in Figure 3A. In block 420, the terminal device 110 obtains invalid character identification results for the text. As mentioned above, the terminal device 110 can perform invalid character identification locally or receive invalid character identification results directly from a remote device. The invalid character identification results may be a set of invalid characters identified in the text, or they may indicate that no invalid characters were identified in the text.
[0068] In block 430, the terminal device 110 detects whether or not it will enter invalid character positioning mode. If it is not detected that it will enter invalid character positioning mode, the terminal device 110 can continue to wait. In response to the detection that it has entered invalid character positioning mode, such as when the user selects a corresponding function presented on the editing page 300 in Figure 3B, in block 440, the terminal device 110 determines whether or not there are invalid characters that have been automatically identified based on the invalid character identification result.
[0069] If automatically identified invalid characters exist, in block 450, the terminal device 110 prominently displays the set of automatically identified invalid characters. As shown in Figure 3E, the set of automatically identified invalid characters can be displayed in the text on the editing page 300.
[0070] In block 440, if the character recognition result is determined to indicate that there are no automatically identified invalid characters, then there are no characters that are automatically highlighted after entering invalid character positioning mode. In such cases, process 400 proceeds to box 460, where terminal device 110 detects a user selection for invalid characters in invalid character positioning mode. For example, the user can select a set of characters from the presented text as invalid characters. In other words, in invalid character positioning mode, the highlighted invalid characters include invalid characters selected and determined by the user.
[0071] In some embodiments, after the automatically identified invalid character set is prominently displayed in block 450, process 400 may proceed to box 460, where the terminal device 110 continues to detect user selections for invalid characters in invalid character positioning mode. In this case, the user can choose not to select one or more automatically identified invalid characters as target invalid characters because there is no way to delete them. Additionally or alternatively, in such cases, the user may also select one or more other characters as invalid characters.
[0072] In block 470, the terminal device 110 determines whether to prominently display invalid characters based on user selection. Depending on the user's specific selection, certain invalid characters may not be prominently displayed, while other specific invalid characters may be selected and prominently displayed.
[0073] In process 400, boxes 460 and 470 may be executed repeatedly until a confirmation instruction for deletion of a target invalid character is received. In response to such confirmation instruction, an invalid character that is still selected or prominently presented may be determined to be the target invalid character to be deleted.
[0074] In the following sections, in conjunction with Figures 5 and 6A to 6F, we will explain in detail the user selection and prominent display of invalid characters in the example of invalid characters on the editing page.
[0075] Figure 5 shows a flowchart of process 500 for prominently displaying invalid characters according to some other embodiments of the present invention. Process 500 may be implemented in terminal device 110. Process 500 can be considered an exemplary embodiment of boxes 460 and 470 of process 400. In process 500, first assume that one or more invalid characters are already prominently displayed. The currently prominently displayed invalid characters may include one or more automatically identified invalid characters and / or one or more invalid characters selected and determined by the user.
[0076] In block 510, the terminal device 110 determines whether it has received a deselection instruction for one or more invalid characters. The deselection instruction is determined based on user selection. For example, if one or more invalid characters are prominently displayed, the user can deselect a specific invalid character or a specific group of invalid characters within that group, thereby preventing these characters from being considered invalid. As shown in Figure 6A, invalid characters 360-1, 360-2, and 360-3 are prominently displayed on the editing page 300. When the user clicks on the invalid character 360-2 "ano", the terminal device 110 receives a deselection instruction for that invalid character.
[0077] In block 520, in response to receiving a deselection instruction, the terminal device 110 stops or downgrades the prominent presentation of one or more invalid characters that have been deselected. In some embodiments, in response to a deselection instruction, the terminal device 110 further removes one or more deselected invalid characters from the set of invalid characters, meaning that these characters are no longer considered invalid characters.
[0078] In one embodiment, the terminal device 110 can prevent one or more deselected invalid characters from being prominently displayed, so that the display of these invalid characters is the same as the display of other characters in the text. Figure 6B shows an example of the prominent display of a deselected invalid character. Specifically, in Figure 6A, after receiving a deselection instruction for invalid character 360-2, the character "ano" is no longer prominently displayed, as shown in Figure 6B.
[0079] In another embodiment, the terminal device 110 reduces the prominence of deselected invalid characters to less than that of other non-deselected invalid characters by downgrading the prominence of the deselected invalid characters. In some examples, deselected invalid characters are still presented conspicuously compared to other characters in the text to indicate to the user that these characters have been determined to be invalid characters (e.g., automatically identified invalid characters). Modes of downgrading the prominence of a character may include canceling some prominent presentation modes (where invalid characters are presented conspicuously in various modes), presenting them conspicuously according to another mode (where the degree of prominence in that mode is lower, e.g., by lowering the saturation of the background), and moderating any other prominent presentation.
[0080] Figure 6C shows the downgrading of the prominent presentation for an invalid character whose selection has been deselected. After receiving a deselection instruction for invalid character 360-2 in Figure 6A, the strikethrough line for invalid character 360-2 is canceled, as shown in Figure 6C, but the color background remains.
[0081] By providing a somewhat conspicuous display of invalid characters that have been deselected, it becomes easier for users to find these invalid characters again if they make a mistake or other error.
[0082] As mentioned above, when one or more invalid characters are deselected, for example when a deselection instruction for one or more invalid characters is received, the number of invalid characters displayed can be changed. For example, in the examples in Figures 6B and 6C, after the selection of invalid character 360-2 is deselected, character instruction 662 may be displayed on the editing page 300 to indicate the updated number of invalid characters. Alternatively, the number of invalid characters (e.g., "2") may also be displayed on the "Delete Confirmation" option 672.
[0083] Returning to process 530 in Figure 5, the terminal device 110 determines whether it has received a selective restoration instruction for one or more invalid characters. If it has received a selective restoration instruction, in block 540, the terminal device 110 restores the prominent presentation of the invalid characters from a stopped or downgraded state.
[0084] As shown in Figure 6C, the user can conveniently reposition and select invalid character 360-2 as needed. As shown in Figure 6D, upon receiving a user re-selection of invalid character 360-2, the prominent presentation of that invalid character 360-2 is restored to the same extent as the prominent presentation of other invalid characters. Of course, in such cases, the invalid character may be indicated as a restored invalid character in a different prominent presentation manner, but this is not limited to the present.
[0085] Furthermore, since the number of invalid characters increases after invalid character 360-2 is re-selected, in the example of Figure 6D, a character indication 664 for the number of updates for invalid characters may also be displayed on the editing page 300. Additionally, the number of invalid characters (e.g., "3") may also be displayed on the "Delete Confirmation" option 674.
[0086] In some embodiments, under various circumstances, the terminal device 110 determines in block 550 whether or not another character has been detected as an invalid character. For example, if a deselection instruction is detected in block 510, or if a selection restoration instruction is not detected in block 530, or after restoring the prominent display of an invalid character, the terminal device 110 can continue to determine whether or not another character has been detected as an invalid character. Although the steps in each box of the flowchart shown in Figure 5 have been described in order, it should be understood that these steps can be performed in a different order or in parallel. For example, the steps in boxes 510, 530, and 550 can be performed in parallel.
[0087] During the process of being in invalid character positioning mode, the terminal device 110 determines whether one or more other characters in the text have been selected as invalid characters based on user selection. For example, the user can select one or more characters in the text that are not prominently displayed as invalid characters.
[0088] In response to the detection that one or more other characters have been selected as invalid characters, in block 560, the terminal device 110 prominently displays the selected one or more invalid characters.
[0089] As shown in Figure 6E, the user selects a character that is not prominently displayed, such as character 660-1 "あ", within the text area 310. The terminal device 110 detects this user selection and decides to select the character as an invalid character. As shown in Figure 6F, the terminal device 110 prominently displays character 660-1 on the editing page 300. At this time, since the number of invalid characters increases from 2 to 3, the editing page 300 may also display a character indicator 666 indicating the number of updates for the invalid characters. Furthermore, the number of invalid characters (e.g., "3") may also be displayed on the "Delete Confirmation" option 676.
[0090] In some embodiments, when audio 114 is not playing, user selections such as deselection, selection restoration, and / or other invalid characters can be detected. In some examples, as shown in Figures 3D and 6A-6F, when the prominent display of invalid characters begins, instruction information 364 can be displayed on the editing page 300 to indicate that clicking on the highlighted portion while the audio is paused will allow the invalid characters to be retained or removed. As shown in Figures 6A-6F, the editing page 300 displays a playback pause indicator 663 to indicate that audio 114 is in a playback pause state.
[0091] If it is not detected that one or more other characters have been selected as invalid characters in block 550, the terminal device 110 may also determine that there is currently no need to prominently display another invalid character. During the process of being in invalid character positioning mode, the terminal device 110 can perform detection of boxes 510, 530, and 550 multiple times in succession.
[0092] Regardless of whether the user further edits the invalid characters, after detecting a delete confirmation instruction, the currently prominently displayed invalid characters are determined to be the target invalid characters to be deleted, and based on these target invalid characters, the audio portion corresponding to these target invalid characters can be deleted from audio 114. As shown in Figure 7A, when the user selects the “Delete Confirm” option 372, the currently prominently displayed invalid characters 360-1, 360-3, and 660-1 are confirmed as target invalid characters. Therefore, the audio portion corresponding to these target invalid characters is deleted from audio 114, and these target invalid characters are also deleted from the text. As shown in Figure 7B, the updated text can be displayed in the text area 310 of the editing page 300, where the target invalid characters 360-1, 360-3, and 660-1 are no longer displayed. In addition, voiceprint representation information 720 and time-length information 722 related to the updated audio can also be displayed.
[0093] In some embodiments, if invalid characters are obtained that have been selected by the user and decided to be deleted, the text corresponding to the user-selected invalid characters and audio 114 may be provided to tune the invalid character recognition model. For example, the text corresponding to the user-selected character "あ" and audio 114 in Figures 6E and 6F may be provided to train the character recognition model. This character recognition model may be a model used by the terminal device 110 or the remote computing device 112 to automatically identify invalid characters. The provided invalid characters and the text corresponding to audio 114 can enrich and expand the training dataset of the invalid character recognition model, thereby allowing the invalid character recognition model to evolve to have stronger recognition capabilities.
[0094] In some embodiments, training of the invalid character recognition model can be resumed after a sufficient amount of separate training data has been collected. In some embodiments, training of the invalid character model may be performed on terminal device 110, computing device 112, or other model training device. Embodiments of the present invention are not limited in this respect.
[0095] Figure 8 shows a schematic block diagram of the configuration of an audio editing apparatus 800 according to a specific embodiment of the present invention. The apparatus 800 may be implemented as a terminal device 18 or contained within a terminal device 18. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0096] As shown in the figure, the device 800 includes a highlighting module 810 for highlighting one or more invalid characters contained in the text corresponding to the audio in a predefined mode for audio. The device 800 further includes an instruction detection module 820 for detecting a delete confirmation instruction for at least one target invalid character among the one or more invalid characters, and an audio delete module 830 for deleting at least one audio portion corresponding to at least one target invalid character from the audio in response to the detection of a delete confirmation instruction.
[0097] In some embodiments, the device 800 further comprises an invalid character identification module for identifying a first invalid character from text.
[0098] In some embodiments, the device 800 further comprises an invalid character determination module for determining a second invalid character in the text based on user input.
[0099] In some embodiments, the device 800 further comprises a data supply module for providing a second invalid character and text for training an invalid character recognition model, the data supply module being trained to identify invalid characters from input text.
[0100] In some embodiments, the instruction detection module includes an invalid character removal module for removing a third invalid character from one or more invalid characters in response to receiving a deselection instruction for a third invalid character among one or more invalid characters.
[0101] In some embodiments, the device 800 further includes a stop or downgrade module for stopping or downgrading the prominent presentation for the fourth invalid character in response to receiving a deselection instruction for the fourth invalid character among one or more invalid characters.
[0102] In some embodiments, the apparatus 800 further comprises a number presentation module for presenting a first number of one or more invalid characters.
[0103] In some embodiments, the device 800 further comprises a number determination module for determining a second number of invalid characters that have not been deselected from the one or more invalid characters in response to receiving a deselection instruction for at least one of the one or more invalid characters, and a number modification module for modifying the presented first number to the second number.
[0104] In some embodiments, the device 800 further comprises a text determination module for obtaining updated text by deleting at least one target invalid character from the text in response to the detection of a deletion confirmation instruction for at least one target invalid character, and a text presentation module for presenting the updated text.
[0105] In some embodiments, the apparatus 800 further comprises an information module for presenting relevant audio information after at least one audio portion has been removed, wherein the relevant information includes at least one of duration and sound wave representation.
[0106] Figure 9 shows a block diagram of a computing device 900 that can carry out one or more embodiments of the present invention. It should be understood that the computing device 900 shown in Figure 9 is illustrative and should not limit the functionality and scope of the embodiments described herein. The computing device 900 shown in Figure 9 may be used to implement the terminal device 110 of Figure 1.
[0107] As shown in Figure 9, the computing device 900 is in the form of a general-purpose electronic device. The components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage devices 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be an actual or virtual processor and can perform various processes based on a program stored in memory 920. In a multiprocessor system, the parallel processing capability of the computing device 900 is improved by having multiple processing units execute computer executable instructions in parallel.
[0108] The computing device 900 typically includes multiple computer storage media. Such media may include, but are not limited to, volatile and non-volatile media, removable and non-removable media, and may be any obtainable media accessible by the computing device 900. Memory 920 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or a specific combination thereof. Storage device 930 may be removable or non-removable media, and may include machine-readable media such as flash memory drives, magnetic disks, or any other media, and may be used to store information and / or data (e.g., training data for training) and may be accessible within the computing device 900.
[0109] The computing device 900 may further include other removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 9, a removable magnetic disk drive for reading from or writing to a non-volatile magnetic disk (e.g., a "floppy disk") and a removable optical disk drive for reading from or writing to a non-volatile optical disk may be provided. In these cases, each drive may be connected to a path (not shown) by one or more data medium interfaces. The memory 920 may also include a computer program product 925 having one or more program modules, which are configured to perform various methods or operations of various embodiments of the present invention.
[0110] The communication unit 940 implements communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 900 may be implemented as a single computing cluster or as multiple computing machines, which can communicate via communication connections. Therefore, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0111] The input device 950 may be one or more input devices such as a mouse, keyboard, or trackball. The output device 960 may be one or more output devices such as a display, speaker, or printer. The computing device 900 may also, if necessary, communicate with one or more external devices (not shown) such as a storage device or display device via the communication unit 940, or it may communicate with one or more devices that enable a user to interact with the computing device 900, or the computing device 900 may communicate with any device (such as a netcard or modem) that communicates with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0112] According to an exemplary implementation of the present invention, a computer-readable storage medium is provided in which one or more computer instructions are stored, and the one or more computer instructions are executed by a processor to implement the above method. According to an exemplary implementation of the present invention, a computer program product is further provided, the computer program product being tangibly stored in a non-temporary computer-readable medium and including computer-executable instructions that are executed by a processor to implement the above method.
[0113] Herein, each aspect of the present invention has been described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products implemented by the present invention. It should be understood that each box in the flowcharts and / or block diagrams, and each combination of boxes in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0114] These computer-readable program instructions may be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine that, when these instructions are executed by the computer or other programmable data processing device's processing unit, generates a device for implementing the functions / operations specified in one or more boxes in a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, and these instructions may cause the computer, programmable data processing device, and / or other device to operate in a particular manner so that the computer-readable medium on which the instructions are stored constitutes a product containing instructions for implementing each aspect of the functions / operations specified in one or more boxes in a flowchart and / or block diagram.
[0115] By loading computer-readable program instructions into a computer, other programmable data processing device, or other device, a series of operational steps are performed on the computer, other programmable data processing device, or other device to generate a process implemented by the computer, so that the instructions executed on the computer, other programmable data processing device, or other device implement the function / operation specified in one or more boxes in a flowchart and / or block diagram.
[0116] The flowcharts and block diagrams in the drawings illustrate the implementable architectures, functions, and operations of several implementable systems, methods, and computer program products according to the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and a module, program segment, or part of an instruction contains one or more executable instructions for implementing a specified logical function. In some implementations as replacements, the functions represented in the boxes may occur in a different order than those shown in the drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or in reverse order depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented by a special-purpose hardware-based system that performs a specified function or operation, or by a combination of special-purpose hardware and computer instructions.
[0117] While the various implementations of the present invention have been described above, the above descriptions are illustrative, not exhaustive, and not limited to the implementations disclosed. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of each implementation described. The choice of terms used herein is intended to best interpret the principles, practical applications, or improvements to the technology in the market of each implementation, or to enable those skilled in the art to understand each implementation disclosed herein.
Claims
1. In a predefined mode for audio, one or more invalid characters contained in the text corresponding to the audio are highlighted, To detect a deletion confirmation instruction for at least one target invalid character among the one or more invalid characters mentioned above, In response to the detection of the deletion confirmation instruction, the following steps are taken: delete at least one audio portion from the audio corresponding to the at least one target invalid character; and delete the at least one target invalid character from the text to obtain updated text. How to edit audio.
2. The method further includes identifying a first invalid character from the text, wherein the first invalid character is determined by automatic identification. The method for editing audio according to claim 1.
3. The further includes determining a second invalid character in the text based on user input, The method for editing audio according to claim 1.
4. To train an invalid character recognition model, the present invention provides the second invalid character and the text, further comprising the invalid character recognition model being trained to identify invalid characters from the input text. The method for editing audio according to claim 3.
5. Detecting the aforementioned deletion confirmation instruction means In response to receiving a deselection instruction for a third invalid character among the one or more invalid characters, the third invalid character is removed from the one or more invalid characters. The method for editing audio according to claim 1.
6. In response to receiving a deselection instruction for the fourth invalid character among the one or more invalid characters, the system further includes stopping or downgrading the prominent presentation of the fourth invalid character. The method for editing audio according to claim 1.
7. Further comprising presenting a first number of the one or more invalid characters, The method for editing audio according to claim 1.
8. In response to receiving a deselection instruction for at least one of the one or more invalid characters, the one or more To determine the second number of invalid characters that have not been deselected from among the invalid characters, Further including modifying the presented first number to the second number, The method for editing audio according to claim 7.
9. Presenting related information of the audio after deleting at least one of the aforementioned audio portions, further comprising the relevant information including at least one of duration and sound wave representation, The method for editing audio according to claim 1.
10. In a predefined mode for audio, a highlighting module for highlighting one or more invalid characters contained in the text corresponding to the audio, An instruction detection module for detecting a deletion confirmation instruction for at least one target invalid character among the one or more invalid characters, An audio deletion module for deleting at least one audio portion corresponding to the at least one target invalid character from the audio in response to the detection of the deletion confirmation instruction, A text determination module for obtaining updated text by deleting at least one target invalid character from the text in response to the detection of the deletion confirmation instruction, A text presentation module for presenting the updated text is provided. A device for audio editing.
11. An invalid character identification module for identifying a first invalid character from the text, wherein the first invalid character is determined by automatic identification, further comprising the invalid character identification module, The apparatus for audio editing according to claim 10.
12. The system further comprises an invalid character determination module for determining a second invalid character in the text based on user input. The apparatus for audio editing according to claim 10.
13. A data supply module for providing the second invalid character and the text for training an invalid character recognition model, further comprising a data supply module in which the invalid character recognition model is trained to identify invalid characters from the input text. The apparatus for audio editing according to claim 12.
14. The instruction detection module is, The system includes an invalid character removal module for removing the third invalid character from the one or more invalid characters in response to receiving a deselection instruction for the third invalid character among the one or more invalid characters. The apparatus for audio editing according to claim 10.
15. The system further comprises a prominent display stop or downgrade module for stopping or downgrading prominent display of the fourth invalid character in response to receiving a deselection instruction for the fourth invalid character among the one or more invalid characters, The apparatus for audio editing according to claim 10.
16. The system further comprises a number presentation module for presenting a first number of the one or more invalid characters. The apparatus for audio editing according to claim 10.
17. It is an electronic device, At least one processing unit, The device comprises at least one memory coupled to the at least one processing unit and used to store instructions executed by the at least one processing unit, and for causing the electronic device to perform the method according to any one of claims 1 to 8 when the instructions are executed by the at least one processing unit, The aforementioned electronic device.
18. A computer program is stored, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 8. Computer-readable storage medium.
Citation Information
Patent Citations
Video processing method, device, electronic equipment and storage medium
CN113613068A
Voice reproduction device, voice reproduction method, and program
JP2017111339A
Audio data processing apparatus, audio data processing method, program, and recording medium
JP2018195896A
Voice recognition device and system
JP2019095644A
Method and User Interface for Creating an Audio Recording Using a Document Paradigm
US20090082887A1