Method, apparatus, device and storage medium for audio editing - Patents.com
The method and apparatus facilitate efficient audio editing by prominently displaying invalid characters for user selection and confirmation, automating the deletion process to improve accuracy and reduce errors.
Patent Information
- Application Number
- JP2024563831
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-06
- Filing Date
- 2023-05-05
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-05-05
AI Technical Summary
Conventional audio editing methods require users to repeatedly listen to audio to locate and delete unwanted characters or words, leading to inefficiency and errors such as omissions and erroneous deletions.
An audio editing method and apparatus that prominently presents invalid characters in text corresponding to audio, allowing users to select and confirm deletion, automatically removing the corresponding audio portions.
Enhances audio editing efficiency by enabling one-click identification and deletion of invalid characters, reducing redundant operations and minimizing errors.
Smart Images

Figure 2025515613000001_ABST
Abstract
Description
[Technical field]
[0001] [Cross-reference to related applications] This application claims priority to a Chinese invention patent application filed on May 6, 2022, entitled "Method, Apparatus, Device, and Storage Medium for Audio Editing" and bearing application number 202210488246.2.
[0002] [Technical field] FIELD OF THE DISCLOSURE Exemplary embodiments of the present invention relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for audio editing. [Background technology]
[0003] Audio data is a common information interaction method in all aspects of people's lives, work, social interactions, etc. Nowadays, people can produce and obtain audio data more and more conveniently, and can also share recorded audio. In order to output high-quality audio, it is expected to perform various editing operations on the audio data, such as adjusting the volume, speed, timbre, etc. In some cases, it is also expected to delete words that are not expected to appear from the audio data. Summary of the Invention
[0004] According to an exemplary embodiment of the present invention, a solution for audio editing is provided.
[0005] In a first aspect of the present invention, there is provided a method of audio editing, the method including prominently presenting, in a predefined mode for the audio, one or more invalid characters included in text corresponding to the audio, the method further including detecting a deletion confirmation indication for at least one target invalid character of the one or more invalid characters, and deleting from the audio at least one audio portion corresponding to the at least one target invalid character in response to detecting the deletion confirmation indication.
[0006] In a second aspect of the present invention, there is provided an apparatus for editing audio, the apparatus comprising: a conspicuous presentation module for conspicuously presenting one or more invalid characters included in a text corresponding to the audio in a predefined mode for the audio, an indication detection module for detecting a deletion confirmation indication for at least one target invalid character of the one or more invalid characters, and an audio deletion module for deleting at least one audio portion corresponding to the at least one target invalid character from the audio in response to the deletion confirmation indication being detected.
[0007] In a third aspect of the present invention there is provided an electronic device comprising at least one processing unit and at least one memory coupled to the at least one processing unit for use in storing instructions executed by the at least one processing unit which, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0008] In a fourth aspect of the present invention, there is provided a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method of the first aspect.
[0009] It should be understood that the contents described in the summary of the present invention are not intended to limit the main or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will be readily understood from the following description. [Brief description of the drawings]
[0010] The above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent from the following detailed description taken in conjunction with the drawings, in which like or similar symbols indicate like or similar elements, and in which: [Figure 1] FIG. 1 illustrates a schematic diagram of an exemplary environment in which embodiments of the present invention may be implemented. [Diagram 2] FIG. 2 illustrates a flowchart of a process for audio editing according to some embodiments of the present invention. [Figure 3A] FIG. 2 shows a schematic diagram of an example of an edit page interaction for audio editing according to some embodiments of the present invention. [Figure 3B] FIG. 2 shows a schematic diagram of an example of an edit page interaction for audio editing according to some embodiments of the present invention. [Figure 3C] FIG. 2 shows a schematic diagram of an example of an edit page interaction for audio editing according to some embodiments of the present invention. [Figure 3D] FIG. 2 shows a schematic diagram of an example of an edit page interaction for audio editing according to some embodiments of the present invention. [Figure 3E] FIG. 2 shows a schematic diagram of an example of an edit page interaction for audio editing according to some embodiments of the present invention. [Figure 3F] FIG. 2 shows a schematic diagram of an example of an edit page interaction for audio editing according to some embodiments of the present invention. [Figure 4] FIG. 1 illustrates a flowchart of a process for prominently presenting invalid characters according to some embodiments of the present invention. [Diagram 5] FIG. 10 illustrates a flowchart of a process for prominently presenting invalid characters according to some alternative embodiments of the present invention. [Figure 6A] 1A-1C are schematic diagrams illustrating examples of user selections for invalid characters on an edit page according to some embodiments of the present invention. [Figure 6B] 1A-1C are schematic diagrams illustrating examples of user selections for invalid characters on an edit page according to some embodiments of the present invention. [Figure 6C] 1A-1C are schematic diagrams illustrating examples of user selections for invalid characters on an edit page according to some embodiments of the present invention. [Figure 6D]1A-1C are schematic diagrams illustrating examples of user selections for invalid characters on an edit page according to some embodiments of the present invention. [Figure 6E] 1A-1C are schematic diagrams illustrating examples of user selections for invalid characters on an edit page according to some embodiments of the present invention. [Figure 6F] 1A-1C are schematic diagrams illustrating examples of user selections for invalid characters on an edit page according to some embodiments of the present invention. [Figure 7A] FIG. 2 shows a schematic diagram of an exemplary page presentation of an audio edit according to some embodiments of the present invention. [Figure 7B] FIG. 2 shows a schematic diagram of an exemplary page presentation of an audio edit according to some embodiments of the present invention. [Figure 8] FIG. 1 shows a block diagram of an apparatus for audio editing according to some embodiments of the present invention. [Figure 9] FIG. 1 shows a block diagram of a device capable of implementing several embodiments of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] In the following, the embodiments of the present invention will be described in more detail with reference to the drawings. Although the drawings show specific embodiments of the present invention, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, but rather, these embodiments are provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes and are not used to limit the protection scope of the present invention.
[0012] In describing embodiments of the present invention, the term "comprising" and similar terms are intended to be open-ended inclusions such as "including, but not limited to." The term "based on" should be understood as "based at least in part on." The terms "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit and implicit definitions may be included below.
[0013] It is understood that any data related to the technical solution (including but not limited to the data itself, the acquisition of the data, or the use of the data) should comply with the corresponding laws and regulations and related specified requirements.
[0014] It is understood that before using the technical solutions disclosed in each embodiment of the present invention, the types, scope of use, usage scenarios, etc. of personal information related to the present invention should be notified to users in an appropriate manner in accordance with relevant laws and regulations, and consent from users should be obtained.
[0015] For example, in response to receiving an unsolicited request from a user, presenting information is sent to the user to explicitly present to the user that the requested operation requires the acquisition and use of the user's personal information, so that the user can independently choose whether or not to provide the personal information to software or hardware, such as an electronic device, application, server, or storage medium, that performs the operation of the technical solution of the present invention, based on the presenting information.
[0016] As an optional, non-limiting implementation, the method of transmitting the presentation to the user in response to receiving the user's unsolicited request may be, for example, by utilizing a pop-up window in which the presentation may be displayed in the form of text and which may further include a selection control for the user to "agree" or "not agree" to providing the personal information to the electronic device.
[0017] It will be appreciated that the notification and user authorization process described above is merely a general outline and is not intended to limit the implementation of the present invention, and other means that comply with relevant laws and regulations may be applied to the implementation of the present invention.
[0018] 1 shows a schematic diagram of an exemplary environment 100 in which an embodiment of the present invention can be implemented. In the exemplary environment 100, an audio editing application 112 for editing audio 114 may be installed on a terminal device 110. For example, the audio editing application 112 may edit the audio 114 based on an operation of a user 102. In this specification, the audio 114 to be edited may be in any audio format and may have any suitable audio length. As an example, the audio 114 may be a podcast, an audio corresponding to a short video, a radio drama, an audiobook, a conference or interview recording, an audiobook course, an audio note, etc.
[0019] In some embodiments, audio 114 may be collected by an audio collection device 105 (e.g., a device with a microphone) and provided to an audio editing application 112 to be edited. For example, the audio collection device 105 may collect audio from at least a user 104. In some embodiments, the audio editing application 112 may provide an audio recording function for recording the audio 114 collected by the audio collection device 105. In some embodiments, the audio 114 edited by the audio editing application 112 may be from any other data source, such as audio 114 downloaded or received from another device. Embodiments of the invention are not limited in this respect.
[0020] Although a user 102 performing editing operations on audio 114 and a user 104 outputting audio 114 are shown, it will be understood that these users may be the same user and are not limited herein. Additionally, although shown as separate devices, it will be understood that the audio collection device 105 can be integrated with the terminal device 110. In other implementations, the audio collection device 105 can be communicatively connected to the terminal device 110 by other means to collect and provide the audio 114.
[0021] The terminal device 110 may be any type of mobile, fixed, or portable terminal, including a mobile cell phone, a desktop computer, a laptop computer, a mobile cell phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communications system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio receiver, an e-book device, a gaming device, or any combination of the foregoing, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal device 110 may also support any type of interface to a user, such as "wearable" circuitry.
[0022] In some embodiments, the terminal device 110 can communicate with a remote computing device 122 to implement editing of the audio 114. For example, the computing device 122 can perform storage functions, specific analysis tasks, etc. on the audio 114 to extend the storage and processing capabilities of the terminal device 110. The computing device 122 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc. In the example shown in FIG. 1, the computing device 122 can be located in the cloud environment 120.
[0023] It should be understood that the configuration and functionality of the environment 100 is described for illustrative purposes only and does not imply any limitation on the scope of the invention. For example, the terminal device 110 may not communicate with the remote computing device 122. Also, for example, the user 104 and the audio collection device 105 may be omitted.
[0024] In an audio editing scenario, it may be expected to remove characters or words in the audio that are not expected to appear, such as characters or words that are meaningless or useless to the expression in the audio. In this specification, such characters or words may be referred to as "invalid characters", and may also be referred to as "invalid words", "useless words", "words with no real meaning", or "useless words", in which "invalid characters" may be text units of different sizes, such as a single character, word, or group of words, and may have different sizes in different natural languages. In some examples, the invalid characters may include modal particles, mantras, etc. that appear in the spoken expression, such as "a", "ya", "on", "e", "kono", "ano", etc. in the Chinese expression, and these meaningless words are considered as invalid expressions. In some examples, the invalid characters may additionally or alternatively include other characters or words that are not expected to appear in the audio, such as sensitive words. In different application scenarios, the sensitive words that are not expected to appear may be different, which is specified as necessary.
[0025] In conventional solutions, to delete characters or words that are not expected to appear in the audio, an audio editor must repeatedly listen to the audio to find and accurately locate the characters or words to be deleted, and then select and delete the corresponding audio portion. Such an editing process is inefficient and prone to many problems, such as omissions and erroneous deletions (e.g., the audio portion to be deleted is too long or too short).
[0026] According to an embodiment of the present invention, an improved audio editing solution is proposed, in which one or more invalid characters present in a text base corresponding to the audio are determined and prominently displayed, so that a user can select a specific invalid character or a specific number of invalid characters therein and confirm whether to delete them. After a deletion confirmation instruction for the invalid characters is detected, audio portions corresponding to the invalid characters confirmed to be deleted from the audio are automatically deleted.
[0027] The solution can support convenient deletion of invalid characters in audio, and can greatly improve the efficiency of audio editing. From the user's perspective, it can implement one-click identification and deletion of invalid characters, avoiding redundant operations and saving time in audio editing. It provides the user with a prominent indication of potential invalid characters that can be deleted, and can effectively avoid erroneous deletion, omission, etc.
[0028] In the following, with continuing reference to the drawings, some exemplary embodiments of the present invention will be described.
[0029] 2 shows a flow chart of a process 200 for audio editing according to some embodiments of the present invention. The process 200 may be implemented in a terminal device 110. For ease of explanation, the process 200 will be described with reference to the environment 100 of FIG.
[0030] In block 210, the terminal device 110 prominently presents invalid character sets in the text corresponding to the audio 114 in a predefined mode for the audio.
[0031] In embodiments of the present invention, corresponding text is identified from the audio 114 to aid in editing the audio 114. In some embodiments, automatic speech recognition (ASR) techniques can be used to identify the corresponding text from the audio 114. The text identification can be performed at the terminal device 110. In other embodiments, the text identification can be performed by a remote computing device, such as the computing device 122 in the environment 100. The terminal device 110 can also receive the text from the computing device 122.
[0032] In some embodiments of the present invention, a predefined mode is provided in which an invalid character set within text can be located and prominently presented, the invalid character set including one or more invalid characters. For ease of explanation, the predefined mode is hereinafter referred to as an "invalid character location mode." In some embodiments, the invalid character location mode can be entered in response to a user selection.
[0033] In an embodiment of the present invention, the invalid characters to be prominently presented are determined on a text basis. In some embodiments, the invalid characters to be prominently presented may include one or more invalid characters that are automatically identified from the text. Automatically identifying invalid characters can save a user's identification time for invalid characters. In particular, compared with the method of locating invalid characters by listening to audio, automatic identification can more quickly present the presence of invalid characters to a user. In this way, after being triggered to enter the invalid character location mode, the invalid characters identified from the text can be automatically and quickly prominently presented.
[0034] In some other embodiments described below, the prominently presented invalid characters may additionally or alternatively include one or more invalid characters selected and determined by a user, for example, allowing a user to select one or more characters from the presented text as invalid characters, allowing a user to more easily and accurately identify invalid characters in the text as compared to locating invalid characters by listening to audio.
[0035] In some embodiments, the automatic identification of invalid characters may be performed at the terminal device 110. In other embodiments, the automatic identification of invalid characters may be performed by a remote computing device, such as the computing device 122 in the environment 100. The terminal device 110 may also obtain the automatically identified set of invalid characters from the computing device 122.
[0036] Various methods can be used to automatically identify invalid characters in text. In some embodiments, a preselected invalid character list can be created and maintained that includes common invalid characters such as "a", "ya", "on", "e", "kono", "ano", etc., and / or other characters or words that are not expected to appear in audio, such as sensitive words. Each character in the text corresponding to the audio 114 can be matched against the invalid character list to determine which invalid characters are present in the text. It should be understood that only non-limiting examples of invalid characters are listed here, and that in different languages and application scenarios, more, fewer, or other invalid characters may be included in the invalid character list.
[0037] Alternatively or additionally, in some embodiments, an invalid character identification model can be constructed and trained, the model being configured to identify invalid characters from input text. Such invalid character identification models can be constructed and trained based on various machine learning or deep learning algorithms. The input of the invalid character identification model can include text, and the output can include an identification result. The identification result indicates whether invalid characters are present in the text, and if present, further includes an indication of the identified invalid characters.
[0038] The training data for training such an invalid character identification model may include a sample text, and may further include annotation information of invalid characters in the sample text. Note that, the invalid character identification model may be constructed using a machine learning or deep learning model suitable for text processing, and the model may be trained using an appropriate training algorithm of machine learning or deep learning. The embodiment of the present invention does not specifically limit the configuration and training process of the invalid character identification model.
[0039] It will be appreciated that whether based on the invalid character list or based on the invalid character identification model, invalid character identification can be performed at the terminal device 110's local or remote computing device 122. In some embodiments, invalid character identification can be initiated after receiving a trigger to enter the invalid character location mode. In some embodiments, invalid character identification can be performed asynchronously, for example, the terminal device 110 or the computing device 112 can identify an invalid character set from the text corresponding to the audio 114 after obtaining the audio 114 and record these identified invalid characters. Then, after entering the invalid character location mode, the previously identified invalid characters can be quickly and prominently presented.
[0040] In some embodiments, editing of the audio 114 may be performed in the audio editing application 112, such as removing portions of the audio that correspond to invalid characters. For example, the audio editing application 112 may provide an edit page for the audio 114. The audio editing application 112 may provide an invalid character location mode, in which invalid character sets within the text are highlighted on the edit page. In some embodiments, the text may be presented on the edit page, and the invalid character sets may be highlighted during the presentation of the text.
[0041] Conspicuous presentation of invalid characters means that the invalid characters are displayed differently from other characters in the text. Conspicuous presentation for invalid characters can be implemented using one or various conspicuous presentation aspects. As an example, conspicuous presentation aspects can include adding a strikethrough (i.e., drawing a line through the middle of the character) or an underline to the invalid character, changing the format of the invalid character (e.g., color, font size, typeface, and / or weight) to distinguish the invalid character from other characters, overlaying a background pattern of a particular color or shape on the invalid character, adding a special shape or annotation on the invalid character, any other aspect that can make the invalid character conspicuous, etc.
[0042] In some implementations, the invalid character may be prominently presented by altering the presentation of other characters other than the invalid character when other characters in the text are simultaneously presented, for example by changing the formatting (e.g., color, font size, typeface and / or weight) of the other characters or by making the other characters invisible or at least partially invisible.
[0043] In some embodiments, the invalid character may be prominently displayed in a single manner, e.g., only a line of deletion may be added to the invalid character. In some embodiments, the invalid character may be prominently displayed in multiple manners, e.g., a line of deletion and a background pattern of a particular color may be added simultaneously.
[0044] The conspicuous manner of the invalid characters can be selected according to the actual application, and the embodiment of the present invention does not limit the conspicuous manner.
[0045] To better understand some embodiments of the present invention, the following description will be further illustrated with reference to user interface diagrams.
[0046] 3A shows a schematic diagram of an example of an edit page 300 interaction for audio editing according to some embodiments of the present invention. It should be understood that the page shown in FIG. 3A and the pages of the other figures described below are merely exemplary, and that a variety of page designs may exist in practice. Each graphic element within a page may have a different arrangement and a different visual representation, one or more elements therein may be omitted or substituted, and one or more other elements may be present. Embodiments of the present invention are not limited in this respect.
[0047] In the edit page 300, content corresponding to the audio 114 is presented in the page region 310. Although certain text has been presented in the figures for purposes of interpretation and explanation, such text does not constitute any limitation on the embodiments of the present invention. In the edit page 300, audio information related to the audio 114 (also referred to as related information of the audio 114) may further be presented, including sound wave representation information 320 and duration information 322 of the audio 114. In other embodiments, one or more of these pieces of audio information may not be presented.
[0048] The edit page 300 further provides one or more selectable edit functions. In the example of FIG. 3A, a function 330 labeled "Remove words with no real meaning with one click" indicates a function for entering an invalid character location mode. FIG. 3A further illustrates other exemplary edit functions including a split function 342 for splitting the audio 114 into one or more audio segments, a volume adjustment function 344 for adjusting the volume of the audio 114, a speed adjustment function 346 for adjusting the speed of the audio 114, a delete function 348 for deleting one or more audio segments of the audio 114, etc. The edit page 300 further presents a playback indicator 363 that indicates that the audio is playing. In some implementations, the user can position one or several characters in the text or drag a progress control bar 312 to position the start position of the audio playback.
[0049] It should be understood that the text annotations of feature 330 and the other editing features shown are both examples, and that edit page 300 may provide more, fewer, or other editing features.
[0050] In response to detecting a user selection for the function 330, such as a user click selection for the function 330 in Fig. 3B, the terminal device 110 or audio editing application 112 enters a null character location mode. Note that for purposes of interpretation and explanation, in Fig. 3B and in some subsequent examples, user selection based on touch gestures is illustrated. However, it should be understood that there may be other means of receiving user selection, such as mouse selection, voice control, etc., depending on the capabilities of the terminal device 110.
[0051] In some implementations, switching to the invalid character location mode enables the terminal device 110 to identify and locate invalid characters within the text presented in the page area 310. As previously mentioned, invalid character identification may be performed local to the terminal device 110 or at the remote computing device 112, and may be performed after being triggered to enter the invalid character location mode, or may be performed in advance.
[0052] 3C, a positioning wait indication 350 may be provided to indicate locating the invalid character in the page area 310. In some cases, invalid character identification may require a certain amount of time, or locating the invalid character in the edit page 300 and rendering the invalid character prominently may also require a certain amount of time. The positioning wait indication 350 may present the current operation of the terminal device 110 to the user.
[0053] After the invalid characters are determined, as shown in Fig. 3D, invalid character 360-1 "e", invalid character 360-2 "ano", and invalid character 360-3 "on" are prominently displayed in page region 310. In this example, the invalid characters are prominently displayed by adding deletion lines and color background patterns.
[0054] In some embodiments, in addition to prominently displaying the invalid characters, additional information of the invalid characters may be presented. The additional information may include at least the number of invalid characters prominently displayed. As shown in FIG. 3D, in the edit page 300, a character indication 362 for the number of invalid characters prominently displayed is presented, and the number of invalid characters (e.g., the number "3") is also displayed on the "confirm delete" option 372. Such presentation allows the user to quickly understand the total number of invalid characters in the understood text, which is particularly useful when the text is longer or when there are many identified invalid characters. In some embodiments, the number of invalid characters presented may be dynamically modified as new invalid characters are successively selected and / or invalid characters are deselected by the user, as described subsequently below.
[0055] By prominently presenting the invalid characters, the user can accurately understand the characters that may be deleted, and can further confirm whether to delete one or more of these invalid characters according to editing needs. Returning to the process 200 of FIG. 2, in block 220, the terminal device 110 detects a deletion confirmation indication for at least one target invalid character of the invalid character set. The at least one target invalid character indicates an invalid character that is confirmed to be deleted. In some embodiments, the deletion confirmation indication can also be detected based on a user selection.
[0056] In some implementations, a confirm delete option for the invalid character may be presented for the user to select. For example, in the example of FIG. 3E, a "confirm delete" option 372 is provided, selection of which triggers a confirm delete prompt.
[0057] In some embodiments, as described below, the user may also selectively confirm the automatically identified invalid characters and / or supplement more invalid characters.
[0058] If it is determined that one or more prominently presented characters do not require deletion, such as by determining that these characters are "non-invalid characters" based on user selection, the remaining invalid characters are determined as target invalid characters to be deleted.
[0059] 2, in block 230, the terminal device 110 determines whether a delete confirmation indication is detected. In response to detecting a delete confirmation indication for the at least one target invalid character, in block 240, the terminal device 110 obtains updated audio by deleting at least one audio portion corresponding to the at least one target invalid character from the audio 114. If a delete confirmation indication for the at least one target invalid character is not detected, the terminal device 110 may continue to wait.
[0060] In some embodiments, the terminal device 110 can determine at least one audio portion corresponding to at least one target invalid character in the audio 114 based on a temporal correspondence between the audio 114 and the text. The correspondence between the audio 114 and the text can indicate an audio portion corresponding to each text character or text string in the text, and can indicate time stamp information of the corresponding audio portion, including, for example, a start time and an end time. In this way, after determining one or more target invalid characters to be removed, these audio portions can be located in the audio 114 by determining time stamp information of the corresponding audio portion based on the correspondence.
[0061] After deleting audio portions corresponding to one or more target invalid characters from audio 114, the updated audio may have a shorter duration. The portions of the audio before and after the deleted portions may be concatenated to form the updated audio. In some embodiments, the updated audio itself may be stored locally or remotely by terminal device 110 as a separate audio file.
[0062] In some embodiments, in addition to deleting the audio portion, an updated text corresponding to the updated audio may be obtained by further deleting one or more identified target invalid characters from the text corresponding to the audio 114. The updated text does not include the deleted target invalid characters. In some embodiments, the updated text may be further presented. In some embodiments, information related to the updated audio (also referred to as related information), such as duration information and / or sound wave representation information, may be further presented. When the audio is updated, the audio information may be updated accordingly.
[0063] For example, if the user selects the "Confirm Delete" option 372 in Figure 3E, the currently prominently presented invalid characters 360-1, 360-2, and 360-3 are confirmed as target invalid characters. Thus, the audio portions corresponding to these target invalid characters are deleted from the audio 114, and these target invalid characters are also deleted from the text. As shown in Figure 3F, the updated text can be submitted in the text area 310 of the edit page 300, where the target invalid characters are no longer submitted.
[0064] In addition, the edit page 300 in Fig. 3F further presents related information of the updated audio, such as voiceprint expression information 324 and duration information 326 shown in Fig. 3F. The updated related information allows the user to visually see the deletion result of the invalid characters in the audio. After deleting the audio portion corresponding to the invalid characters, when the user selects play audio, the updated audio without the invalid characters can be heard.
[0065] As previously discussed, after entering the invalid character location mode, the automatically identified invalid characters are prominently presented, and, additionally or alternatively, the user may selectively confirm whether the automatically identified invalid characters can be deleted and / or select additional invalid characters for deletion, such embodiments being described in more detail below.
[0066] 4 shows a flow chart of a process 400 for prominently presenting invalid characters according to some embodiments of the present invention. The process 400 may be implemented in the terminal device 110. The process 400 of FIG. 4 generally illustrates the display of invalid characters determined based on automatic identification and manual user selection.
[0067] In block 410, the terminal device 110 presents text corresponding to the audio 114, such as the text shown in FIG. 3A. In block 420, the terminal device 110 obtains an invalid character identification result for the text. As previously described, the terminal device 110 can perform the invalid character identification locally or receive the invalid character identification result directly from a remote device. The invalid character identification result can be a set of invalid characters identified in the text or can indicate that no invalid characters were identified in the text.
[0068] In block 430, the terminal device 110 detects whether to enter the invalid character locating mode. If it is not detected that the invalid character locating mode is entered, the terminal device 110 can continue to wait. In response to detecting that the invalid character locating mode is entered, such as when the user selects a corresponding function presented on the edit page 300 in FIG. 3B, in block 440, the terminal device 110 determines whether there is an invalid character that is automatically identified based on the invalid character identification result.
[0069] If automatically identified invalid characters are present, the terminal device 110 prominently presents the automatically identified invalid character set in block 450. The automatically identified invalid character set can be displayed in the text on the edit page 300, as shown in FIG.
[0070] If it is determined in block 440 that the character identification result indicates that there are no automatically identified invalid characters, then no characters will be automatically prominently presented after entering the invalid character location mode. In such a case, the process 400 proceeds to box 460, where the terminal device 110 detects a user selection of invalid characters in the invalid character location mode. For example, it allows the user to select a set of characters from the presented text as invalid characters. In other words, in the invalid character location mode, the invalid characters that are prominently presented include invalid characters selected and determined by the user.
[0071] In some embodiments, after prominently presenting the set of automatically identified invalid characters in block 450, process 400 may proceed to box 460, where terminal device 110 continues to detect user selections for invalid characters in the invalid character location mode. In this case, the user may determine that one or more of the automatically identified invalid characters are not to be deleted and therefore not selected as target invalid characters. Additionally or alternatively, in such a case, the user may also select one or more other characters as invalid characters.
[0072] In block 470, the terminal device 110 determines the prominence for the invalid characters based on the user selection. Depending on the specific selection of the user, the particular invalid character may not be prominenced and a particular different invalid character may be selected to be prominenced.
[0073] In process 400, boxes 460 and 470 may be executed repeatedly until a delete confirmation instruction for the target void character is received, in response to which the currently still selected or prominently presented void character may be determined to be the target void character to be deleted.
[0074] In the following, examples of user selection and conspicuous presentation of invalid characters on an edit page are described in detail in conjunction with FIG. 5 and FIGS. 6A-6F.
[0075] 5 shows a flowchart of a process 500 for prominently presenting invalid characters according to some alternative embodiments of the present invention. Process 500 may be implemented in terminal device 110. Process 500 may be considered an example embodiment of boxes 460 and 470 of step 400. In process 500, it is first assumed that one or more invalid characters have already been prominently presented. The currently prominently presented invalid characters may include one or more invalid characters that have been automatically identified and / or one or more invalid characters selected and determined by a user.
[0076] In block 510, the terminal device 110 determines whether or not it has received a deselection instruction for one or more invalid characters. The deselection instruction is determined based on a user selection. For example, for one or more invalid characters that are prominently displayed, the user can deselect a specific invalid character or a specific number of invalid characters, so that these characters are no longer considered as invalid characters. As shown in FIG. 6A, invalid characters 360-1, 360-2, and 360-3 are prominently displayed on the edit page 300. When the user clicks on the invalid character 360-2 "that", the terminal device 110 receives a deselection instruction for the invalid character.
[0077] In response to receiving the deselection instruction, the terminal device 110 ceases or downgrades the prominence of the deselected invalid character(s) at block 520. In some embodiments, in response to the deselection instruction, the terminal device 110 also removes the deselected invalid character(s) from the set of invalid characters, meaning that these characters are no longer considered invalid characters.
[0078] In one embodiment, the terminal device 110 can prevent one or more deselected invalid characters from being prominently presented, and the presentation of these invalid characters will be the same as the presentation of other characters in the text. Figure 6B shows an example of the pause vs. prominent presentation of a deselected invalid character. Specifically, in Figure 6A, after receiving a deselection instruction for the invalid character 360-2, the character "ano" is no longer prominently presented, as shown in Figure 6B.
[0079] In another embodiment, the terminal device 110 downgrades the conspicuous presentation of the deselected invalid characters, so that the conspicuousness of the deselected invalid characters is lower than that of other invalid characters that are not deselected. In some examples, the deselected invalid characters are still presented conspicuously compared to other characters in the text, to indicate to the user that these characters have been determined as invalid characters (e.g., automatically identified invalid characters). The downgrading of the conspicuous presentation may include canceling some conspicuous presentation aspects (when the invalid characters are presented conspicuously in various aspects), presenting the invalid characters conspicuously according to another aspect (where the conspicuousness of the aspect is lower, for example, by reducing the saturation of the background pattern), and downgrading any other conspicuous presentation.
[0080] 6C illustrates a downgrading of prominence for a deselected invalid character. After receiving a deselect instruction for the invalid character 360-2 in FIG. 6A, the deletion line for the invalid character 360-2 is cancelled, but the color background pattern still remains, as shown in FIG. 6C.
[0081] By providing some degree of prominence for deselected invalid characters, it becomes easier for the user to find those invalid characters again if he or she makes a mistake or the like.
[0082] As discussed above, the number of invalid characters presented may change when one or more invalid characters are deselected, such as upon receiving a deselection indication for one or more invalid characters. For example, in the examples of Figures 6B and 6C, after invalid character 360-2 is deselected, character indication 662 may be presented on edit page 300 to indicate the updated number of invalid characters. The number of invalid characters (e.g., the number "2") may also be displayed on "Confirm Delete" option 672.
[0083] 5, the terminal device 110 determines whether a selection restore instruction is received for one or more invalid characters. If a selection restore instruction is received, in block 540, the terminal device 110 restores prominence for the invalid characters from the suspended or demoted state.
[0084] As shown in Fig. 6C, the user can conveniently position and select the invalid character 360-2 again as necessary. As shown in Fig. 6D, when receiving a user's reselection of the invalid character 360-2, the conspicuous presentation of the invalid character 360-2 is restored to the same degree as the conspicuous presentation of the other invalid characters. Of course, in such a case, the invalid character may be indicated as a restored invalid character again in a different conspicuous presentation manner. This is not limited here.
[0085] Also, after invalid character 360-2 is reselected, because the number of invalid characters has now increased, in the example of Figure 6D, edit page 300 may further provide an updated number of characters indication 664 for the invalid characters. Also, the number of invalid characters (e.g., the number "3") may be displayed on "Confirm Delete" option 674.
[0086] In some embodiments, in various cases, the terminal device 110 determines whether another character has been detected as an invalid character in block 550. For example, if a deselection instruction is detected in block 510, or if a selection restoration instruction is not detected in block 530, or after restoring the prominent presentation of the invalid character, the terminal device 110 may continue to determine whether another character has been detected as an invalid character. Although the steps of each box in the flowchart shown in FIG. 5 are described in sequence, it should be understood that these steps may be performed in a different order or in parallel. For example, the steps of boxes 510, 530, and 550 may be performed in parallel.
[0087] In the process of being in the invalid character locating mode, the terminal device 110 determines whether another one or more characters in the text are selected as invalid characters based on the user selection, for example, allowing the user to select one or more characters that are not prominently presented in the text as invalid characters.
[0088] In response to detecting that one or more additional characters have been selected as invalid characters, at block 560, the terminal device 110 prominently presents the selected one or more additional invalid characters.
[0089] As shown in FIG. 6E, the user selects a character that is not prominently presented, such as the character 660-1 “あ” in the text area 310. The terminal device 110 detects such a user selection and determines to select the character as an invalid character. As shown in FIG. 6F, the terminal device 110 prominently presents the character 660-1 on the edit page 300. At this time, since the number of invalid characters increases from 2 to 3, a character indication 666 of the updated number of invalid characters may be presented on the edit page 300. Additionally, the number of invalid characters (e.g., the number “3”) may also be displayed on the “Confirm Delete” option 676.
[0090] In some implementations, when the audio 114 is in a non-playing state, a user selection such as deselection, selection restoration, and / or selection of a separate override character may be detected. In some examples, as shown in Figures 3D and 6A-6F, when the prominent presentation of the override character is initiated, an instruction 364 may be presented in the edit page 300 to indicate that the override character may be kept or removed by clicking on the highlighted portion of the audio in the paused state. As shown in Figures 6A-6F, the edit page 300 presents a playback pause indicator 663 to indicate that the audio 114 is in a playback paused state.
[0091] If it is not detected that another one or more characters have been selected as an invalid character in block 550, the terminal device 110 may determine that there is currently no need to prominently present another invalid character. In the process of being in the invalid character location mode, the terminal device 110 may perform the detection of boxes 510, 530, and 550 multiple times in succession.
[0092] Regardless of whether the user further edits the invalid characters, after detecting the delete confirmation instruction, the currently prominently presented invalid characters can be determined as target invalid characters to be deleted, and audio portions corresponding to these target invalid characters can be deleted from the audio 114 based on these target invalid characters. As shown in FIG. 7A, when the user selects the "delete confirm" option 372, the currently prominently presented invalid characters 360-1, 360-3, and 660-1 are confirmed as target invalid characters. Thus, audio portions corresponding to these target invalid characters are deleted from the audio 114, and these target invalid characters are also deleted from the text. As shown in FIG. 7B, the updated text can be presented in the text area 310 of the edit page 300, where the target invalid characters 360-1, 360-3, and 660-1 are no longer presented. Also, voiceprint representation information 720 and duration information 722 related to the updated audio can be further presented.
[0093] In some embodiments, when an invalid character is selected by a user and determined to be deleted, text corresponding to the invalid character and audio 114 based on the user selection may be provided to train the invalid character identification model. For example, text corresponding to the character "あ" selected by the user in FIG. 6E and FIG. 6F and audio 114 may be provided to train the character identification model. The character identification model may be a model used by the terminal device 110 or the remote computing device 112 to automatically identify invalid characters. The provided text corresponding to the invalid character and audio 114 may enrich and expand the training dataset of the invalid character identification model, so that the invalid character identification model may evolve to have stronger identification ability.
[0094] In some embodiments, training of the invalid character identification model may be resumed after collecting a sufficient amount of additional training data. In some embodiments, training of the invalid character model may be performed on the terminal device 110, the computing device 112, or other model training device. Embodiments of the invention are not limited in this respect.
[0095] 8 shows a block diagram of a schematic configuration of an apparatus 800 for audio editing according to a specific embodiment of the present invention. The apparatus 800 can be implemented as or included in the terminal device 18. Each module / component in the apparatus 800 can be implemented by hardware, software, firmware, or any combination thereof.
[0096] As shown in the figure, the apparatus 800 comprises a conspicuous presentation module 810 for conspicuously presenting one or more invalid characters included in a text corresponding to the audio in a predefined mode for the audio, the apparatus 800 further comprises an indication detection module 820 for detecting a deletion confirmation indication for at least one target invalid character of the one or more invalid characters, and an audio deletion module 830 for deleting at least one audio portion corresponding to the at least one target invalid character from the audio in response to the deletion confirmation indication being detected.
[0097] In some embodiments, the apparatus 800 further comprises an invalid character identification module for identifying a first invalid character from the text.
[0098] In some embodiments, the apparatus 800 further comprises an invalid character determination module for determining a second invalid character in the text based on the user input.
[0099] In some embodiments, the apparatus 800 further comprises a data providing module for providing a second invalid character and text to train the invalid character identification model, where the invalid character identification model is trained to identify invalid characters from the input text.
[0100] In some embodiments, the indication detection module includes an invalid character removal module for removing the third invalid character from the one or more invalid characters in response to receiving a deselection indication for a third invalid character among the one or more invalid characters.
[0101] In some embodiments, the apparatus 800 further comprises a deactivating or downgrading module for deactivating or downgrading prominent presentation for the fourth invalid character in response to receiving a deselection instruction for the fourth invalid character of the one or more invalid characters.
[0102] In some embodiments, the apparatus 800 further comprises a number submission module for submitting a first number of the one or more invalid characters.
[0103] In some embodiments, the apparatus 800 further comprises a number determination module for determining a second number of non-deselected invalid characters among the one or more invalid characters in response to receiving a deselection instruction for at least one invalid character among the one or more invalid characters, and a number modification module for modifying the presented first number to the second number.
[0104] In some embodiments, the apparatus 800 further comprises a text determination module for, in response to detecting a deletion confirmation indication for the at least one target invalid character, deleting the at least one target invalid character from the text to obtain an updated text, and a text presentation module for presenting the updated text.
[0105] In some embodiments, the apparatus 800 further comprises an information module for presenting related information of the audio after removing at least one audio portion, the related information including at least one of a duration and a sound wave representation.
[0106] 9 illustrates a block diagram of a computing device 900 capable of implementing one or more embodiments of the present invention. It should be understood that the computing device 900 illustrated in FIG. 9 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The computing device 900 illustrated in FIG. 9 may be used to implement the terminal device 110 of FIG. 1.
[0107] As shown in Fig. 9, the computing device 900 is in the form of a general purpose electronic device. The components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be a real or virtual processor and may perform various processes based on programs stored in the memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel, thereby improving the parallel processing capabilities of the computing device 900.
[0108] Computing device 900 typically includes a number of computer storage media. Such media may be any obtainable media accessible by computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 may be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 may be removable or non-removable media and may include machine-readable media, such as a flash memory drive, a magnetic disk, or any other media, and may be used to store information and / or data (e.g., training data for training) and may be accessible within computing device 900.
[0109] The computing device 900 may further include other removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected by one or more data media interfaces to a path (not shown). The memory 920 may include a computer program product 925 having one or more program modules, which are configured to perform various methods or operations of various embodiments of the invention.
[0110] The communication unit 940 implements communication with other computing devices over a communication medium. Additionally, the functionality of the components of the computing device 900 may be implemented as a single computing cluster or multiple computing machines that can communicate over a communication connection. Thus, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0111] The input device 950 may be one or more input devices such as a mouse, a keyboard, a trackball, etc. The output device 960 may be one or more output devices such as a display, a speaker, a printer, etc. The computing device 900 may further communicate, as necessary, with one or more external devices (not shown) such as a storage device, a display device, etc. via the communication unit 940, one or more devices that allow a user to interact with the computing device 900, or the computing device 900 communicates with any device (net card, modem, etc.) that communicates with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0112] According to an exemplary implementation of the present invention, a computer-readable storage medium having one or more computer instructions stored thereon is provided, the one or more computer instructions being executed by a processor to implement the above-mentioned method. According to an exemplary implementation of the present invention, a computer program product is further provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions executed by a processor to implement the above-mentioned method.
[0113] Aspects of the present invention have been described herein with reference to flowchart and / or block diagrams of methods, apparatus (systems) and computer program products implemented by the present invention. It will be understood that each box in the flowchart and / or block diagrams, and combinations of boxes in the flowchart and / or block diagrams, can be implemented by computer readable program instructions.
[0114] These computer readable program instructions may be provided to a processing unit of a general purpose computer, special purpose computer, or other programmable data processing apparatus to generate a machine such that, when the instructions are executed by a processing unit of the computer or other programmable data processing apparatus, they generate an apparatus for implementing the functions / operations specified in one or more boxes in the flowcharts and / or block diagrams. These computer readable program instructions may be stored on a computer readable storage medium such that the instructions cause the computer, programmable data processing apparatus, and / or other device to operate in a particular manner such that the computer readable medium on which the instructions are stored constitutes an article of manufacture including instructions that implement each aspect of the functions / operations specified in one or more boxes in the flowcharts and / or block diagrams.
[0115] Loading the computer-readable program instructions into a computer, other programmable data processing apparatus, or other device causes the computer, other programmable data processing apparatus, or other device to perform a series of operational steps to produce a computer-implemented process, such that the instructions executing on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes in the flowcharts and / or block diagrams.
[0116] The flowcharts and block diagrams in the figures illustrate possible architectures, functions, and operations of the systems, methods, and computer program products that may be implemented according to the present invention. In this regard, each box in the flowcharts or block diagrams may represent a module, program segment, or part of instructions, which includes one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions depicted in the boxes may occur in a different order than depicted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or may be executed in a reverse order depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a special purpose hardware-based system that performs the specified functions or operations, or by a combination of special purpose hardware and computer instructions.
[0117] Although each implementation of the present invention has been described above, the above description is illustrative, not exhaustive, and is not limited to each disclosed implementation. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of each described implementation. The selection of terms used in this specification is intended to best interpret the principles, practical applications, or improvements to technology in the marketplace of each implementation, or to enable those skilled in the art to understand each implementation disclosed in this specification.
Claims
1. prominently presenting, in a predefined mode for the audio, one or more null characters included in a text corresponding to the audio; detecting a deletion confirmation indication for at least one target invalid character among the one or more invalid characters; in response to detecting the deletion confirmation indication, deleting from the audio at least one audio portion corresponding to the at least one target invalid character. How to edit audio.
2. further comprising identifying a first invalid character from the text.
2. The method of audio editing according to claim 1.
3. determining a second invalid character in the text based on user input; 2. The method of audio editing according to claim 1.
4. providing the second invalid character and the text to train an invalid character identification model, the invalid character identification model being trained to identify invalid characters from the input text.
4. The method of audio editing according to claim 3.
5. Detecting the deletion confirmation instruction includes: in response to receiving a deselection instruction for a third invalid character among the one or more invalid characters, removing the third invalid character from the one or more invalid characters.
2. The method of audio editing according to claim 1.
6. and in response to receiving a deselection instruction for a fourth void character of the one or more void characters, ceasing or downgrading prominence for the fourth void character.
2. The method of audio editing according to claim 1.
7. providing a first number of the one or more null characters.
2. The method of audio editing according to claim 1.
8. in response to receiving a deselection instruction for at least one invalid character of the one or more invalid characters, determining a second number of the invalid characters that are not deselected; and modifying the submitted first number to the second number.
8. The method of audio editing according to claim 7.
9. in response to detecting the deletion confirmation indication for the at least one target void character, deleting the at least one target void character from the text to obtain updated text; and presenting the updated text.
2. The method of audio editing according to claim 1.
10. presenting related information of the audio after removing the at least one audio portion, the related information including at least one of a duration and a sound wave representation.
2. The method of audio editing according to claim 1.
11. a prominence module for prominently presenting, in a predefined mode for the audio, one or more invalid characters included in a text corresponding to the audio; an indication detection module for detecting a deletion confirmation indication for at least one target invalid character among the one or more invalid characters; an audio deletion module for deleting at least one audio portion corresponding to the at least one target invalid character from the audio in response to the deletion confirmation indication being detected. Equipment for audio editing.
12. and an invalid character identification module for identifying a first invalid character from the text. Apparatus for audio editing according to claim 11.
13. and an invalid character determination module for determining a second invalid character in the text based on a user input. Apparatus for audio editing according to claim 11.
14. a data providing module for providing the second invalid character and the text to train an invalid character identification model, the invalid character identification model being trained to identify invalid characters from input text; Apparatus for audio editing according to claim 13.
15. The instruction detection module includes: an invalid character removal module for removing a third invalid character from the one or more invalid characters in response to receiving a deselection instruction for the third invalid character among the one or more invalid characters; Apparatus for audio editing according to claim 11.
16. and a prominence deactivation or deactivation module for deactivating or deactivating prominence for a fourth void character in response to receiving a deselection instruction for the fourth void character of the one or more void characters. Apparatus for audio editing according to claim 11.
17. and a number generating module for generating a first number of the one or more invalid characters. Apparatus for audio editing according to claim 11.
18. a text determination module for, in response to detecting the deletion confirmation indication for the at least one target void character, deleting the at least one target void character from the text to obtain updated text; a text presentation module for presenting the updated text. Apparatus for audio editing according to claim 11.
19. 1. An electronic device comprising: At least one processing unit; and at least one memory coupled to said at least one processing unit and adapted to store instructions to be executed by said at least one processing unit, said instructions, when executed by said at least one processing unit, causing said electronic device to execute the method according to any one of claims 1 to 9. The electronic device.
20. A computer program is stored, the computer program being executed by a processor to implement the method according to any one of claims 1 to 9. A computer-readable storage medium.
Citation Information
Patent Citations
Video processing method, device, electronic equipment and storage medium
CN113613068A
Voice reproduction device, voice reproduction method, and program
JP2017111339A
Audio data processing apparatus, audio data processing method, program, and recording medium
JP2018195896A
Voice recognition device and system
JP2019095644A
Method and User Interface for Creating an Audio Recording Using a Document Paradigm
US20090082887A1