An audio editing method, system, device, and storage medium

By collaborating with middleware and audio editing models, the system identifies audio content and provides editing suggestions, solving problems such as slips of the tongue and style mismatch in audio editing, thus achieving efficient and intelligent audio editing.

CN119207367BActive Publication Date: 2026-01-02BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411246707.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-01-02
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

Existing audio editing systems cannot effectively handle issues such as slips of the tongue, missing content, and mismatched voice styles that occur during voice recording, leading to the need for re-recording and wasting time.

Method used

Through the collaboration of middleware and audio editing models, audio content is identified, editing suggestions are generated, and editing materials are provided using content and template libraries to achieve automatic editing of audio content and style adjustments.

Benefits of technology

It improves the efficiency and accuracy of audio editing, provides a convenient and intelligent audio editing experience, reduces manual intervention, and enhances the integrity and style adaptability of audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207367B_ABST
    Figure CN119207367B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio editing method, system, device and storage medium, relating to the technical field of computers, and particularly to the technical fields of data processing, large language models, audio editing and the like. The method comprises: inputting, by a middleware, to-be-edited audio and editing instructions into an audio editing model; determining, by the audio editing model, corresponding editing suggestions based on the to-be-edited audio and the editing instructions, and feeding back the editing suggestions to the middleware; determining, by the middleware, corresponding editing materials according to the editing suggestions, constructing prompt words according to the editing suggestions and the editing materials, and inputting the prompt words and the editing materials into the audio editing model; and editing, by the audio editing model, the to-be-edited audio based on the prompt words and the editing materials to obtain edited audio. The present disclosure can improve the efficiency and accuracy of audio editing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer technology, and particularly relates to the technical fields of data processing, large language model, audio editing, etc. BACKGROUND

[0002] With the wide application of intelligent terminals, more and more users record oral broadcasts, that is, the user orally broadcasts a content, and the intelligent terminal records and saves the oral broadcast content. In oral broadcast recording, there are often errors, missing content, or voice style and audio content do not match, etc. In this case, it is necessary to re-record, resulting in a large amount of time waste. How to realize automatic editing or modification of audio content is a technical problem to be solved. SUMMARY

[0003] The present disclosure provides an audio editing method, system, device and storage medium.

[0004] According to an aspect of the present disclosure, an audio editing method is provided, applied to an audio editing system including a middleware and an audio editing model, the method comprising:

[0005] The middleware inputs the audio to be edited and the editing instruction into the audio editing model;

[0006] The audio editing model determines the corresponding editing suggestion based on the audio to be edited and the editing instruction, and feeds back the editing suggestion to the middleware;

[0007] The middleware determines the corresponding editing material according to the editing suggestion, constructs the prompt word according to the editing suggestion and the editing material, and inputs the prompt word and the editing material into the audio editing model;

[0008] The audio editing model edits the audio to be edited based on the prompt word and the editing material to obtain the edited audio.

[0009] An audio editing system includes a middleware and an audio editing model; wherein,

[0010] The middleware is configured to input the audio to be edited and the editing instruction into the audio editing model, receive the editing suggestion returned by the audio editing model, determine the corresponding editing material according to the editing suggestion, construct the prompt word according to the editing suggestion and the editing material, and input the prompt word and the editing material into the audio editing model;

[0011] The audio editing model is configured to determine the corresponding editing suggestion based on the audio to be edited and the editing instruction, and feed back the editing suggestion to the middleware; and is further configured to receive the prompt word and the editing material, and edit the audio to be edited based on the prompt word and the editing material to obtain the edited audio.

[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein

[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method according to any of the embodiments of the present disclosure.

[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.

[0018] The present disclosure realizes the deep integration of user editing intention and audio processing technology through the close cooperation of middleware and audio editing model. This method not only improves the efficiency and accuracy of audio editing work, but also creates a convenient and intelligent audio creation and editing environment for users, making the audio editing process more smooth and intuitive.

[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:

[0021] Figure 1 is an implementation flowchart of an audio editing method 100 according to an embodiment of the present disclosure;

[0022] Figure 2 is a schematic diagram of the content delivery method in the audio editing method proposed in the embodiments of the present disclosure;

[0023] Figure 3 is a structural schematic diagram of an audio editing system 300 according to an embodiment of the present disclosure;

[0024] Figure 4 shows a schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0025] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, in which various details are set forth to facilitate understanding of the present disclosure. However, it should be apparent to those of ordinary skill in the art that the embodiments described herein can be practiced without such details. In other instances, well-known methods, procedures, components, and networks have not been described in detail so as not to unnecessarily obscure aspects of the present disclosure.

[0026] The "and / or" of the embodiments of the present disclosure means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first", "second", herein mean to refer to a plurality of similar technical terms and to distinguish them, and are not meant to limit the order or to limit to only two, for example, the first feature and the second feature mean to refer to two categories / two features, the first feature can be one or more, and the second feature can also be one or more.

[0027] With the wide application of intelligent terminals, more and more users record oral broadcasts, and the user orally broadcasts a content, and the intelligent terminal records and saves the oral broadcast content. In oral broadcast recording, there are often errors, missing content, or voice style and audio content do not match, etc. In this case, it is necessary to re-record, resulting in a lot of time waste; if there is an audio editor that can help automatically correct audio problems or edit audio content, it will greatly improve efficiency. At the same time, the editor can also be used as a background sound removal tool, background music generation tool, etc.

[0028] The existing audio editing system mainly uses a single expert model to implement a single task; this kind of expert model can implement audio editing in a single field, and cannot cover a wider field; for other fields outside the field it is good at, the audio editing effect of the expert model still has defects, and it is easy to edit audio errors, resulting in loss of original audio quality or content.

[0029] To solve the above problems, the present disclosure provides an audio editing method and system. The audio editing method and system can be applied to electronic devices such as mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, game consoles, e-book readers, multimedia playback devices, and wearable devices. The electronic device receives a to-be-edited audio and an editing instruction input by a user, and automatically edits the to-be-edited audio according to the editing instruction. Alternatively, the audio editing method and system can also be applied to a cloud server, which receives a to-be-edited audio and an editing instruction sent by a user through a terminal device, automatically edits the to-be-edited audio according to the editing instruction, and returns the edited audio to the terminal device.

[0030] Figure 1 FIG. 1 is a flowchart of an implementation of an audio editing method 100 according to an embodiment of the present disclosure, which includes the following steps:

[0031] S110, the middleware inputs the to-be-edited audio and the editing instruction into the audio editing model;

[0032] S120, the audio editing model determines the corresponding editing suggestion based on the to-be-edited audio and the editing instruction, and feeds back the editing suggestion to the middleware;

[0033] S130, the middleware determines the corresponding editing material according to the editing suggestion, constructs a prompt word according to the editing suggestion and the editing material, and inputs the prompt word and the editing material into the audio editing model;

[0034] S140, the audio editing model edits the to-be-edited audio based on the prompt word and the editing material to obtain the edited audio.

[0035] In the audio editing method provided by the present disclosure, the middleware and the audio editing model work closely together to achieve the deep integration of user editing intention and audio processing technology. This method not only improves the efficiency and accuracy of audio editing, but also provides users with a more convenient and intelligent audio editing experience.

[0036] In some embodiments, the audio editing model identifies the to-be-edited audio to determine the corresponding text of the to-be-edited audio; and the audio editing model determines the corresponding editing suggestion based on the audio text and the audio editing instruction.

[0037] An audio editing model receives an audio file to be edited, which can contain various sound elements such as the voice of the main speaker, the voices of other speakers, background music, and environmental noise, etc. In order to accurately extract the voice of the main speaker, the audio editing model will preprocess the audio to be edited, including filtering out background noise using a filter, adjusting the overall volume level of the audio through gain, and identifying human voice in the audio using a voice activity detection (VAD) method.

[0038] The core of the VAD method is to analyze the subtle features in the audio in depth to accurately distinguish between human voice segments and non-active segments that do not contain human voice (such as long periods of silence, environmental noise, or other non-speech background sounds, etc.). When building an audio editing model, a voice detection module can be included in the audio editing model. Based on the VAD method, the voice detection module can automatically identify the human voice part in the audio, so that the audio editing model can focus more on the processing and utilization of human voice.

[0039] To achieve this goal, the voice detection module can use machine learning or deep learning techniques. The voice detection module can be pre-trained. During training, the voice detection module, as part of the audio editing model, can be trained together with other parts of the audio editing model; or the voice detection module can be trained, and after the voice detection module is trained, the audio editing model is trained, and during training, the parameters of the voice detection module are frozen, and only the parameters of other parts of the audio editing model are adjusted.

[0040] During the training of the voice detection module, a diverse audio dataset can first be prepared, which should contain rich human voice samples and various types of non-human voice interference (such as natural environmental noise, mechanical equipment sound, background music, etc.). Then, through feature extraction techniques, multi-dimensional feature vectors that can represent human voice features are extracted from each audio sample, such as time domain features and frequency domain features, etc. These feature vectors are used as input to train the voice detection module, so that the module can effectively distinguish between human voice samples and non-human voice samples. After the human voice samples and non-human voice samples in the audio to be edited are distinguished, the audio editing model performs channel separation on the human voice samples. Specifically, the audio editing model has strong analysis capabilities and can analyze the characteristics of the human voice samples in the audio to be edited in depth based on the pre-stored voice information of the main speaker as a reference. Through this detailed analysis process, the model can accurately extract the unique voice of the main speaker from the complex audio environment, achieving efficient extraction and separation of the voice.

[0041] The voice of the main speaker is processed frame by frame, and each frame of audio waveform is converted into audio features by using linear prediction cepstral coefficients (LPCC) and mel-frequency cepstral coefficients (MFCC) and input into the audio editing model.

[0042] By using the acoustic module and the language module in the audio editing model, each frame of audio information can be converted into text.

[0043] The acoustic module is used to process the audio features of each frame of audio and output phoneme information in the audio, that is, the smallest unit of speech, which carries pronunciation information of the speech. For example, "ma" in Chinese pinyin, in this syllable, "m" and "a" are two different pronunciation actions, so they are two different phonemes. The acoustic module can be obtained through model training. Before training the acoustic module, a large amount of audio data needs to be collected, which covers different speakers, speech rates, tones, etc. The features of the audio data are extracted, and the audio features are correspondingly labeled with phonemes, and the labeling results are input into the model for training.

[0044] Based on the factor information recognized by the acoustic module, the corresponding characters or words are searched in the character library. For Chinese, the corresponding Chinese characters are searched according to the pinyin; for English, the corresponding words are searched according to the phonetic symbols.

[0045] The language module is used to evaluate the association probability of each character or word with the phoneme to determine the final recognized audio text. The language module can be obtained through model training. Before training the language module, a large amount of text data needs to be collected, which covers news reports, literary works, social media posts, etc. The text data is preprocessed, including word segmentation, deduplication, etc. The content in the processed text data is labeled with the corresponding phonemes, and the labeling results are input into the model for training.

[0046] The audio editing model determines the corresponding editing suggestion according to the recognized audio text and the editing instruction. For example, a user records an audio about a historical event, and the recognized audio text content is "September 2, 1945, is the day when the war ended", and the user's editing instruction is "help me correct the wrong word in this sentence". According to the audio text and the editing instruction, the audio editing model outputs the editing suggestion as "the wrong word in this sentence is 'zhang', which can be modified to 'zhan'".

[0047] Or, the user records an audio of self-introduction, the audio text content after identification is "Hello everyone, I am XX", and the editing instruction provided by the user is "Help me make the sentence smooth". The audio editing model outputs the editing suggestion as "The audio can be modified as 'Hello everyone, I am XX' and 'Hello, I am XX'". The user can select the appropriate editing suggestion according to the actual needs.

[0048] In the embodiments of the present application, the audio content can be accurately converted into text, providing a reliable basis for subsequent editing work. At the same time, the audio editing model can intelligently analyze and give reasonable editing suggestions, reducing the possibility of human judgment errors.

[0049] In some embodiments, the editing instruction includes at least one of a content editing instruction and a style editing instruction.

[0050] In some embodiments, the editing suggestion includes at least one of an editing suggestion for the audio content and an editing suggestion for the audio style.

[0051] In the embodiments of the present application, the audio editing model can understand the content editing instruction and the style editing instruction of the user. The content editing instruction includes a series of instructions for modifying or supplementing the content said by the user.

[0052] The style editing instruction refers to an instruction for changing the audio style, wherein the audio style includes but is not limited to music style (such as classical, rock, jazz, electronic music, etc.), emotional expression style (such as happy, sad, calm, etc.), era style (such as retro, modern, future, etc.), and specific scene atmosphere (such as movie theater sound effect, outdoor natural sound, etc.). The audio editing model can determine the audio style that best matches the audio to be edited according to the style editing instruction provided by the user, or automatically analyze and match, and then perform subsequent audio editing work.

[0053] The audio editing model can generate corresponding editing suggestions according to the content editing instruction and the style editing instruction of the user. By using the above method, the editing of the audio content and its style can be realized.

[0054] In some embodiments, the audio editing model modifies the text based on the text and the content editing instruction; and generates an editing suggestion for the audio content based on the modified text and the time stamp of each part in the text; and / or,

[0055] The audio editing model determines the field corresponding to the text; and generates an editing suggestion for the audio style according to the field and the style editing instruction.

[0056] For editing suggestions of audio content, for example, a user records an audio segment of "I went to the market yesterday and bought some pots", and the user provides an instruction of "correct the misspelling in the audio". The audio editing model identifies the audio text, finds the misspelling from the text, and makes a modification in the text, and outputs the editing suggestion as "timestamp: 1703213598, the misspelling 'chao' can be modified to 'chao'; timestamp 1703213612, the misspelling 'guo' can be modified to 'guo'".

[0057] For editing suggestions of audio style, for example, a user records a short course audio, and the audio content can be "Today we learn about Newton's first law in physics", and the user provides an instruction of "match the audio with a suitable style". The audio editing model identifies the audio text, determines from the audio text that the audio is used in the field of education, and according to the field, the audio editing model outputs the style editing suggestion for the audio as "the audio is a teaching audio, and the audio style can be set to 'quiet'".

[0058] By using the above method, the accuracy and editing efficiency of audio editing can be improved. At the same time, the method supports field-specific and stylized audio editing, and enhances user experience.

[0059] In some embodiments, the editing material includes at least one of content that needs to be supplemented and an audio template;

[0060] The middleware determines the corresponding editing material according to the editing suggestion, including:

[0061] The middleware finds the content that needs to be supplemented from a pre-set content library according to the editing suggestion for the audio content; and / or,

[0062] The middleware finds the corresponding audio template from a pre-set audio template library according to the editing suggestion for the audio style.

[0063] By using the above method, the required content or template can be found and supplemented from the pre-set content library or audio template library, greatly reducing the user's manual search and screening time, thereby improving the editing efficiency of the audio content.

[0064] In some embodiments, the prompt word contains at least one of a first indication, a second indication and a third indication, wherein the first indication is used to indicate that the audio editing model modifies at least one of the style and background sound of the audio to be edited according to the audio template; the second indication is used to indicate that the audio editing model expands the audio to be edited based on the content that needs to be supplemented; and the third indication is used to indicate that the audio editing model corrects errors in the audio to be edited.

[0065] By using the above method, the accuracy of audio editing can be improved, the integrity of audio content can be enhanced, and the audio automated editing process is promoted, and the need for manual intervention is reduced.

[0066] The embodiments of the present disclosure also provide an audio editing system, which comprises middleware and an audio editing model; wherein the audio editing model can be a general large language model and is applied to multiple fields. Figure 2 FIG. 1 is a schematic diagram of a content delivery method in an audio editing method according to an embodiment of the present disclosure.

[0067] S201, input the audio to be edited and the editing instruction into the middleware.

[0068] After the audio to be edited and the editing instruction are input into the middleware, the middleware can preprocess the audio, such as noise reduction, volume adjustment, etc., and convert the audio to be edited and the editing instruction into a format that can be understood by the audio editing model, so that after the audio to be edited and the editing instruction are input into the audio editing model, the corresponding task can be performed.

[0069] S202, input the audio to be edited and the editing instruction into the audio editing model by the middleware.

[0070] S203, the audio editing model identifies the audio to be edited and generates an audio text.

[0071] According to the specific content of the audio text and the editing instruction provided by the user, each aspect of the audio is analyzed in depth, and a series of targeted editing suggestions are generated based on the analysis results and fed back to the middleware. These suggestions not only cover how to improve the audio content, but also include how to adjust the audio style to meet the specific emotional or atmosphere requirements.

[0072] S204, the middleware queries related content from the content library or the audio template library according to the feedback editing suggestions.

[0073] Specifically, if the editing suggestions contain modifications or expansions of the audio content, the middleware queries related content from the content library based on the modified or expanded content; if the editing suggestions contain modifications of the audio template, the middleware queries related templates from the audio template library based on the modification suggestions of the audio template.

[0074] S205, the content library or the audio template library feeds back the related materials queried by the middleware to the middleware.

[0075] S206, the middleware inputs the prompt words and the editing materials into the audio editing model.

[0076] S207, the audio editing model processes the audio to be edited according to the received prompt words and editing materials, and obtains the edited audio.

[0077] Based on Figure 2 In the specific mode of content transfer in the audio editing method shown, this process can be deeply understood and practiced through the following examples:

[0078] Example 1: A user is recording an explanatory audio about a historical event with the content "During a period of 'zhang' from 1939 to 1945, at that time, well, people experienced a lot of difficulties, and then, well, they started looking for solutions". From the above audio content, it can be seen that the user wrongly corresponded the time with the historical event, there are typos in the audio, and the sentence expression is not smooth. The user can input the audio editing instruction "Correct the typos in the audio and make this sentence smooth" into the audio editing model.

[0079] After receiving the instruction "Correct the typos in the audio and make this sentence smooth" and identifying the audio content, the audio editing model can give suggestions for the edited content of the audio as "At timestamp 1703211378, change the typo 'zhang' to 'zhan'; at timestamp 1703211250, modify the historical event corresponding to 1939 - 1945".

[0080] After receiving the editing suggestions provided by the audio editing model as "At timestamp 1703211378, change the typo 'zhang' to 'zhan'; at timestamp 1703211250, modify the historical event corresponding to 1939 - 1945", the middleware queries the historical event that fits the time period of "1939 - 1945" from the content library, and the editing material feedback by the content library to the middleware is "During the period from 1939 to 1945, the world experienced World War II".

[0081] According to the actual situation of this example, the middleware sends the queried editing material and the third instruction to the audio editing model, and the result after being processed by the audio editing model is "During World War II from 1939 to 1945, people encountered unprecedented hardships, and then, they began to actively seek various solutions to cope with the difficulties".

[0082] Example 2: A user is recording audio content "We came to AA County today", and the user can provide the instruction "At timestamp 170321256, supplement the location information of AA County" to the audio editing model.

[0083] After receiving the instruction "At timestamp 170321256, supplement the information of AA County" and identifying the audio content, the audio editing model can give suggestions for the edited content of the audio as "Relevant geographical information can be supplemented before 'AA County'".

[0084] The middleware receives the editing suggestion provided by the audio editing model as "relevant geographic information can be supplemented before 'AA County'". The middleware queries the geographic information about "AA City" from the content library. The content library feeds back the editing material "AA County, located in FF City, YY Province" to the middleware.

[0085] According to the actual situation of the present example, the middleware sends the queried editing material and the second indication to the audio editing model. The result processed by the audio editing model is "we came to AA County in FF City, YY Province today".

[0086] In Example Three, the user records an audio congratulating the opening of a store. The user can input the style editing instruction "change the style of this audio to a happy style" to the audio editing model.

[0087] The audio editing model receives the instruction "change the style of this audio to a happy style" and identifies the audio content. The audio editing model can then provide the suggestion "happy style is suitable for the audio".

[0088] The middleware receives the editing suggestion provided by the audio editing model as "happy style is suitable for the audio". The middleware queries the audio template of the happy style from the audio template library. Since the audio template library feeds back more than one audio template to the middleware, the middleware can evaluate the queried audio templates according to the audio content. The evaluation process can be based on multiple factors, such as the similarity between the audio template and the original audio content, the applicability of the audio template, the style matching degree, and the user preference, etc. Through comprehensive analysis of these factors, the middleware can intelligently select the audio template that is most suitable for the current audio content.

[0089] According to the actual situation of the present example, the middleware sends the queried editing material and the first indication to the audio editing model. The audio editing model fuses the audio template with the audio to be edited to obtain the edited audio.

[0090] After the above steps, the editing material may not be queried. In the case of query failure, these situations can be recorded, including the keywords, time, and other information of the query, and these information is fed back to the maintenance personnel to supplement the data missing situation.

[0091] Under extreme conditions (such as high noise environment), the audio editing model may not be able to identify the audio recorded by the user, so the audio editing model needs to be fine-tuned. The model fine-tuning process includes:

[0092] Collect audio data recorded by the user under extreme conditions. These data cover different extreme conditions. The collected audio data is carefully annotated, and the time period of the main speaker, the speaking content of the main speaker, and other information are clearly marked;

[0093] According to the annotation information, a series of prompts are designed, which can guide the audio editing model to focus on the voice channel characteristics of the main speaker and prompt the audio editing model to identify the speech content of the main speaker. The prompts can include descriptions of audio characteristics (such as "extract clear human voice in a noisy environment"), specific indications of target voice channels (such as "focus on the voice channel of the main speaker"), and expected output formats (such as "separate and clarify the audio of the main speaker").

[0094] The intermediate obtains the editing suggestion of the audio editing model for the audio data under extreme conditions, and inputs the audio data under extreme conditions, the designed prompt, and the information provided by the intermediate into the audio editing model. The model continuously attempts to separate the voice of the main speaker and the speech content of the main speaker from the audio data, and continuously optimizes its internal parameters through detailed comparison with the annotation information to achieve more accurate editing effect. The present disclosure also provides an audio editing system, Figure 3 is a structural schematic diagram of an audio editing system 300 according to an embodiment of the present disclosure, which comprises an intermediate 310 and an audio editing model 320: wherein,

[0095] The intermediate 310 is configured to input the audio to be edited and the editing instruction into the audio editing model, receive the editing suggestion returned by the audio editing model, determine the corresponding editing material according to the editing suggestion, construct a prompt word according to the editing suggestion and the editing material, and input the prompt word and the editing material into the audio editing model.

[0096] The audio editing model 320 is configured to determine the corresponding editing suggestion based on the audio to be edited and the editing instruction, and feed back the editing suggestion to the intermediate. The audio editing model 320 is also configured to receive the prompt word and the editing material, and edit the audio to be edited based on the prompt word and the editing material to obtain the edited audio.

[0097] In some embodiments, the audio editing model 320 is configured to,

[0098] identify the audio to be edited to determine the text corresponding to the audio to be edited;

[0099] determine the corresponding editing suggestion based on the text and the editing instruction.

[0100] In some embodiments, the editing instruction comprises at least one of a content editing instruction and a style editing instruction.

[0101] In some embodiments, the editing suggestion comprises at least one of an editing suggestion for audio content and an editing suggestion for audio style.

[0102] In some embodiments, the audio editing model 320 is configured to,

[0103] based on the text and the content editing instruction, modifying the text; and generating an editing suggestion for the audio content based on the modified text and the time stamp of each part in the text; and / or,

[0104] determining a field corresponding to the text; and generating an editing suggestion for the audio style according to the field and the style editing instruction.

[0105] In some embodiments, the editing material includes at least one of content that needs to be supplemented and an audio template;

[0106] The middleware 310 is configured to find the content that needs to be supplemented from a pre-set content library according to the editing suggestion for the audio content; and / or,

[0107] find a corresponding audio template from a pre-set audio template library according to the editing suggestion for the audio style.

[0108] In some embodiments, the prompt word contains at least one of a first indication, a second indication and a third indication:

[0109] The first indication is used to indicate that the audio editing model modifies at least one of the style and the background sound of the audio to be edited according to the audio template;

[0110] The second indication is used to indicate that the audio editing model expands the audio to be edited based on the content that needs to be supplemented;

[0111] The third indication is used to indicate that the audio editing model corrects errors in the audio to be edited.

[0112] The specific functions and examples of each module and sub-module of the apparatus of the embodiments of the present disclosure are described above in the related description of the corresponding steps in the method embodiments, and will not be described here.

[0113] In the technical solutions of the present disclosure, the acquisition, storage and application of the user's personal information all comply with the relevant legal regulations and do not violate public order and good customs.

[0114] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0115] Figure 4A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0116] As shown in Figure 4 The device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 402 or a computer program loaded into a random access memory (RAM) 403 from a storage unit 408. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0117] Various components in the device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, a magneto-optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0118] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above, such as the detection method. For example, in some embodiments, the detection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of the detection method described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the detection method by any other appropriate means, such as by means of firmware.

[0119] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0120] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0122] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0123] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0124] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0125] It should be understood that the various forms of flow shown above can be re-ordered, steps added or removed, etc. For example, the steps recited in the present disclosure can be performed in parallel, in series, in a different order, etc., as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.

[0126] The above detailed description does not constitute a limitation of the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. An audio editing method, applied to an audio editing system including middleware and an audio editing model, the method comprising: The middleware inputs the audio to be edited and the editing instructions into the audio editing model; The audio editing model determines corresponding editing suggestions based on the audio to be edited and the editing instructions, and feeds the editing suggestions back to the middleware; The middleware determines the corresponding editing materials based on the editing suggestions, constructs prompt words based on the editing suggestions and the editing materials, and inputs the prompt words and the editing materials into the audio editing model; The audio editing model edits the audio to be edited based on the prompt words and the editing materials to obtain the edited audio.

2. The method according to claim 1, wherein, The audio editing model determines corresponding editing suggestions based on the audio to be edited and the editing instructions, including: The audio editing model identifies the audio to be edited in order to determine the text corresponding to the audio to be edited. The audio editing model determines the corresponding editing suggestions based on the text and the editing instructions.

3. The method according to claim 2, wherein, The editing instructions include at least one of content editing instructions and style editing instructions.

4. The method according to claim 2 or 3, wherein, The editorial suggestions include at least one of editorial suggestions for audio content and editorial suggestions for audio style.

5. The method according to claim 4, wherein, The audio editing model determines the corresponding editing suggestions based on the text and the editing instructions, including: The audio editing model modifies the text based on the text and the content editing instructions; and generates editing suggestions for the audio content based on the modified text and the timestamps of each part of the text; and / or, The audio editing model determines the domain corresponding to the text; and generates the editing suggestions for the audio style based on the domain and the style editing instructions.

6. The method according to claim 5, wherein, The editing materials include the content that needs to be added, and at least one of the audio templates; The middleware determines the corresponding editing materials based on the editing suggestions, including: The middleware searches for the content that needs to be supplemented from a pre-set content library based on the editing suggestions for the audio content; And / or, The middleware searches for the corresponding audio template from a pre-set audio template library based on the editing suggestions for the audio style.

7. The method according to claim 6, wherein, The prompt word contains at least one of the following: A first instruction, wherein the first instruction is used to instruct the audio editing model to modify at least one of the style and background sound of the audio to be edited according to the audio template; The second instruction is used to instruct the audio editing model to expand the audio to be edited based on the content that needs to be supplemented; The third instruction is used to instruct the audio editing model to correct errors in the audio to be edited.

8. An audio editing system, the system comprising middleware and an audio editing model; wherein, The middleware is used to input the audio to be edited and the editing instructions into the audio editing model, and to receive the editing suggestions returned by the audio editing model; Based on the editing suggestions, corresponding editing materials are determined; prompt words are constructed based on the editing suggestions and the editing materials; and the prompt words and the editing materials are input into the audio editing model. The audio editing model is used to determine the corresponding editing suggestions based on the audio to be edited and the editing instructions, and to feed the editing suggestions back to the middleware; it is also used to receive the prompt words and the editing materials, and to edit the audio to be edited based on the prompt words and the editing materials to obtain the edited audio.

9. The system according to claim 8, wherein, The audio editing model is used for: Identify the audio to be edited to determine the text corresponding to the audio; Based on the text and the editing instructions, the corresponding editing suggestions are determined.

10. The system according to claim 9, wherein, The editing instructions include at least one of content editing instructions and style editing instructions.

11. The system according to claim 9 or 10, wherein, The editorial suggestions include at least one of editorial suggestions for audio content and editorial suggestions for audio style.

12. The system according to claim 11, wherein, The audio editing model is used for: Based on the text and the content editing instructions, the text is modified; and based on the modified text and the timestamps of each part of the text, the editing suggestions for the audio content are generated. And / or, Determine the domain corresponding to the text; Based on the domain and the style editing instructions, the editing suggestions for the audio style are generated.

13. The system according to claim 12, wherein, The editing materials include the content that needs to be added, and at least one of the audio templates; The middleware is used for: Based on the editing suggestions for the audio content, search for the content that needs to be supplemented from a pre-set content library; and / or, Based on the editing suggestions for audio style, the corresponding audio template is searched from the pre-set audio template library.

14. The system according to claim 13, wherein, The prompt word contains at least one of the following: A first instruction, wherein the first instruction is used to instruct the audio editing model to modify at least one of the style and background sound of the audio to be edited according to the audio template; The second instruction is used to instruct the audio editing model to expand the audio to be edited based on the content that needs to be supplemented; The third instruction is used to instruct the audio editing model to correct errors in the audio to be edited.

15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Audio editing method and device, electronic equipment and storage medium

    CN113724686A

  • Audio processing method and electronic equipment

    CN117294990A