A method, apparatus, device, medium and program product for generating an audio file
By selecting and editing audio segments in the audio file generation method, and using a music generation model to generate target audio segments with similar styles, the problem of music generation deviation caused by unclear user text descriptions is solved, thus improving the user experience.
Patent Information
- Application Number
- CN202411826761.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-12-11
AI Technical Summary
In existing technologies, when users generate music through text, the text may be unclear or fail to accurately describe the expected music generation, resulting in a discrepancy between the generated music and the user's expectations, leading to a poor user experience.
An audio file generation method is provided, which displays an editing control on the playback page and displays an audio editing page in response to interactive operations. Users can select a reference audio segment from a preset audio file and use a music generation model to generate a target audio segment with a similar style based on the reference audio segment.
By selecting audio segments to generate target audio segments, the system meets users' expectations for music generation, reduces the difficulty of music creation, and improves the user experience.
Smart Images

Figure CN119653174B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to computer technology, and more particularly to a method, apparatus, device, medium, and program product for generating audio files. Background Technology
[0002] Computer technology is widely used in music processing, and more and more users are creating music through music clients. Currently, users can generate music by inputting text into the music client. However, this method requires users to input text to prompt the model to generate music. If the text is unclear or fails to accurately describe the expected music generation, the generated music may deviate from the user's expectations, resulting in a poor user experience. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, medium, and program product for generating audio files, which can generate audio that meets user expectations and improve user experience.
[0004] In a first aspect, embodiments of this disclosure provide a method for generating audio files, including:
[0005] An editing control is displayed on the playback page, which is used to play preset audio files;
[0006] In response to an interactive operation on the editing control, an audio editing page is displayed, the audio editing page including audio information of the preset audio file and a generation control;
[0007] In response to a selection operation on the audio information, a reference audio segment in the preset audio file is determined;
[0008] In response to an interactive operation on the generation control, a target audio segment is generated based on the reference audio segment, and a target audio file is determined based on the target audio segment.
[0009] Secondly, embodiments of this disclosure also provide an audio file generation apparatus, the apparatus comprising:
[0010] An editing control display module is used to display editing controls on a playback page, wherein the playback page is used to play a preset audio file;
[0011] The editing page display module is used to display an audio editing page in response to interactive operations on the editing control. The audio editing page includes audio information of the preset audio file and generation controls.
[0012] An audio segment selection module is used to determine a reference audio segment in the preset audio file in response to a selection operation on the audio information.
[0013] An audio generation module is used to respond to interactive operations on the generation control, generate a target audio segment based on the reference audio segment, and determine a target audio file based on the target audio segment.
[0014] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0015] One or more processors;
[0016] Storage device for storing one or more programs.
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating audio files as described in any embodiment of this disclosure.
[0018] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the audio file generation method as described in any embodiment of this disclosure.
[0019] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the method for generating audio files as described in any embodiment of this disclosure.
[0020] This disclosure provides a method for generating audio files. An editing control is displayed on a playback page, which plays a preset audio file. In response to interactive operations on the editing control, an audio editing page is displayed, including audio information of the preset audio file and a generation control. In response to a selection operation on the audio information, a reference audio segment in the preset audio file is determined. In response to an interactive operation on the generation control, a target audio segment is generated based on the reference audio segment. A target audio file is determined based on the target audio segment, and the target audio file is played. The technical solution of this disclosure, by selecting a reference audio segment in the currently playing preset audio file and generating a target audio segment with a similar style based on the reference audio segment, addresses the user's expectations for music generation. Since audio contains richer knowledge than text, the target audio segment generated based on the user-selected reference audio segment can meet the user's expectations for music generation, reducing the difficulty of music creation and improving the user experience. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0022] Figure 1 A flowchart illustrating a method for generating an audio file according to an embodiment of this disclosure;
[0023] Figure 2 This is a schematic diagram of a playback page provided in an embodiment of the present disclosure;
[0024] Figure 3 This is a schematic diagram of an audio editing page provided in an embodiment of the present disclosure;
[0025] Figure 4 This is a schematic diagram of another audio editing page provided in an embodiment of this disclosure;
[0026] Figure 5 This is a schematic diagram of yet another audio editing page provided in an embodiment of the present disclosure;
[0027] Figure 6 A schematic flowchart illustrating another method for generating audio files provided in this embodiment of the present disclosure;
[0028] Figure 7 This is a schematic diagram of yet another audio editing page provided in an embodiment of the present disclosure;
[0029] Figure 8 This is a schematic diagram of the structure of an audio file generation apparatus provided in an embodiment of the present disclosure;
[0030] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0032] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0033] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0035] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0036] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0037] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0038] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0039] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0040] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0041] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0042] Figure 1 This is a flowchart illustrating an audio file generation method provided in an embodiment of the present disclosure. This embodiment is applicable to music creation. The method can be executed by an audio file generation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.
[0043] like Figure 1 As shown, the method includes:
[0044] S110. Display editing controls on the playback page, wherein the playback page is used to play preset audio files.
[0045] The preset audio files include historically generated audio files. For example, users can create their own music through the music client, upload their music to the server, or share their music. The preset audio files may include music created and uploaded by users through the music client.
[0046] The playback page includes the album art, lyrics, author, and editing controls for the currently playing preset audio file. The editing controls trigger a music generation event to generate music in a similar style. The music generation event generates a new audio segment based on the selected audio segment. Clicking the editing control triggers this event. If the music generation event is triggered, the music editing page is displayed. For example, the playback page displays the album art corresponding to the preset audio file, and the editing controls are displayed in the corresponding position on the album art. For instance, the editing controls can be displayed at the bottom of the album art. By displaying the editing controls in the corresponding position on the album art, the entry point for the similar music generation service is clearly shown, facilitating the triggering of the music generation event for the currently playing audio file while listening to music. This eliminates the need to search and select an audio file from the playlist and then trigger the similar music generation event for the selected audio file, simplifying the user experience.
[0047] Figure 2 This is a schematic diagram of a playback page provided in an embodiment of this disclosure. Figure 2 As shown, the music client's playback page 210 displays the music cover and lyrics in the corresponding position on the music cover. The lyrics scroll synchronously according to the music playback time, and an editing control 220 is also displayed at the bottom of the music cover.
[0048] S120. In response to an interactive operation on the editing control, an audio editing page is displayed, the audio editing page including audio information of the preset audio file and a generation control.
[0049] In this embodiment, the audio editing page represents the selection page for audio segments. The audio editing page may include audio information and generation controls. The audio information may represent the audio duration of a preset audio file. For example, the audio information may include an audio bar, a progress bar, or a timestamp for lyrics. The audio bar may include an audio waveform. The audio waveform may represent the relationship between the volume and time of the preset audio file.
[0050] In this embodiment of the disclosure, the interactive operation is an operation such as clicking on the editing control, voice command, visual gaze, or gesture.
[0051] For example, in response to a click operation of the editing control, the audio editing page is displayed. The audio editing page also includes a lyrics panel and a text control, wherein the lyrics panel is used to display the lyrics of the preset audio file. The audio bar is located between the lyrics panel and the text control. Through this embodiment of the disclosure, the lyrics of the preset audio file can be displayed intuitively, making it easier for users to discover creative inspiration from the lyrics of the preset audio file. In addition, displaying the text control on the audio editing page allows for text input operations on the text control, determining prompts for the music generation model based on the input target text, so that the target audio segment generated by the music generation model better matches the user's expectations.
[0052] Optionally, the sub-segments corresponding to the audio bar are converted to a playback state based on the playback time of the preset audio file. A lyrics panel is displayed above the audio bar. The lyrics of the preset audio file are scrolled through in the lyrics panel, and the corresponding lyrics are converted to a selected state based on the playback time of the preset audio file. Optionally, if the audio bar is an audio waveform, the sub-segments of the audio bar may include audio waveform segments. For example, based on the playback time of the preset audio file, the audio waveform segment corresponding to the currently playing audio segment is adjusted to a target color. The display effect of the lyrics corresponding to the audio segment in playback state is also adjusted. Optionally, text processing such as bolding can be applied to the lyrics corresponding to the audio segment in playback state.
[0053] Figure 3 This is a schematic diagram of an audio editing page provided in an embodiment of this disclosure. Figure 3 As shown, an audio editing page 310 is displayed above the playback page 300. Optionally, the audio editing page 310 can be a floating window. The audio editing page 310 includes an audio bar 320, a lyrics panel 330, a text control 340, and a generation control 350, etc. Among them, the lyrics in the lyrics panel 330 correspond to the sub-segments of the audio bar 320.
[0054] A generation control is used to trigger a music generation event. If a music generation event is detected, a reference audio segment is fed into a pre-trained generative model to generate a target audio segment with a similar musical style based on the reference audio segment. The pre-trained generative model may include a music generation model that identifies musical features such as timbre, style, and arrangement of the input audio segment to generate a new audio segment similar to the input audio segment. Optionally, the diffusion model can be supervised and fine-tuned to obtain the music generation model.
[0055] S130. In response to the selection operation for the audio information, a reference audio segment in the preset audio file is determined.
[0056] The reference audio segment represents the selected audio segment in the preset audio file. The reference audio segment can be used as a prompt word and input into the music generation model to generate the target audio segment.
[0057] For example, in response to a drag operation on the audio information, a reference audio segment in the preset audio file is determined based on the audio position corresponding to the drag operation. The reference audio segment is played in a loop, and the audio information corresponding to the reference audio segment is converted into a selected state. This embodiment of the disclosure allows for the selection of a reference audio segment simply by dragging the audio information, providing an intuitive way to select audio segments, enriching the interaction methods, and simplifying the audio segment selection process.
[0058] Optionally, based on the duration requirements of the input audio segment in the music generation model, time-related prompts can be displayed at the corresponding positions in the audio information to indicate the duration range of the selected reference audio segment to the user. The duration of the reference audio segment between the audio positions corresponding to the drag operation must be within the aforementioned duration range. If the duration of the reference audio segment between the audio positions corresponding to the drag operation exceeds the upper limit of the aforementioned duration range, an audio segment that meets the duration range is extracted from the reference audio segment and used as the reference audio segment. If the duration of the reference audio segment between the audio positions corresponding to the drag operation is less than the lower limit of the duration range, the reference audio segment corresponding to the current drag operation is ignored, and a duration insufficient prompt is displayed. For example, if the duration range is greater than or equal to t1 and less than or equal to t2, and the duration of the reference audio segment is t2+a seconds, the t2-second audio data of the reference audio segment can be extracted from the beginning position of the reference audio segment as the final reference audio segment. Optionally, the t2-second audio data of the reference audio segment can also be extracted from the end position of the reference audio segment as the final reference audio segment.
[0059] The audio position includes the audio time at the start and end of the drag operation. For example, if the preset audio file is 90 seconds long, and the file is dragged from second 30 to second 60, then 30 and 60 seconds are used as the start and end times of the reference audio segment, and the corresponding audio segment is extracted from the preset audio file based on these times. Optionally, a rectangle can be used to highlight the audio information corresponding to the reference audio segment, indicating that the audio information is selected. And / or, the rectangle can be made bold. And / or, a background can be added to the audio information within the rectangle. And / or, the color of the rectangle can be adjusted. And / or, the color of the background can be adjusted, etc.
[0060] Figure 4 This is a schematic diagram of another audio editing page provided in an embodiment of this disclosure. Figure 4 As shown, the audio editing page 410 includes audio information from preset audio files, represented by an audio bar 420. Dragging the audio bar 420 from the first audio position 430 to the second audio position 440, and using a rectangular frame to select a sub-segment of the audio bar 420 between the first audio position 430 and the second audio position 440, converts the selected sub-segment into a selected state. The lyrics corresponding to the selected sub-segment are then played in a loop on the lyrics panel 450. The red-highlighted area in the audio bar 420 indicates the currently playing audio.
[0061] Optionally, if the audio information of the preset audio file includes a progress bar, the progress bar represents the total duration and the duration already played of the preset audio file. A reference audio segment in the preset audio file can be selected by dragging the progress bar.
[0062] Optionally, the preset audio file is divided into at least two audio segments based on the lyrics of the preset audio file. The audio information of the preset audio file includes the start and end times of each audio segment. In response to the input audio start and end times, a reference audio segment in the preset audio file is determined.
[0063] Optionally, if the audio information of the preset audio file includes the total audio duration and the played time, in response to the input audio segment duration, a reference audio segment in the preset audio file is determined based on the played time and the audio segment duration. If the total audio duration of the preset audio file is 120 seconds, the played time is s seconds, the audio segment duration is n seconds, and s+n seconds is less than or equal to 120 seconds, then an audio segment of length n seconds is extracted starting from the (s+1)th second of the preset audio file as the reference audio segment. If s+n exceeds 120 seconds, an audio selection error message is displayed.
[0064] S140. In response to an interactive operation on the generation control, generate a target audio segment based on the reference audio segment, and determine a target audio file based on the target audio segment.
[0065] In this embodiment, after selecting a reference audio segment, if a click operation on the generation control is detected, the reference audio segment is input into a music generation model. The music generation model learns at least one musical attribute from the reference audio segment, including timbre, style, arrangement, and lyrics, to generate a target audio segment with a similar musical style. The reference audio segment and the target audio segment are concatenated to obtain a target audio file, which is then played. To ensure natural transitions between audio segments, a cross-gradient can be applied to the reference and target audio segments. The target audio file is played to verify whether the generated audio meets expectations.
[0066] Optionally, in response to a text input operation on the text control, the lyrics panel in the audio editing page is hidden, and the positions of the audio bar and the text control are adjusted. The target text is determined based on the text input operation, wherein the target text includes the lyrics of the target audio segment and / or the lyrics description of the target audio segment.
[0067] The lyrics description of the target audio segment is used to describe the lyrics of the target audio segment. For example, the lyrics description of the target audio segment includes at least one of the following: description of the lyrics structure, description of the lyrics content, and description of the lyrics style. Optionally, the lyrics description may include lyrics expressing a walk on the beach, including one verse and two choruses. Optionally, the number of sentences included in the verse may also be limited, as may the number of sentences included in the choruses.
[0068] For example, in response to a click on a text control, the lyrics panel in the audio editing page is hidden, and the audio bar and text control are moved to the top of the audio editing page to display the keyboard at the bottom of the audio editing page, allowing text to be entered into the corresponding position of the text control via the keyboard.
[0069] The step of generating a target audio segment based on the reference audio segment in response to an interactive operation on the generation control includes: generating the target audio segment based on the reference audio segment and the target text in response to an interactive operation on the generation control.
[0070] If the target text includes lyrics describing the target audio segment, target lyrics for the target audio segment are generated based on the lyrics description. The target audio segment is generated based on at least one musical attribute from the reference audio segment, including timbre, style, and arrangement, as well as the target lyrics.
[0071] For example, after selecting a reference audio segment and entering lyrics in the corresponding position of the text control, if a click operation of the generation control is detected, the reference audio segment and lyrics are input into the music generation model. The music generation model then generates the target audio segment based on at least one of the musical attributes of the reference audio segment, such as timbre, style, and arrangement, as well as the input lyrics.
[0072] Optionally, after selecting a reference audio segment and entering lyrics description in the corresponding position of the text control, if a click operation of the generation control is detected, the reference audio segment and lyrics description are input into the music generation model, which then generates target lyrics for the target audio segment based on the lyrics description. Then, the audio generation model generates the target audio segment based on at least one musical attribute of the reference audio segment, including timbre, style, and arrangement, as well as the target lyrics.
[0073] Through the embodiments disclosed herein, music can be composed based on lyrics written by a user, eliminating the need for the user to have composition skills, thus reducing the difficulty of music creation. Alternatively, lyrics can be input into a music generation model to generate lyrics, further reducing the difficulty of lyric creation.
[0074] Optionally, if no target text is entered and a click operation of the generation control is detected, the reference audio segment is input into the music generation model, and the music generation model generates a model audio segment based on at least one of the timbre, style, and arrangement of the reference audio segment and the lyrics of the reference audio segment.
[0075] Optionally, the audio editing page further includes a time option. In response to a selection operation on the time option, a target time is determined. The step of generating a target audio segment based on the reference audio segment in response to an interaction operation on the generation control includes: generating the target audio segment based on the reference audio segment and the target time in response to an interaction operation on the generation control. By selecting a target time, the audio duration of the target audio segment generated by the music generation model can be constrained.
[0076] Figure 5 This is a schematic diagram of yet another audio editing page provided in an embodiment of this disclosure. Figure 5 As shown, the audio editing page also includes a time option 510. The time option 510 is located between the generation control 520 and the text control 530. In response to a selection operation on the time option 510, the selected time option 510 is set as the target time. Clicking the text control 530 hides the lyrics panel 540 and the time option 510 in the audio editing page, and controls the position of the audio bar 550, the text control 530, and the generation control 520 to move vertically upwards along the audio editing page. A keyboard 560 is displayed at the bottom of the audio editing page.
[0077] Optionally, the lyrics of the target audio segment are entered in the text control. After selecting a target time, the reference audio segment, the lyrics of the target audio segment, and the target time are input into the music generation model. The music generation model then generates the target audio segment based on the reference audio segment, the lyrics of the target audio segment, and the target time. The target time is used to constrain the duration of the target audio segment generated by the music generation model.
[0078] Optionally, if the duration of the lyrics of the target audio segment is less than the target time, the target audio segment includes at least two lyrics, and the target audio segment corresponding to the target time is obtained by looping the lyrics.
[0079] The technical solution of this disclosure embodiment displays an editing control on a playback page. The playback page is used to play a preset audio file. In response to the interactive operation of the editing control, an audio editing page is displayed. The audio editing page includes audio information of the preset audio file and a generation control. In response to a selection operation on the audio information, a reference audio segment in the preset audio file is determined. In response to an interactive operation on the generation control, a target audio segment is generated based on the reference audio segment. A target audio file is determined based on the target audio segment, and the target audio file is played. The technical solution of this disclosure embodiment, by selecting a reference audio segment in the currently playing preset audio file and generating a target audio segment with a similar style based on the reference audio segment, can meet the user's music generation expectations, reduce the difficulty of music creation, and improve the user experience, since audio contains richer knowledge than text.
[0080] Figure 6 This is a flowchart illustrating another method for generating an audio file provided by an embodiment of the present disclosure. Based on the above embodiments, this embodiment further defines the method of dividing the audio information into at least two audio information segments based on the lyrics of the preset audio file, and determining the reference audio segment in the preset audio file according to the audio position corresponding to the drag operation in response to a drag operation on the audio information.
[0081] like Figure 6 As shown, the method includes:
[0082] S610. Display editing controls on the playback page, wherein the playback page is used to play preset audio files.
[0083] S620. In response to an interactive operation on the editing control, an audio editing page is displayed, the audio editing page including audio information of the preset audio file and a generation control.
[0084] S630, In response to the selection operation for the audio information segment, a reference audio segment in the preset audio file is determined based on the selected audio information segment.
[0085] In this embodiment of the disclosure, the audio information is divided into at least two audio information segments according to the lyrics structure of a preset audio file. At least two audio information segments are displayed on the audio editing page. Each audio information segment can be represented as a rectangle, and the corresponding audio waveform is displayed within the rectangle. A reference audio segment is determined based on the selected audio information segment through a selection operation on at least one audio information segment.
[0086] In this embodiment of the disclosure, the selected audio information segment is a consecutive audio information segment among at least two audio information segments to ensure the continuity of the content of the generated target audio segment.
[0087] S640, loop the reference audio segment, and convert the audio information segment corresponding to the reference audio segment into a selected state.
[0088] For example, after selecting a reference audio segment, the reference audio segment is played in a loop to wait for the user to input the target text or target time. The properties of the rectangle corresponding to the audio information segment of the reference audio segment are adjusted to indicate that the audio information segment is selected.
[0089] Figure 7 This is a schematic diagram of yet another audio editing page provided in an embodiment of this disclosure. Figure 7 As shown, the audio editing page 710 includes at least two audio waveform segments 720. The audio waveform segments 720 correspond to lyric sections. The audio waveform segments 720 are displayed between the lyric panel 730 and the text control 740. A generation control 750 is displayed at the bottom of the audio editing page 710. A time option 760 is also displayed between the text control 740 and the generation control 750. In response to a selection operation on an audio waveform segment 720, the rectangle corresponding to the selected audio waveform segment 720 is thickened to indicate that the audio waveform segment 720 is selected.
[0090] S650, In response to an interactive operation on the generation control, a target audio segment is generated based on the reference audio segment, and a target audio file is determined based on the target audio segment.
[0091] The technical solution of this disclosure divides audio information into at least two audio information segments based on the lyrics structure of a preset audio file. A reference audio segment is determined through a selection operation on the audio information segments, and then a target audio segment is generated based on the reference audio segment. Since the selected audio information frequency band is a continuous audio information segment, the continuity of the generated target audio segment can be ensured, improving the audio generation quality.
[0092] Figure 8 This is a schematic diagram of an audio file generation device provided in an embodiment of the present disclosure. The device can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, a PC, or a server. Figure 8 As shown, the device includes: an editing control display module 810, an editing page display module 820, an audio segment selection module 830, and an audio generation module 840.
[0093] The editing control display module 810 is used to display editing controls on the playback page, wherein the playback page is used to play a preset audio file;
[0094] The editing page display module 820 is used to display an audio editing page in response to interactive operations on the editing control. The audio editing page includes audio information of the preset audio file and generation controls.
[0095] The audio segment selection module 830 is used to determine a reference audio segment in the preset audio file in response to a selection operation on the audio information.
[0096] The audio generation module 840 is configured to, in response to an interactive operation on the generation control, generate a target audio segment based on the reference audio segment, and determine a target audio file based on the target audio segment.
[0097] Optionally, the audio segment selection module 830 is specifically used for:
[0098] In response to a drag operation on the audio information, a reference audio segment in the preset audio file is determined based on the audio position corresponding to the drag operation;
[0099] The reference audio segment is played in a loop, and the audio information corresponding to the reference audio segment is converted into a selected state.
[0100] Furthermore, if the audio information is divided into at least two audio information segments based on the lyrics of the preset audio file, then the step of determining the reference audio segment in the preset audio file according to the audio position corresponding to the drag operation in response to the drag operation includes:
[0101] In response to the selection operation for the audio information segment, a reference audio segment in the preset audio file is determined based on the selected audio information segment.
[0102] Optionally, the edit control display module 810 is specifically used for:
[0103] The playback page displays the album art corresponding to the preset audio file, and the editing controls are displayed in the corresponding position on the album art.
[0104] Optionally, the audio information includes an audio bar, and the editing page display module 820 is specifically used for:
[0105] In response to a click on the editing control, the audio editing page is displayed. The audio editing page also includes a lyrics panel and a text control. The lyrics panel is used to display the lyrics of the preset audio file, and the audio bar is located between the lyrics panel and the text control.
[0106] Optionally, it also includes:
[0107] In response to a text input operation on the text control, the lyrics panel in the audio editing page is hidden, and the positions of the audio bar and the text control are adjusted. The target text is determined based on the text input operation, wherein the target text includes the lyrics of the target audio segment and / or the lyrics description of the target audio segment.
[0108] The audio generation module 840 is specifically used for:
[0109] In response to an interactive operation on the generation control, the target audio segment is generated based on the reference audio segment and the target text.
[0110] Further, generating the target audio segment based on the reference audio segment and the target text includes:
[0111] If the target text includes lyrics descriptions of the target audio segment, generate target lyrics for the target audio segment based on the lyrics descriptions;
[0112] The target audio segment is generated based on at least one musical attribute from the reference audio segment, including its timbre, style, and arrangement, as well as the target lyrics.
[0113] Optionally, the audio editing page also includes a time option;
[0114] The device further includes a time determination module for determining a target time in response to a selection operation for the time option;
[0115] The audio generation module 840 is specifically used for:
[0116] In response to an interactive operation on the generation control, the target audio segment is generated based on the reference audio segment and the target time.
[0117] Optionally, the audio generation module 840 is also specifically used for:
[0118] The reference audio segment and the target audio segment are connected to obtain the target audio file, and the target audio file is played.
[0119] The audio file generation apparatus provided in this disclosure can execute the audio file generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0120] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0121] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 9 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 9 The diagram below shows the structure of the terminal device or server 900. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0122] like Figure 9 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An edit / output (I / O) interface 905 is also connected to the bus 904.
[0123] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0124] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0125] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0126] The electronic device provided in this embodiment and the audio file generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0127] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the audio file generation method provided in the above embodiments.
[0128] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0129] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0130] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0131] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0132] An editing control is displayed on the playback page, which is used to play preset audio files;
[0133] In response to an interactive operation on the editing control, an audio editing page is displayed, the audio editing page including audio information of the preset audio file and a generation control;
[0134] In response to a selection operation on the audio information, a reference audio segment in the preset audio file is determined;
[0135] In response to an interactive operation on the generation control, a target audio segment is generated based on the reference audio segment, and a target audio file is determined based on the target audio segment.
[0136] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0138] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0139] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0141] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0142] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0143] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method of generating an audio file, characterized by, The method comprises the following steps: displaying an editing control on a playing page, wherein the playing page is used for playing a preset audio file; in response to an interaction operation on the editing control, displaying an audio editing page, wherein the audio editing page comprises audio information of the preset audio file and a generation control, and the audio information represents an audio duration of the preset audio file; in response to a selection operation on the audio information, determining a reference audio segment in the preset audio file; in response to an interaction operation on the generation control, generating a target audio segment according to the reference audio segment, and determining a target audio file according to the target audio segment; wherein the generating of the target audio segment according to the reference audio segment comprises: inputting the reference audio segment into a music generation model, learning at least one music attribute of the reference audio segment, such as tone, style, arrangement and lyrics, through the music generation model, and generating a target audio segment similar in music style to the reference audio segment; the determining of the target audio file according to the target audio segment comprises:
2. The method of claim 1, wherein, connecting the reference audio segment and the target audio segment to obtain the target audio file. The determining of the reference audio segment in the preset audio file in response to the selection operation on the audio information comprises: in response to a drag operation on the audio information, determining the reference audio segment in the preset audio file according to an audio position corresponding to the drag operation; 3. The method of claim 2, wherein, cyclically playing the reference audio segment, and converting the audio information corresponding to the reference audio segment into a selected state. If the audio information is divided into at least two audio information segments based on the lyrics of the preset audio file, the determining of the reference audio segment in the preset audio file in response to the drag operation on the audio information comprises:
4. The method of claim 1, wherein, in response to a selection operation on the audio information segment, determining the reference audio segment in the preset audio file according to the selected audio information segment. The displaying of the editing control on the playing page comprises:
5. The method of claim 1, wherein, displaying a music cover corresponding to the preset audio file on the playing page, and displaying the editing control at a corresponding position of the music cover. The audio information comprises an audio bar, and the displaying of the audio editing page in response to the interaction operation on the editing control comprises:
6. The method of claim 5, wherein, in response to a click operation on the editing control, displaying the audio editing page, wherein the audio editing page further comprises a lyrics panel and a text control, the lyrics panel is used for displaying lyrics of the preset audio file, and the audio bar is located between the lyrics panel and the text control. Further comprising: in response to a text input operation on the text control, hiding the lyrics panel in the audio editing page, adjusting positions of the audio bar and the text control, and determining a target text according to the text input operation, wherein the target text comprises lyrics of the target audio segment and / or lyrics description of the target audio segment; the generating of the target audio segment according to the reference audio segment in response to the interaction operation on the generation control comprises: In response to the interaction operation on the generation control, the target audio segment is generated according to the reference audio segment and target text.
7. The method of claim 6, wherein, The generating the target audio segment according to the reference audio segment and target text comprises: If the target text comprises a lyric description of the target audio segment, the target lyric of the target audio segment is generated according to the lyric description; The target audio segment is generated according to at least one of the timbre, style, and arrangement of the reference audio segment and the target lyric.
8. The method of claim 1, wherein, The audio editing page further comprises a time option; The method further comprises: in response to a selection operation on the time option, determining a target time; The generating the target audio segment according to the reference audio segment in response to the interaction operation on the generation control comprises: In response to the interaction operation on the generation control, the target audio segment is generated according to the reference audio segment and target time.
9. An apparatus for generating an audio file, characterized by Comprise: An editing control display module is configured to display an editing control on a playing page, wherein the playing page is configured to play a preset audio file; An editing page display module is configured to display an audio editing page in response to an interaction operation on the editing control, wherein the audio editing page comprises audio information of the preset audio file and a generation control, and the audio information represents an audio duration of the preset audio file; An audio segment selection module is configured to determine a reference audio segment in the preset audio file in response to a selection operation on the audio information; An audio generation module is configured to generate a target audio segment according to the reference audio segment in response to an interaction operation on the generation control, and determine a target audio file according to the target audio segment; The generating the target audio segment according to the reference audio segment comprises: inputting the reference audio segment into a music generation model, learning at least one of the timbre, style, arrangement, and lyrics of the reference audio segment through the music generation model, and generating a target audio segment similar to the music style of the reference audio segment; The determining the target audio file according to the target audio segment comprises: connecting the reference audio segment and the target audio segment to obtain the target audio file.
10. An electronic device, comprising: The electronic device comprises: one or more processors; a storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating an audio file according to any one of claims 1-8.
11. A storage medium containing computer-executable instructions, wherein: The computer executable instructions, when executed by a computer processor, are used to execute the method for generating an audio file according to any one of claims 1-8.
12. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method for generating an audio file according to any one of claims 1-8.
Citation Information
Patent Citations
Audio editing method, electronic equipment and storage medium
CN114023301A
Information processing method and device, electronic equipment and storage medium
CN115065840A
Multimedia data sharing method and device, equipment and medium
CN115643244A
Music generation method, device and system and storage medium
CN117012170A
Music generation method and device, electronic equipment and storage medium
CN119107922A
Cited By
Audio file generation method, device, and medium
EP4760721A1