Media data editing method and device, electronic equipment, storage medium and product
By integrating audio editing functionality into the media content creation page, the problem of workflow interruption caused by switching audio tools in multimedia content creation has been solved, enabling seamless collaborative editing of audio and visual materials, and improving creation efficiency and user experience.
Patent Information
- Application Number
- CN202511299554.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, users need to frequently switch audio editing tools during the multimedia content creation process, which leads to interruptions in the creation process, high costs of context switching, and lengthy operation paths, making it impossible to achieve seamless collaborative editing of audio elements and visual materials.
By integrating audio editing functionality into the media content creation page, and dynamically displaying the corresponding media data editing page in response to user actions, seamless integration of audio processing and the main creation workflow is achieved, allowing users to perform real-time synchronous editing of audio data in a unified environment.
It significantly reduces context switching costs, improves creative efficiency and the synergy between audio elements and visual materials, provides a seamless, integrated creative experience, and enhances the production quality and user satisfaction of multimedia content.
Smart Images

Figure CN121037601A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer processing, and in particular, to a media data editing method and device, electronic equipment, storage medium and product. BACKGROUND
[0002] In the current field of media content creation, especially in the production process of multimedia content such as short videos and podcasts, users generally rely on independent applications or modular tools with scattered functions for audio processing and content integration.
[0003] Existing solutions usually require users to trigger audio editing in the main creation environment, and then be forced to switch to reuse existing audio templates or import processed audio files, rather than realizing real embedded audio creation and adjustment. This operation mode causes the creation process to be forcibly interrupted, resulting in high context switching cost, long operation path, and difficulty in real-time collaborative editing. Not only does it significantly reduce the creation efficiency, but it also makes the synchronization adjustment of audio elements and visual materials complex, making it impossible to achieve a truly seamless integrated creation experience, ultimately affecting the production quality of multimedia content and user satisfaction. SUMMARY
[0004] The present disclosure provides a media data editing method, device, electronic equipment, storage medium and product to realize seamless connection of creation pages by deeply integrating audio editing functions, reduce operation interruption and switching cost, improve editing efficiency and audio-visual collaboration, and ultimately enhance user experience and content quality.
[0005] In a first aspect, the embodiments of the present disclosure provide a media data editing method, which comprises:
[0006] displaying a media content creation page; wherein the media content creation page comprises a function list for generating media data, and the function list at least includes an audio editing function item;
[0007] in response to a triggering operation on any audio editing function item, displaying a media data editing page corresponding to the triggered audio editing function item, to determine first audio data corresponding to the triggered audio editing function based on editing operations in the media data editing page;
[0008] in response to detecting an editing completion event of the first audio data, determining first media content based on the edited first audio data.
[0009] In a second aspect, the embodiments of the present disclosure also provide a media data editing device, which comprises:
[0010] The creation page display module is configured to display a media content creation page, wherein the media content creation page comprises a function list for generating media data, and the function list comprises at least an audio editing function item;
[0011] The function item editing module is configured to, in response to a triggering operation on any audio editing function item, display a media data editing page corresponding to the triggered audio editing function item, to determine first audio data corresponding to the triggered audio editing function based on an editing operation in the media data editing page.
[0012] The media content generation module is configured to, in response to detecting an editing completion event of the first audio data, determine first media content based on the edited first audio data.
[0013] In a third aspect, the embodiments of the present disclosure further provide an electronic device, which comprises:
[0014] one or more processors;
[0015] a storage device configured to store one or more programs,
[0016] when the one or more programs are executed by the one or more processors, the one or more processors implement the media data editing method according to any of the embodiments of the present disclosure.
[0017] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions for executing the media data editing method according to any of the embodiments of the present disclosure when executed by a computer processor.
[0018] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product comprising a computer program for implementing the media data editing method according to any of the embodiments of the present disclosure when executed by a processor.
[0019] The technical scheme of the embodiments of the present disclosure is that, in response to detecting a triggering operation of a user for generating media data, a media content creation page is displayed, wherein the media content creation page includes a function list for generating media data, and the function list at least includes an audio editing function item. Then, in response to a triggering operation of any audio editing function item, a media data editing page corresponding to the triggered audio editing function item is displayed, so as to determine first audio data corresponding to the triggered audio editing function based on an editing operation in the media data editing page. Finally, in response to detecting an editing completion event of the first audio data, first media content is determined based on the edited first audio data. The embodiments of the present disclosure effectively solve the problem of interruption of the creation process caused by fragmentation of function modules by deeply integrating the audio editing function into the media content creation page. When the user triggers any audio editing function item, the corresponding media data editing page can be presented in a unified creation environment, realizing seamless connection of audio processing and main creation process, significantly reducing the context switching cost, shortening the operation path, enabling real-time synchronization of audio data to the main project, greatly improving the creation efficiency, ensuring accurate collaborative editing of audio elements and visual materials, and finally providing an integrated and smooth creation experience for the user, significantly enhancing the quality and user satisfaction of multimedia content output. BRIEF DESCRIPTION OF DRAWINGS
[0020] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0021] Figure 1 A flowchart of a media data editing method provided by the embodiments of the present disclosure;
[0022] Figure 2 A media data editing page related to the embodiments of the present disclosure;
[0023] Figure 3 A flowchart of another media data editing method provided by the embodiments of the present disclosure;
[0024] Figure 4 A material editing panel related to the embodiments of the present disclosure;
[0025] Figure 5 A media content creation page when a user imports different data content, related to the embodiments of the present disclosure;
[0026] Figure 6 A flowchart of another media data editing method provided by the embodiments of the present disclosure;
[0027] Figure 7 A display content schematic diagram in a media content creation page when a user triggers a rewrite function item according to an embodiment of the present disclosure;
[0028] Figure 8 A background execution logic schematic diagram for generating second audio data according to an embodiment of the present disclosure;
[0029] Figure 9 A flow schematic diagram of another media data editing method according to an embodiment of the present disclosure;
[0030] Figure 10 A display content schematic diagram in a media content creation page when a user triggers a sound color changing function item according to an embodiment of the present disclosure;
[0031] Figure 11 Still another display content schematic diagram in a media content creation page when a user triggers a sound color changing function item according to an embodiment of the present disclosure;
[0032] Figure 12 A display content schematic diagram in a media content creation page when a user triggers a melody changing function item according to an embodiment of the present disclosure;
[0033] Figure 13 A flow schematic diagram of another media data editing method according to an embodiment of the present disclosure;
[0034] Figure 14 A display content schematic diagram in a media content creation page when a user triggers a script writing function item according to an embodiment of the present disclosure;
[0035] Figure 15 A structure schematic diagram of a media data editing apparatus according to an embodiment of the present disclosure;
[0036] Figure 16 A structure schematic diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION
[0037] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided to more thoroughly and completely understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0038] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0039] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0040] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0041] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0042] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0043] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0044] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0045] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0046] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0047] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0048] Before introducing this technical solution, an exemplary application scenario can be provided. The technical solution provided by this disclosure can be applied to scenarios where audio editing and media content generation are seamlessly integrated in a unified creation environment. For example, this technical solution can be applied to the creation of media content on a short video platform. On the short video platform, the user first enters the media content creation page, clicks the "Rewrite" function, and a media data editing page with the original text and timecode immediately slides out on the right. The user changes "Limited-time offer" to "Today's direct price reduction" and generates a new reading with one click; then clicks the "Change Voice" function, and switches to a card-style voice library in the same location. The user changes the male voice to a "vibrant female voice" to obtain the updated audio; finally, clicks the "Speed Change" function, and a speed slider pops up at the bottom of the page. The user adjusts the speech speed from 1× to 1.2× while retaining the pitch, and the final first audio data is immediately output and the sequence is automatically filled back. The entire process is completed continuously within the corresponding expanded dedicated editing page, and the text, voice, and speech speed can be modified without jumping to other windows.
[0049] Figure 1 This is a flowchart illustrating a media data editing method provided in an embodiment of the present disclosure. This embodiment can be applied to any situation where audio editing and media content generation need to be seamlessly integrated in a unified creation environment. The method can be executed by a media data editing device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.
[0050] like Figure 1 As shown, the method in this embodiment may specifically include:
[0051] S110. Display the media content creation page; wherein the media content creation page includes a list of functions for generating media data, and the list of functions includes at least an audio editing function.
[0052] The media content creation page is an integrated user interface. Its core function is to provide users with an entry point for media data generation and editing. This page displays various audio editing functions in a structured list, allowing users to intuitively select the desired operation and trigger the corresponding specialized editing environment to perform specific data generation or modification tasks.
[0053] Media data refers to comprehensive content carriers that contain auditory, visual, or combined audiovisual information, recorded, stored, and processed digitally. The function list refers to structured graphical user interface elements that centrally display all available audio editing functions in an ordered or categorized manner. Audio editing functions are specific interactive elements integrated into the function list of the media content creation interface. They are abstract encapsulations and visual representations of audio processing capabilities, transforming underlying audio processing technology into explicit, user-triggerable commands through predefined operational logic. This allows users to enter a dedicated editing environment by activating the function, thereby enabling targeted generation, modification, or optimization of audio data.
[0054] In this embodiment, when a user launches the audio / video editing application and clicks the media data generation control, the system detects the user's trigger action to generate media data and displays a main workbench interface on the device screen. This interface is the media content creation page. The center of this page features a preview canvas for viewing the final product, while the top or sides contain horizontally arranged or layered menus composed of icons and text representing different audio editing functions—this is the function list. By visually perceiving this page layout, the user clearly understands that they are at a starting point where they can begin synthesizing or modifying multimedia content such as video and audio by clicking on these function items.
[0055] Optionally, the media content creation page includes at least one audio track for displaying audio data. At least one audio track is in an empty track state before any audio data is configured. Audio data refers to sound information encoded and stored in digital form, and is the fundamental raw material constituting the auditory elements in multimedia content. An audio track is an independent channel layer in the multimedia editing interface used to carry and linearly arrange audio data. Essentially, it is a visual timeline container that graphically represents the duration, positional relationship, and attribute status of audio signals, allowing users to perform overall operations on the data within a single audio track. An empty track state means that the audio track exists as a pre-allocated but inactive containerized resource in the media content creation page. This is manifested by the fact that the audio track channel does not yet carry any valid audio data entities, and visually it is usually presented as a blank timeline track or an invalid area with a placeholder mark.
[0056] In this embodiment, in the initial architecture of the media content creation page, one or more visual channels dedicated to presenting audio information are pre-set, namely audio tracks. These channels remain empty tracks before being assigned specific audio content by the user—that is, they exist as blank containers with a timeline dimension but no actual data carrier. Their function is to provide a structured framework and clear visual placeholders for subsequent audio editing operations, which not only ensures the integrity of interface elements, but also reserves scalable operation space for dynamic injection and hierarchical management of audio data.
[0057] Optionally, the media content creation page also includes at least one media track for displaying media data, including image data and / or video data. At least one media track is in an empty track state when no media data is configured. Media data refers to non-audio media resources that can serve as basic editing elements in the multimedia creation environment, including static image data and / or dynamic video data. Media data is used in conjunction with audio data to synthesize the final multimedia content product. A media track is an independent channel layer in the multimedia editing interface dedicated to carrying and linearly organizing visual media resources (such as images and videos). It is a visual container with temporal attributes, representing the duration, positional relationship, and overlay level of visual materials through a graphical timeline, allowing users to crop, move, or add visual effects to materials on a single track. An empty track state indicates that the track has not yet been assigned any valid visual data content.
[0058] In this embodiment, in the overall architecture of the media content creation page, one or more independent channel layers, namely material tracks, are pre-allocated to carry visual media resources (such as image sequences or video streams). These channels are in an empty track state before receiving specific visual data configured by the user—that is, as structured blank areas with timeline attributes but no actual visual content carriers. Their function is to provide a basic container framework for the import, chronological arrangement, hierarchical management and special effects compositing of image and video materials, while maintaining the functional integrity and operational predictability of the editing interface through visual placeholders.
[0059] Optionally, the audio editing functions include one or more of the following: rewriting function to adjust the text content corresponding to the audio data, changing the timbre of the audio data, changing the melody of the audio data, performing voice separation on the audio data, adjusting the playback speed of the audio data, voice reading function, text writing assistance function, sound effect addition function, and sound source addition function.
[0060] The rewrite function is an interactive entry point within the audio editing functions, specifically designed for modifying the text content corresponding to the audio data. For example, the rewrite function can convert the speech information contained in the audio into an editable text carrier, allowing users to perform semantic reconstruction, sentence structure adjustment, or word replacement on the text. Then, through speech synthesis technology, the modified text is used to regenerate new audio data that conforms to the timing and intonation characteristics of the original audio, thereby achieving the secondary creative purpose of indirectly controlling the expression of audio content through text-level editing.
[0061] The "Timbre Change" function is an interactive entry point within the audio editing functions specifically designed for modifying the timbre characteristics of vocals or instruments in audio data. For example, the timbre change function allows for real-time analysis and reconstruction of audio spectral characteristics, enabling users to alter the harmonic structure, formant distribution, and timbre attributes of the sound by adjusting parameters or selecting preset modes while maintaining the original melody, rhythm, and content. This generates new audio data with different timbre expressiveness (such as gender transformation, age change, or instrument replacement) but consistent content.
[0062] The "Change Melody" function is an interactive entry point specifically designed for modifying the melody structure in audio data. For example, it can analyze the pitch, rhythm, and harmony features of the original audio using music information retrieval technology, and then reorganize, vary, or stylize the melody sequence based on music theory rules or machine learning models. This allows users to regenerate new audio data with different melodic lines while maintaining musical harmony, by adjusting parameters or applying templates, while preserving the original audio timbre and content.
[0063] Among them, the vocal separation function is an interactive entry point in the audio editing function set up to isolate and extract the vocal components from background music and ambient sound effects in mixed audio data. For example, it can use deep learning or spectrum analysis algorithms to decompose audio signals into multiple sound sources, allowing users to intelligently identify, separate, and independently output vocal tracks and non-vocal elements in the original audio by adjusting parameters or selecting modes, thereby generating new audio data segments that retain only pure vocals or isolated background sounds.
[0064] The speed adjustment function is an interactive entry point within the audio editing functions specifically designed for modifying the playback time reference of audio data. For example, digital signal processing algorithms (such as time stretching or resampling techniques) can be used to non-linearly compress or expand the time axis of the audio waveform, allowing users to change the overall playback speed of the audio by adjusting parameters or using sliders while maintaining the original pitch and timbre characteristics. This results in generating new audio data with a faster or slower tempo but unchanged pitch.
[0065] Among them, the voice reading function is an interactive entry point in the audio editing function set up specifically for converting text content into synthesized speech. For example, it can use speech synthesis technology to perform linguistic analysis and acoustic modeling on the input text to generate humanized audio data with natural rhythm and clear sound quality. Users can customize the voice output effect by adjusting parameters such as speech rate, intonation, emotional tone, or speaker, thereby directly creating standardized reading audio materials without the need for manual recording.
[0066] Among them, the copywriting assistance feature is an interactive entry point within the audio editing features specifically designed to assist in generating or optimizing text content. For example, it can intelligently analyze the user's original text or semantic needs using natural language processing technology, and provide creative support such as text expansion, abbreviation, stylization conversion, or keyword optimization based on a pre-trained language model. Ultimately, it outputs standardized copy that meets the needs of specific scenarios, providing a processable text foundation for subsequent voice reading or audio content creation.
[0067] The sound effects addition function is an interactive entry point within the audio editing functions, specifically designed to overlay preset or custom sound effects onto existing audio data. For example, digital audio mixing technology can be used to temporally align and parametrically blend ambient sounds, skeuomorphic sounds, or artistic sound elements with the original audio. Users can add a sense of scene, dramatic tension, or auditory embellishment to the audio by selecting sound effect library resources, adjusting effect intensity, and spatial parameters, thereby generating new composite audio data with auditory layers.
[0068] The audio source addition function is an interactive entry point within the audio editing functions specifically designed for importing or generating external audio resources and integrating them into the creative project. For example, through file system access or real-time audio stream capture technology, users can import locally stored audio files, third-party audio source library materials, or raw audio signals input from hardware devices into the current editing environment. Users can then browse, preview, and select these audio data as independent audio tracks or mixed materials to insert into the timeline, thereby expanding the basic raw material library of available audio content.
[0069] In this embodiment, the audio editing function set is an integration of tool modules provided on the media content creation page for modifying or generating multi-dimensional attributes of audio data. It covers the entire operational chain from underlying text content and acoustic features to external resource integration, allowing users to trigger different function items to achieve differentiated editing goals such as semantic reconstruction of audio-related text, timbre spectrum adjustment, melody line modification, separation of vocals and background sounds, playback rate scaling, text-to-speech generation, auxiliary copywriting, sound effect overlay, and import of external audio resources. These function items can be called in any combination according to actual needs. The advantage of this setup is that by integrating multi-dimensional audio processing capabilities into a unified function list, it significantly improves the efficiency and flexibility of media creation.
[0070] Based on the above embodiments, see the schematic diagram of the media data editing page. Figure 2 ,like Figure 2 As shown, the blank area in the middle of the page represents a preview canvas for viewing the final product; the rectangle labeled "+Add Material" represents the material track, and the rectangle labeled "+Add Audio" represents the audio track; the area at the bottom of the page represents the function list, and each small rectangle in the function list represents a different audio editing function.
[0071] S120. In response to a trigger operation on any audio editing function, display a media data editing page corresponding to the triggered audio editing function, so as to determine the first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page.
[0072] The triggered audio editing function refers to a specific audio processing module actively selected and activated by the user through clicking, touching, or other interactive methods in the function list of the media content creation page. The media data editing page is a dedicated interactive interface that is dynamically loaded in response to the triggering of a specific audio editing function. It can be understood as a highly customized operating environment for the preset editing goals of the function. By centrally presenting parameter controls, preview areas, and process guidance elements, it encapsulates the complex underlying audio processing process into a visual user task flow, enabling users to generate new audio data that conforms to the function definition through intuitive operation, while ensuring the contextual coherence of the editing process and the overall creation workflow.
[0073] The first audio data refers to the resulting audio file or data stream generated after a user performs a series of editing operations on a specific triggered audio editing function item through the media data editing page. Essentially, it is the output product of the original audio after processing by this function item, retaining both the targeted modification characteristics of the editing intent and the integrity and usability to be integrated into the final media content as an independent audio element.
[0074] Specifically, when a user activates an audio editing function in the function list, a dedicated operation interface matching the predefined editing logic of that function will be immediately invoked, which is the media data editing page. Through the parameter controls, real-time preview, and workflow guidance tools integrated into this interface, the editing operations performed by the user on this interface are transformed into targeted audio processing instructions, and finally an output audio data entity that carries all editing results and meets the preset goals of the function is generated, thus obtaining the first audio data.
[0075] Based on the above embodiments, when a user triggers the rewrite function, a media data editing page dedicated to text modification can be loaded. This page uses speech recognition technology to convert the original audio into editable text and presents a text input box and semantic optimization tools. The text addition, deletion, sentence restructuring, or style adjustment operations performed by the user on this interface will trigger the speech synthesis engine to generate corresponding new audio in real time. Finally, the modified text is rendered into output audio data that combines editing intent and natural speech flow through an acoustic model, which is the first audio data.
[0076] When a user triggers the timbre change function, a media data editing page dedicated to timbre adjustment is loaded. This page uses audio analysis technology to analyze the spectral characteristics of the original audio and presents timbre preset templates, gender conversion sliders, and harmonic parameter controls. The timbre mode or custom parameters selected by the user on this interface will drive the digital signal processing algorithm in real time to reconstruct the audio formants and sound source filters, and finally generate output audio data that retains the original content but has the target timbre characteristics, which is the first audio data.
[0077] When a user triggers the melody change function, a media data editing page dedicated to melody reconstruction is loaded. This page uses music information retrieval technology to analyze the pitch sequence and rhythm pattern of the original audio and presents a melody template library, scale adjustment tools, and rhythmic mapping controls. The melody variation scheme or custom pitch curve editing operation selected by the user on this interface will drive the audio resynthesis algorithm in real time to perform structural replacement and harmonic adaptation of the original melody, and finally generate output audio data that retains the original timbre and lyrics but has a new melodic line, which is the first audio data.
[0078] When a user triggers the pedestrian voice separation function, a media data editing page dedicated to sound source separation is loaded. This page uses deep learning algorithms to perform multi-source analysis on the original mixed audio and presents a slider for adjusting the separation intensity of human voice and background music, a spectrum visualization view, and an option to export independent audio tracks. The user's interactive operation of adjusting separation parameters or selecting output mode on this interface will drive the neural network model in real time to extract human voice components and suppress ambient noise in the audio signal, and finally generate output audio data that retains only pure human voice or background music, which is the first audio data.
[0079] When a user triggers the speed adjustment function, a media data editing page dedicated to timing adjustment is loaded. This page analyzes the waveform characteristics of the original audio through an audio time stretching algorithm and presents a playback rate slider, a pitch hold switch, and a rhythm density visualization tool. When the user drags the rate parameter or selects a preset multiplier on this interface, the phase vocoder will perform time axis compression / expansion processing on the audio samples in real time, ultimately generating output audio data that retains the original pitch but has a new playback speed, which is the first audio data.
[0080] When a user triggers the speed adjustment function, a media data editing page dedicated to timing adjustment is loaded. This page analyzes the waveform characteristics of the original audio through an audio time stretching algorithm and presents a playback rate slider, a pitch hold switch, and a rhythm density visualization tool. When the user drags the rate parameter or selects a preset multiplier on this interface, the phase vocoder will perform time axis compression / expansion processing on the audio samples in real time, ultimately generating output audio data that retains the original pitch but has a new playback speed, which is the first audio data.
[0081] When a user triggers the pedestrian voice reading function, a media data editing page dedicated to text-to-speech synthesis is loaded. This page provides a text input area, a voice style option library (such as gender, age, emotional tone), a speech rate adjustment slider, and real-time preview playback controls through the integrated speech synthesis engine. The user's operation of inputting text content and adjusting voice parameters on this interface will generate a speech waveform with acoustic features in real time, and finally convert the edited text into human-read audio data with natural rhythm, which is the first audio data.
[0082] Optionally, when the audio editing function is triggered in relation to the text reading function, the media data editing page also includes a text editing area and a timbre selection option to read the text in the text editing area aloud based on the selected timbre.
[0083] The text editing area refers to an interactive input box or editable text canvas on the media data editing page specifically used for receiving, displaying, and modifying text content. The timbre selection options refer to a set of parameterized controls on the media data editing page specifically used for configuring the acoustic features of speech synthesis. Essentially, they provide users with a selection interface for differentiated timbre output by using preset acoustic model parameter sets, such as fundamental frequency range, formant distribution, and speaker identifiers.
[0084] In this embodiment, when the triggered audio editing function involves a text reading function, the two core components, text editing area and timbre selection, can be dynamically loaded in the media data editing page. The text editing area provides an interactive interface for text input and modification, while the timbre selection provides acoustic feature configuration capabilities. Finally, based on the text content and timbre parameters determined by the user on the page, the speech synthesis engine is driven to generate reading audio data with specified timbre features.
[0085] When a user triggers the copywriting assistance feature, a media data editing page dedicated to text-assisted creation is loaded. This page provides keyword input boxes, style options (such as formal, colloquial, and poetic), content length sliders, and AI-generated suggestion panels through natural language processing technology. The user's actions of setting text requirements and selecting optimization directions on this interface will drive the language model to generate a draft copy that conforms to semantic logic in real time. Finally, the edited and confirmed text content will be used as the source data for subsequent audio synthesis or modification, which is the text basis corresponding to the first audio data.
[0086] When a user triggers the sound effect addition function, a media data editing page dedicated to sound effect mixing is loaded. This page provides a sound effect library navigation interface, a real-time spectrum display, a volume balance slider, and a spatial parameter panel (such as reverb and sound panning) through audio layering technology. When the user selects sound effect materials and adjusts the overlay parameters on this interface, the multi-track mixing engine will perform time-series alignment and frequency band fusion with the original audio in real time, and finally generate composite audio data that embeds the target sound effect and maintains acoustic harmony, which is the first audio data.
[0087] When a user triggers the audio source addition function, a media data editing page dedicated to importing external audio resources is loaded. This page provides a local / cloud audio library browsing window, a format compatibility detection module, and audio track mapping options through a file system interface. When the user selects the target audio file and configures the insertion position on this interface, the decoder will parse the audio source and unify the sampling rate in real time, ultimately seamlessly integrating the external audio data into standardized audio material that can be used in the current project, which is the first audio data.
[0088] S130, In response to detecting a completion event of editing the first audio data, determine the first media content based on the first audio data that has been edited.
[0089] The "edit completion" event refers to the signal emitted by the user through a specific interactive action after performing all operations related to the current audio editing function on the media data editing page, indicating the termination of the editing state. This specific interactive action could be a confirmation button click, a gesture command, or an automatic save trigger. The "first media content" refers to the composite media product formed by integrating the edited first audio data with other media elements in the current creative project.
[0090] Specifically, when user interaction signals (such as confirmation or timed saving) detect that the first audio data has been edited, the media assembly engine is immediately triggered. This engine aligns the newly generated audio data with the existing visual material tracks on the creation page along the timeline, verifies audio-visual synchronization, and performs format conversion. Ultimately, it generates a structured media content entity with complete playback attributes, integrating the audio editing results with multi-track media elements—this is the first media content. In essence, the first media content is the final deliverable after the user's audio output, processed using specific audio editing functions, is synchronized, layered, and format-packaged with existing visual materials along the timeline. It retains the core modification features of the audio editing operation while possessing the audiovisual harmony and structural integrity required for complete multimedia content.
[0091] For example, taking the rewrite function as the triggered audio editing function item, when a user changes the text "Today the weather is great" to "Today the sky is clear" in the editing page of the rewrite function item and clicks to confirm the export (i.e., triggering the editing completion event), the audio corresponding to the new text (first audio data) can be generated through speech synthesis technology. Then, the audio is automatically embedded into the main audio track of the creation page, and the duration is matched and the audio and video are synchronized with the blue sky and white clouds in the video material. Finally, a complete short video containing the voice-over of "Today the sky is clear" and the corresponding visual content is rendered, which is the first media content.
[0092] Taking the voice change function as an example, when a user adjusts the original children's voice audio to a mature male voice through voice parameters in the voice change function's editing page and clicks to confirm the synthesis (i.e., triggering the editing completion event), a dubbing audio with new voice characteristics (first audio data) will be generated. Then, the audio will be automatically aligned with the preset science video footage in the creation page in terms of duration and lip-sync optimization. Finally, a complete educational video with a mature male voice as the narration track and matching science visual content will be rendered, which is the first media content.
[0093] Taking the "Change Melody" function as an example of the audio editing function that is triggered, when a user adjusts the original lyrical melody to a light and upbeat rhythm through a melody template in the editing page of the "Change Melody" function and clicks to confirm and generate (i.e., triggering the editing completion event), a background music audio with the characteristics of the new melody (first audio data) will be generated. Then, the audio will be automatically synchronized with the preset travel video clips in the creation page in terms of beat and emotion, and finally a complete short video with a light and upbeat melody as background music that perfectly matches the rhythm of the travel scene will be rendered, which is the first media content.
[0094] Taking the voice separation function as an example, when a user successfully separates the human voice from the background music in the mixed audio and clicks to export the human voice track in the editing page of the voice separation function (i.e., triggering the editing completion event), a clean human voice dubbing audio (first audio data) will be generated. Then, the human voice track will be automatically matched with the preset silent product demonstration video in the creation page for precise lip-syncing and duration matching. Finally, a complete advertising video with clear human voice narration as the core and completely synchronized with the product operation screen will be synthesized, which is the first media content.
[0095] Taking the speed-up function as an example of the audio editing function that is triggered, when the user speeds up the playback rate of the original explanatory audio from 1.0x to 1.5x in the editing page of the speed-up function and clicks confirm (i.e. triggering the editing completion event), a speed-up version of the explanatory audio (first audio data) with compressed duration but unchanged pitch will be generated. Then, the audio will be automatically synchronized with the corresponding software operation demonstration video in the creation page at the frame level, so that the video playback rate is synchronously increased to 1.5x and the audio and video are completely matched. Finally, a tutorial video with a shorter duration but higher information density will be generated, which is the first media content.
[0096] Taking the voice reading function as an example, when a user enters product introduction text in the editing page of the voice reading function and selects the "steady male voice" parameter and clicks generate (i.e. triggering the editing completion event), the corresponding professional narration audio (first audio data) will be synthesized. Then, the audio will be automatically matched with the preset product display animation in the creation page in terms of duration and keyframes, and finally a complete product promotion video with professional voice-over will be rendered, which is the first media content.
[0097] Taking the copywriting assistance function as an example, when a user enters the keyword "summer drinks" and selects "lively style" on the copywriting assistance function's editing page and clicks to generate copy (i.e., triggering the editing completion event), the copy text "Icy summer! This drink will instantly cool you down by 10℃" will be generated. Then, the voice reading function will be automatically called to convert it into a vibrant voice-over audio (first audio data). Finally, the audio will be matched with the drink making video material on the creation page in terms of duration and audio-visual synchronization to render and generate a complete product promotion video with creative narration, which is the first media content.
[0098] Taking the audio editing function as an example of adding sound effects, when a user selects the "sizzling oil" sound effect for the frying scene in a food video and adjusts the volume balance in the editing page of the sound effects adding function, and then clicks confirm (which triggers the editing completion event), an enhanced audio (first audio data) embedded with the ambient sound effect will be generated. Then, the audio will be automatically synchronized frame by frame with the cooking scene in the original video, so that the boiling oil scene and the sound effect are completely matched. Finally, a food making tutorial video with a stronger sense of audio and video immersion will be synthesized, which is the first media content.
[0099] Taking the audio editing function as an example of adding a sound source, when a user selects a background music file from a local file in the editing page of the sound source addition function and drags it to the 00:15 position on the timeline and clicks to confirm the import (i.e. triggering the editing completion event), the external audio file will be decoded into a standardized audio track (first audio data). Then, it will automatically synchronize the beat and reduce the volume with the already edited travel video footage in the creation page, so that the music will naturally cut in at 15 seconds and match the rhythm of the sunset transition scene in the video. Finally, a complete travel short film with carefully arranged background music will be generated, which is the first media content.
[0100] The technical solution of this disclosure embodiment, in response to detecting a user's trigger operation on generating media data, displays a media content creation page. The media content creation page includes a function list for generating media data, with at least an audio editing function item. Then, in response to a trigger operation on any audio editing function item, a media data editing page corresponding to the triggered audio editing function item is displayed. Based on the editing operation on the media data editing page, first audio data corresponding to the triggered audio editing function is determined. Finally, in response to detecting an editing completion event of the first audio data, first media content is determined based on the edited first audio data. This disclosure embodiment, by deeply integrating audio editing functions into the media content creation page, effectively solves the problem of interrupted creation processes caused by fragmented functional modules. When a user triggers any audio editing function item, the corresponding media data editing page can be presented in a unified creation environment, achieving seamless connection between audio processing and the main creation process. This significantly reduces context switching costs, shortens operation paths, and enables audio data to be synchronized to the main project in real time. This not only greatly improves creation efficiency but also ensures precise collaborative editing of audio elements and visual materials, ultimately providing users with an integrated and smooth creation experience, significantly enhancing the quality of multimedia content output and user satisfaction.
[0101] Based on obtaining the primary media content, this method further includes: in response to the event of publishing the primary media content, sending the primary media data to a primary platform. Here, the primary platform is distinct from the platform used to edit and generate the primary media content.
[0102] In this embodiment, when a user triggers a publishing instruction for completed media content, such as clicking the "Publish" button, the final generated first media content can be automatically sent to a designated external platform, namely the first platform, through an API interface or file transfer protocol. This platform is independent of the original system currently used for media editing and creation, realizing the physical separation of the editing environment and the publishing environment and automated data flow.
[0103] Figure 3 This is a flowchart illustrating another media data editing method provided in this embodiment. Based on the above embodiments, this embodiment provides a more detailed explanation of how to generate material data. For specific implementation details, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 3 As shown, the method in this embodiment may specifically include:
[0104] S210, Display the media content creation page.
[0105] The media content creation page includes a list of functions for generating media data, and the list includes at least audio editing functions.
[0106] Optionally, the media content creation page also includes a media editing function. This function, when triggered, allows adjustment of the media data displayed in at least one media track, as well as the associated audio data. Specifically, the media editing function refers to an interactive entry point on the media content creation page dedicated to handling the collaborative editing of visual materials and associated audio.
[0107] In this embodiment of the disclosure, by adding a material editing function to the media content creation page, a unified operation entry point is provided for users to collaboratively process visual materials and associated audio. When this function is triggered, the image / video data (such as cropping and filter application) in the material track and its time-related audio elements (such as background music and sound effects) can be adjusted in a coordinated manner to ensure that visual modifications and auditory effects maintain semantic and temporal synchronization, and finally output optimized material data with unified audiovisual presentation.
[0108] S220. In response to a trigger operation on a material editing function item, display a material editing panel to determine the first material data based on the editing operation in the material editing panel.
[0109] The media editing panel is a dedicated interface that dynamically loads in response to media editing function triggers. It can be viewed as a composite editing environment integrating visual media parameter controls and associated audio adjustment tools. By synchronously presenting image / video editing options (such as cropping boxes, filter sliders, and duration cutters) and their associated audio adjustment modules (such as volume curves and sound effect synchronizers) from the media track, users can modify visual media and their associated audio elements in a coordinated manner, ensuring optimized media data with both audiovisual harmony in the output.
[0110] In this embodiment, the material editing panel includes a smart screen type option and a digital human type option. The smart screen type option is used to determine the screen type of the first material data generated after being triggered, and the digital human type option is used to determine the display object after being triggered.
[0111] Among them, the intelligent image type option refers to the interactive control in the material editing panel that is dedicated to automatically generating visual content. In essence, it uses computer vision algorithms and scene understanding technology to convert audio content (such as voice rhythm, emotional tone or text semantics) into corresponding visual style parameters (such as landscape, city, cartoon, etc.). This allows users to intelligently generate standardized visual materials that are highly matched with the audio by selecting preset image types, thus realizing the semantic collaborative creation of audio and visual elements.
[0112] Among them, the Digital Human Type option is a dedicated function module in the material editing panel that provides the selection of virtual digital human images. Through the preset digital human body model and motion parameter library, users can select digital human characters with different characteristics (such as appearance, clothing, and movement patterns), and it supports driving the lip movements, facial expressions, and body movements of digital humans through audio data to achieve automatic synchronization and matching between virtual images and audio content.
[0113] In this embodiment, when a user triggers the material editing function, a dedicated material editing panel is displayed. This panel includes an intelligent visual type option and a digital human type option. The intelligent visual type option allows the user to select a visual style template that matches the audio content to generate material data for the corresponding visual type. The digital human type option provides the selection and configuration function for virtual digital human avatars to generate digital human demonstration materials synchronized with the audio. Ultimately, these operations output the first material data that meets the user's needs. The beneficial effect of this design is that it significantly improves the creation efficiency and consistency of multimedia content through integrated visual material generation tools. Specifically, the intelligent visual type option automatically matches the audio content and visual style through algorithms to ensure semantic coordination between the visuals and audio; the digital human type option provides a standardized calling interface for virtual avatars to achieve automated lip-syncing and motion matching between the digital human and the audio, ultimately lowering the professional production threshold while ensuring the quality and synergistic effect of the audiovisual materials.
[0114] Based on the above embodiments, optionally, the specific implementation method of determining the first material data based on the editing operation in the material editing panel may include: in response to the triggering operation of any picture style type associated with the picture type option, generating the first material data corresponding to the selected picture style type according to the sentence segmentation result corresponding to the first audio data.
[0115] Among them, the visual style type refers to the classification and identification of image or video content using pre-set visual feature templates. The sentence segmentation result refers to the sequence of independent sentence or phrase units divided according to semantics and pauses after the structured analysis of the text content corresponding to the audio using natural language processing technology. Each unit is marked with a timestamp and is used to achieve fine-grained segmentation matching between audio content and visual materials.
[0116] In this embodiment, after the user selects a certain visual style, the system automatically generates visual content matching the selected style for each sentence or paragraph based on the text segmentation results of the first audio data (i.e., paragraphs divided by semantics and timestamps). This achieves a precise temporal and semantic correspondence between audio and visuals, ultimately outputting a complete set of material data that meets the style requirements. The purpose of this setup is to significantly improve the efficiency and quality of visual material creation through an intelligent audio-visual matching mechanism. By automatically generating visuals matching the selected style for each semantic paragraph based on the audio segmentation results, it ensures that the visual content precisely corresponds to the audio rhythm, emotion, and theme. This avoids the tedious manual frame-by-frame editing and guarantees the professional-grade audiovisual harmony of the final product.
[0117] For example, when a user selects the "ink wash painting" visual style, the text segments after the sentences in the first audio data can be extracted, such as "Green mountains faintly visible, waters stretching far away" corresponding to 0-3 seconds, and "Autumn ends in Jiangnan, grass still withers" corresponding to 3-6 seconds. Then, dynamic visual materials that conform to the characteristics of ink wash painting are automatically generated for each time period, such as the effect of landscape blurring and ink diffusion, and finally combined into ink wash style video materials that are completely synchronized with the audio content and timeline.
[0118] Another example is a diagram of the material editing panel. Figure 4 ,like Figure 4 As shown in (a), the media editing panel includes Smart Screen Type and Digital Human Type options. When the user selects a Smart Screen option, multiple options can be displayed under the selected Smart Screen option. When the user selects a specific option and extracts the text segment after the sentence in the first audio data, such as "The kitten is sleeping" corresponding to 0-3 seconds, a media screen of the kitten sleeping can be generated, as shown in (a). Figure 4 As shown in (b).
[0119] Based on the above embodiments, optionally, the specific implementation of determining the first material data based on the editing operation in the material editing panel may further include: responding to the triggering operation of any digital human associated with the digital human type option, driving the selected digital human according to the first audio data to obtain the first material data.
[0120] Among them, digital humans refer to virtual character models that have highly human-like appearance, expressions, body movements and voice interaction capabilities. They can automatically generate lip-sync, facial expression changes and body movements based on audio or text input to simulate the audiovisual performance of real humans.
[0121] In this embodiment, when a user selects a digital human type, the voice information (such as pitch, rhythm, and word segmentation timestamps) in the first audio data can be used to automatically drive the digital human to generate corresponding lip movements, facial expressions, and body movements. This ensures that the virtual character's performance is completely synchronized with the audio content, ultimately outputting a digital human demonstration video as visual material data. The purpose of this setup is to significantly reduce the production threshold and cost of virtual human videos through automated digital human driving technology. Users only need to select the digital human type and provide audio to automatically generate high-quality digital human video materials with synchronized lip movements, natural expressions, and matching movements, without the need for professional motion capture equipment or manual parameter adjustments. This ensures accurate synchronization of audiovisual content and significantly improves the efficiency and expressiveness of multimedia creation.
[0122] For example, when a user selects the "virtual teacher" digital human type, the explanation content (such as math course audio) in the first audio data can be extracted, and the digital human model can be automatically driven to generate corresponding lip movements, blackboard gestures and facial expressions, and finally output a video of a virtual teacher giving a lesson, whose lip movements and actions are completely consistent with the audio content.
[0123] S230. In response to a trigger operation on any audio editing function, display a media data editing page corresponding to the triggered audio editing function, so as to determine the first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page.
[0124] S240. In response to detecting a completion event of editing the first audio data, determine the first media content based on the edited first audio data and the first material data.
[0125] In this embodiment, once the first audio data is detected as having been edited, the processed audio is automatically time-aligned, synchronized with the first material data (intelligent visuals or digital human videos) generated through the material editing panel, and format-encapsulated. Finally, a complete multimedia work with high-quality audio and matching visual elements is synthesized, which is the first media content. The automated synthesis mechanism seamlessly connects audio editing and visual creation. After the user completes the audio processing, the associated intelligent visuals and digital human materials are automatically synchronized, accurately matching the timeline and content theme. There is no need to manually align the audio and video tracks, which not only ensures professional-grade audio-visual synchronization quality but also significantly reduces the complexity and time cost of multimedia content production, achieving efficient one-stop media content production.
[0126] For example, when a user completes the editing of a product explanation audio (first audio data) and clicks confirm, the audio can be automatically time-aligned, volume balanced, and format-composite with the smart visuals (such as product demonstration animations) and / or digital human videos (virtual salesperson explanations) previously generated through the material editing panel, and finally outputs a complete advertising video (first media content) that includes synchronized narration, dynamic product display, and virtual human appearance.
[0127] In this embodiment, when editing media content on the media content creation page, users can import only audio data, only source material data, or import both audio and source material data sequentially. See the diagram below for illustrations of the media content creation page when users import different data content. Figure 5 .like Figure 5 As shown in (a), when a user initially opens the media content creation page, both the media track and the audio track are empty. At this time, the media track can be hidden, and the track area sliding event can be disabled, displaying a placeholder view for the media track; as... Figure 5 As shown in (b), when the user only imports media data, the media track placeholder view is hidden, the media track is displayed, and the media track area sliding event is disabled; Figure 5 As shown in (c), when the user only imports audio data, the audio track placeholder UI is hidden, the audio track clips are displayed, and the audio track area sliding event is disabled; Figure 5 As shown in (d), when a user imports both audio data and material data, the material track and audio track can be imported first, the track area sliding event can be unlocked, the material track placeholder can be hidden, and the audio track placeholder can be hidden.
[0128] In this embodiment, the media content creation page also includes a material editing function. This function, when triggered, adjusts the material data displayed in at least one material track and the audio data associated with it. The method further includes: in response to a trigger operation on the material editing function, displaying a material editing panel to determine first material data based on editing operations within the panel; wherein the material editing panel includes a smart image type option and a digital human type option. The image type option, when triggered, determines the image type of the generated first material data, and the digital human type option, when triggered, determines the display object. This method of generating the first material data provides users with a unified visual creation entry point through integrated material editing functions. The smart image type option enables automated matching of audio content and visual style, and the digital human type option provides standardized invocation of virtual avatars. This allows non-professional users to quickly generate high-quality visual materials that are highly coordinated with the audio, significantly reducing the technical threshold for audio-visual synchronization in multimedia creation and improving content production efficiency and professionalism.
[0129] Figure 6 This is a flowchart illustrating another media data editing method provided in this embodiment. Based on the above embodiments, this embodiment refines the method for determining the first audio data when the triggered audio editing function item is related to the rewrite function item. Specific implementation details can be found in the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 6 As shown, the method in this embodiment may specifically include:
[0130] S310, Display the media content creation page.
[0131] The media content creation page includes a list of functions for generating media data, which includes at least one audio editing function. These audio editing functions include one or more of the following: rewriting (adjusting the text content corresponding to the audio data), changing the timbre (adjusting the audio data's timbre), changing the melody (adjusting the audio data's melody), performing vocal separation on the audio data, adjusting the playback speed of the audio data, providing voice reading, assisting with copywriting, adding sound effects, and adding audio sources.
[0132] In this embodiment, when the triggered audio editing function item is related to the rewrite function item, and at least one audio track on the media content creation page displays the second audio data, the implementation method for determining the first audio data is as described in S320-S340, specifically including:
[0133] S320. In response to the triggering operation of the overwrite function item, display the media data editing page corresponding to the overwrite function item.
[0134] The media data editing page displays at least one first text content and its associated audio playback time after audio recognition of the second audio data. The first text content is arranged in chronological order according to the audio playback time. The second audio data refers to the pre-existing audio file or data stream in the audio track of the media content creation page, serving as the original object for editing operations. The first text content refers to the text sequence generated after speech recognition conversion of the second audio data, arranged in the original audio timeline order. The audio playback time refers to the time position information in the original second audio data corresponding to each text segment in the first text content. Essentially, it uses speech recognition technology to map the timeline coordinates of the audio waveform to the temporal markers of the text content, ensuring that the text editing operation remains synchronized with the temporal structure of the original audio in the form of timecodes (such as start / end timestamps or duration intervals), providing a precise time alignment basis for subsequent audio playback.
[0135] In this embodiment, when the user activates the rewrite function, the corresponding media data editing page is immediately rendered. This page first performs speech recognition on the second audio data, splits the speech information into several first text contents according to the timeline, and displays them in order from early to late according to the audio playback time associated with each segment. This allows the user to intuitively see the text results that correspond one-to-one with the original audio time progress in the visualized text sequence, providing a time reference and semantic units for subsequent text-based re-editing.
[0136] Optionally, the specific implementation of displaying the media data editing page corresponding to the rewrite function item in response to a triggered operation may include:
[0137] When a trigger operation on the rewrite function is detected, and it is determined that the audio track includes the second audio data, the second audio data is sent to the server so that when the voice detection module on the server determines that the second audio data meets the preset conditions, the text extraction module extracts the text information of the second audio data.
[0138] The server refers to a computing system or cluster located at the other end of the network, independent of the user's local terminal. It continuously listens for and responds to requests from the client through a preset communication protocol. It uses its own deployed algorithm modules, storage resources and computing power to perform human voice detection, text extraction and time alignment of text and audio on the uploaded second audio data in sequence, and returns the processing results to the client, thereby completing complex data conversion and synchronization tasks without consuming local device resources.
[0139] The voice detection module is a built-in algorithm unit on the server. It analyzes the acoustic characteristics of the second audio data, distinguishing between human and non-human voice components in terms of spectrum, energy, harmonic structure, and temporal continuity. Based on a preset confidence threshold or rule set, it determines whether the audio segment contains a valid human voice signal. The preset conditions refer to quantitative judgment criteria that are fixed or user-configurable before the voice detection module starts. These criteria exist in the form of acoustic feature thresholds, probability confidence levels, or rule combinations. During operation, each frame or segment of the second audio data is evaluated in real-time. The voice detection module determines the presence of a valid human voice and triggers the subsequent text extraction process only when the audio signal meets or exceeds the criteria in terms of energy magnitude, spectrum distribution, harmonic structure, and temporal continuity. Otherwise, it considers the audio invalid and terminates processing, ensuring that subsequent recognition and alignment steps are only performed on voice content that meets quality requirements. Preferably, the preset condition is that the second audio data is song audio.
[0140] The text extraction module is a speech recognition algorithm component embedded in the server. After the voice detection module confirms that the second audio data meets the preset conditions, it inputs the acoustic signal of the audio segment frame by frame into the joint decoding network of the acoustic model and the language model. By mapping the acoustic feature sequence into a sequence of words and symbols and outputting the corresponding text information, the voice content that originally existed in waveform form is transformed into editable string data for subsequent text alignment and display.
[0141] In practical applications, when a trigger operation on the rewrite function is detected, it can immediately check whether there is a second audio data already loaded in the current audio track. If it is confirmed to exist, the audio data segment is sent to the server in its entirety via the network protocol. After receiving the data, the server first uses the human voice detection module to evaluate its acoustic characteristics frame by frame. Only when the evaluation result meets the preset conditions of energy, spectrum and continuity is it determined that the audio segment contains a valid human voice and then the text extraction module is started to convert the audio signal into the corresponding text information, thereby completing a coherent automated process from audio uploading, human voice verification to text extraction.
[0142] The text alignment module integrated on the server performs time alignment on the text information and the second audio data, and provides feedback to obtain at least one first text content displayed on the media data editing page.
[0143] The text alignment module is a built-in temporal matching algorithm component on the server. It takes the discrete text sequence output by the text extraction module and the continuous time axis of the second audio data as dual inputs. Through acoustic model scoring, phoneme boundary estimation and dynamic programming strategy, it calculates the start and end times of each text or phrase in the audio stream and generates an alignment result with a precise timestamp. This allows the media data editing page to display the first text content segment by segment in the order of audio playback time and ensures that the text and audio rhythm remain synchronized when the user edits or synthesizes the audio.
[0144] In practical applications, the text alignment module integrated on the server can be used to perform time alignment of text information and second audio data. After receiving the text sequence and original second audio data output by the text extraction module, the module uses acoustic models, phoneme boundary detection and dynamic programming algorithms to calculate the precise start and end times of each text or phrase on the audio timeline, generate alignment results with corresponding timestamps, and feed the results back to the media data editing page, so that at least one first text content can be displayed segment by segment according to the audio playback time order, thereby ensuring that the text and audio rhythm remain synchronized, providing a reliable time reference for subsequent editing and synthesis.
[0145] In this embodiment, when the rewriting function is triggered and second audio data exists in the audio track, the audio is immediately uploaded to the server. The voice detection module first verifies that it meets the preset conditions, ensuring that the text extraction module is called to generate text only for valid voice segments. Then, the text alignment module accurately matches the text with the original audio to the timeline and feeds back the alignment result with the timestamp to the media data editing page. This significantly reduces the local computing load while ensuring recognition accuracy, reduces the waste of resources caused by invalid processing, and enables users to obtain an editable text sequence that corresponds one-to-one with the audio playback time in real time, improving the efficiency and synchronization accuracy of subsequent rewriting, synthesis, and dubbing.
[0146] S330. In response to any editing operation on the first text content, adjust the first text content displayed on the media data editing page.
[0147] The editing operations include one or more of the following: word modification, deletion of the first text content, and segment duration adjustment. Word modification refers to the user's interactive action of replacing or modifying words in the identified and displayed first text content. It should be noted that the number of characters in the first text content after the word modification operation is the same as the number of characters before the operation. Deletion of the first text content refers to the user's interactive instruction to remove one or more segments of the first text content that have been identified and are arranged in chronological order according to the audio playback time. Segment duration adjustment refers to the user's interactive action of changing the duration of the audio segment corresponding to a identified and displayed first text content.
[0148] In this embodiment, when a user performs an editing operation on any first text content on the media data editing page, the operation type can be immediately identified and the interface display updated in real time: the word modification operation changes the semantics of the text by replacing, inserting, or deleting words; the first text content deletion operation removes the entire text and its associated timestamp from the sequence and rearranges the remaining segments; and the segment duration adjustment operation changes the mapping relationship between the text and the audio by stretching or compressing the duration of the corresponding audio segment. After the above single or multiple operations are triggered in combination, the text order, timestamp distribution, and audio segment length will be adjusted synchronously, so that the first text content displayed on the media data editing page is completely consistent with the user's editing intention in terms of semantics, timing, and duration, and provides accurate text and time mapping data for subsequent audio regeneration.
[0149] For example, see the diagram showing the content displayed on the media content creation page when a user triggers the rewrite function. Figure 7 . Figure 7(a) A display page showing materials related to kittens and an audio file named "Happy New Year" that the user has imported; when the user triggers the "Recognize Subtitles" function on this display page (shown as control C in the figure), the lyrics in the second audio data can be recognized and displayed, such as... Figure 7 As shown in (b); furthermore, when the user triggers the "Rewrite Function" on the page shown in 7(b), the media data editing page corresponding to the Rewrite Function can be displayed, such as... Figure 7 As shown in (c); in Figure 7 (c) The displayed page includes first text content corresponding to each line of lyrics in the second audio data, and the audio playback time associated with each text content, and these first text contents are arranged in order of audio playback time. Figure 7 (c) On the page, if a user wants to edit a certain lyric, they can click on the control editing area corresponding to the first text content. At this time, the triggered first text content is in an editable state, and the user can perform "add", "delete", or "modify" editing operations on this first text content. Figure 7 (d) shows a diagram of word rewriting operations performed on each piece of the first text. Compared to the original first text, “Happy New Year” is rewritten as “Hello, little cat”, and “Wishing everyone a Happy New Year” is rewritten as “Wishing the little cat grows up well”. The other rewritten content is similar and will not be described in detail here.
[0150] S340, In response to the event that the first text content editing is completed, generate the first audio data based on the adjusted first text content, and return to the media content creation page.
[0151] In this embodiment, in response to the completion of the first text content editing event, the adjusted text sequence and its corresponding timestamp information are immediately sent to the speech synthesis engine. The engine regenerates first audio data that perfectly matches the edited text and is duration-aligned using an acoustic model and vocoder. Simultaneously, the current media data editing page is closed and the media content creation page is restored, allowing the new audio to automatically replace the original second audio data in the audio track, completing the closed loop from text-driven to audio update. The purpose of this setup is to achieve the core advantage of non-destructive audio editing through a two-way real-time linkage mechanism between text and audio: on the one hand, it allows users to indirectly manipulate audio content through intuitive text modifications (such as word changes, deletions, or duration adjustments) while preserving the original audio timeline structure, avoiding the technical hurdle of directly processing audio waveforms; on the other hand, the automatic triggering of the editing completion event seamlessly converts the text changes into next-generation audio data and returns to the main creation page, ensuring the accurate implementation of editing intentions and maintaining workflow continuity across editing stages, significantly improving modification efficiency and operational error tolerance.
[0152] Based on the above embodiments, optionally, the specific implementation of generating first audio data based on the adjusted first text content may include: generating first audio data with the same melody as the second audio data but different text content based on the adjusted first text content and the second audio data using the audio generation service integrated in the server.
[0153] Among them, the audio generation service is a computing module integrated on the server side for converting text sequences into audio with specific acoustic features. It can use speech synthesis technology combined with the melody parameters of the original audio (such as pitch curves and rhythm patterns) to re-synthesize the user-edited text content into an audio data stream that retains the original melody structure but updates the semantic content, thereby achieving precise fusion and regeneration of text changes and audio features.
[0154] In practical implementation, based on the server-integrated audio generation service, the adjusted first text content can be used as the new semantic input, and the melody, rhythm and pitch contours provided by the second audio data can be used as acoustic constraints. The text can be re-encoded into a new audio signal that is consistent with the original melody features through speech synthesis or singing synthesis models, thereby outputting the first audio data whose text content has been changed but whose melody, speed and emotional color are still the same as the original audio, realizing the synchronous generation of text rewriting and melody preservation.
[0155] In this embodiment, the server-side audio generation service is used as the hub. The adjusted text is used as the new semantic input, while the melody, rhythm and pitch contour of the second audio data are locked as acoustic constraints. The automatic synthesis of "text rewriting and melody preservation" can be completed with a single network request, eliminating the tedious steps of re-recording, matching, or post-mixing for users, and significantly reducing time, manpower and equipment costs.
[0156] An exemplary diagram illustrating the background execution logic for generating the second audio data can be found here. Figure 8 The system architecture for executing this media data editing method may include a user, a client, a server, and an audio generation server. The specific application flow is as follows: When the user clicks the trigger control corresponding to the "Rewrite Function Item" on the client, the client uploads the selected second audio data from the current editing page to the server. The server then returns the URL (Uniform Resource Locator) of the audio file. Subsequently, the client sends this URL to the server, which forwards it to the audio generation server and triggers the human voice detection module (i.e.,...). Figure 8 (AED audio detection steps).
[0157] The audio generation server returns the audio event type to the client. Upon receiving the audio event type, the client uploads the audio binary data and submits the subtitle recognition task to the audio generation server. At this point, the audio generation server returns a task identifier code. The client then initiates a polling process based on this task identifier code to query the audio generation server for the subtitle recognition task progress.
[0158] After the audio generation server returns the subtitle recognition results to the client, the client sends the original lyrics and the original audio URL of the second audio data to the audio generation server, requesting forced alignment processing between the lyrics and the audio. After processing, the audio generation server returns the strongly aligned lyrics. At this point, the client's display interface successfully enters the word editing interface.
[0159] After the user finishes editing the text, they click the "Generate Audio" control. The client then sends the strongly aligned lyrics timing data, the original audio URL, and the revised lyrics to the audio generation server. The audio generation server generates the modified new audio and returns its URL. The client retrieves the new audio using this URL, replaces the original audio and subtitle data locally, and finally displays the new audio and subtitles on the client interface.
[0160] S350, In response to detecting a completion event of editing the first audio data, determine the first media content based on the first audio data that has been edited.
[0161] The technical solution of this embodiment, when the triggered audio editing function item is related to the rewrite function item, determines the first audio data as follows: In response to the triggering operation of the rewrite function item, a media data editing page corresponding to the rewrite function item is displayed. Then, in response to the editing operation of any first text content, the first text content displayed in the media data editing page is adjusted. The editing operation includes one or more of the following: word modification operation, first text content deletion operation, and segment duration adjustment operation. Furthermore, in response to the event that the first text content editing is completed, first audio data is generated based on the adjusted first text content, and the user is redirected to the media content creation page. The technical solution provided by this embodiment converts the audio content into a text sequence arranged along the timeline and associates it with the original audio timestamp, allowing users to indirectly control the audio content by directly modifying the text. This avoids the technical hurdles of traditional waveform editing and ensures strict synchronization between semantic modification and audio timing. At the same time, after editing is completed, a new audio is automatically generated and the user returns to the main page, realizing non-destructive editing and a highly efficient closed loop of workflow, significantly improving the efficiency and accuracy of audio content modification.
[0162] Figure 9This is a flowchart illustrating another media data editing method provided in this embodiment. Based on the above embodiments, this embodiment refines the method for determining the first audio data when the triggered audio editing function item is related to the timbre-changing function item. For specific implementation details, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 9 As shown, the method in this embodiment may specifically include:
[0163] S410, Display the media content creation page.
[0164] The media content creation page includes a list of functions for generating media data, and the list includes at least audio editing functions.
[0165] In this embodiment, when the triggered audio editing function is related to the timbre-changing function, and at least one audio track on the media content creation page displays second audio data, the implementation method for determining the first audio data is as described in S420-S460, specifically including:
[0166] S420: In response to the triggering operation of the tone-changing function item, display the media data editing page corresponding to the tone-changing function item.
[0167] The media data editing page includes multiple preset timbre types and custom timbre types.
[0168] The preset timbre type is a set of timbre category tags pre-packaged and built into the media data editing page. Each tag corresponds to a set of trained acoustic features, used to quickly map the original human voice or instrument sound in the second audio data to the spectrum, formants, and pitch features of the target timbre. In this embodiment, the preset timbre type item is associated with multiple optional timbres. A set of trained discrete timbre instances (i.e., optional timbres) is pre-configured for each preset timbre category and stored in the cloud. These instances belong to the same type family in terms of acoustic features but have subtle differences. When the user triggers the type item, the system immediately expands the list of all timbre instances under it, allowing the user to quickly listen, compare, and select the target timbre that best meets their needs within the same category without additional collection or training, achieving one-click switching while maintaining a consistent timbre style.
[0169] Custom timbre type is a user-defined timbre feature that allows users to upload dry audio samples of a specified length that meet the specifications or record in real time. The cloud-based timbre training service extracts and models the acoustic features of the samples and generates a unique timbre file. This file is saved to the user's personal timbre library and can be called up during timbre changing operations to replace the timbre in the original audio, achieving a personalized sound effect that is different from the preset.
[0170] In this embodiment, when a user triggers the timbre-changing function, the corresponding media data editing page is immediately rendered on the current interface or in a newly opened editing area. This page contains two types of timbre sources: one is a pre-trained and categorized preset timbre type, with multiple ready-to-use timbre instances under each type; the other is a custom timbre type entry that allows users to upload samples and generate personalized timbres through cloud training. This presents all available timbre resources centrally on the same interface, allowing users to expand, preview, and replace them as needed. The purpose of this setup is to: provide a categorized and optimized standard timbre template library through preset timbre types, lowering the user's selection threshold and ensuring sound quality; and simultaneously meet personalized needs through custom timbre types, forming a two-level operation structure of categorized navigation and specific selection. This allows non-professional users to quickly locate the target timbre, while professional users can make fine adjustments, significantly improving the accuracy and efficiency of timbre replacement while ensuring intuitive operation.
[0171] Optionally, in this embodiment, before displaying the media data editing page corresponding to the color-changing function item, the method may further include: determining the audio type corresponding to the second audio data; and retrieving the preset timbre type corresponding to the audio type.
[0172] In this embodiment, feature analysis can first be performed on the second audio data to automatically determine its audio type based on its spectral distribution, fundamental frequency range, formant structure and rhythm density. Then, a set of preset timbre types that perfectly match the type can be retrieved and loaded from the preset timbre resource library to ensure that the optional timbres displayed later are consistent with the original audio in terms of acoustic properties, thereby avoiding sound quality distortion or style conflict after timbre replacement.
[0173] For example, when the second audio data is a male rap, by extracting its prominent low-frequency rhythm, fundamental frequency between 80 and 180 Hz and rapid formant changes, the audio type is automatically determined to be "male rap". Then, several timbre cards belonging to the same category as "male rap" are retrieved from the preset timbre library for the user to choose from, realizing quick color changing under type matching.
[0174] Based on this, if the user triggers the preset control corresponding to the preset tone type, then S430-S440 are executed; if the user triggers the preset control corresponding to the custom tone type, then S450-S460 are executed.
[0175] S430, in response to a trigger operation on any preset timbre type, displays multiple selectable timbres associated with the timbre type item.
[0176] In practice, when a user selects a preset timbre type, multiple optional timbre instances associated with that type node are immediately expanded and presented side-by-side in a list or card format on the media data editing page. This allows the user to browse all candidate timbres in the same category without having to switch interfaces, thus enabling them to quickly preview, compare, and determine the final replacement target.
[0177] For example, when a user clicks on the preset tone type "Female Folk Guitar" in the tone-changing interface, six tone cards will immediately expand horizontally below this type: "Warm Nylon Strings," "Bright Steel Strings," "Soft Fingerstyle," "Crisp Strumming," "Deep Pick," and "Sweet Harmony." Each card displays a waveform thumbnail on the left and the tone name, applicable range, and play button on the right. Users can click to listen to each card and compare subtle differences within the same type without leaving the current panel, thus quickly identifying the replacement tone that best suits their creative needs.
[0178] S440, In response to a trigger operation on any selectable timbre, update the timbre in the second audio data with the timbre corresponding to the triggered selectable timbre to obtain the first audio data.
[0179] The triggered tone options are presented in a preset format. The preset format refers to a standardized visual feedback style predefined for the triggered tone options. It uses dynamic elements of the graphical user interface (GUI) (such as highlight colors, border animations, icon changes, or size scaling) to visually enhance the currently selected tone option, providing clear operational feedback and distinguishing it from unselected options, ensuring the visibility of the status and the accuracy of the operation during the interaction process.
[0180] In this embodiment, when a user triggers a selectable timbre, the acoustic parameter set associated with that timbre option is immediately invoked. The spectral characteristics (such as formant distribution and timbre envelope) of the second audio data are reconstructed using a digital signal processing algorithm to generate first audio data that retains the original content but has new timbre characteristics. At the same time, the currently selected timbre option is fed back in real time through preset forms such as highlighting, borders or animations to ensure accurate transmission of the operation intention and intuitive perception of the timbre replacement effect.
[0181] For example, when a user triggers the voice change function, the displayed content on the media content creation page can be found in the following diagram. Figure 10 .like Figure 10 As shown in (a), when the user triggers the "Change Timbre" function on the page, the media data editing page corresponding to the "Change Timbre" function will be displayed, such as... Figure 10As shown in (b), this page displays several preset sound types, including "Favorites," "My," "Popular," "Commentary," "Cute," and "Foreign Language." If the user further activates the "Popular" preset sound type, it will expand to display multiple selectable sounds associated with that type; such as... Figure 10 As shown in (c), when the user selects "Optional tone 1", the icon for this option will be presented in a visual form that is different from other options to clearly indicate the current selection status.
[0182] The purpose of this setup is to significantly improve the intuitiveness and operational efficiency of audio editing through a two-way linkage mechanism of real-time timbre replacement and visual feedback. On the one hand, it allows users to directly trigger the timbre algorithm to reconstruct the spectrum of the original audio with a single click, achieving professional-grade sound quality processing while hiding technical complexity. On the other hand, it visually enhances the selected item through preset shapes, instantly confirming the user's operational intent and reducing the risk of misoperation, forming a closed-loop experience of "click-feedback-generation," which optimizes the smoothness of human-computer interaction while ensuring the accuracy of timbre replacement.
[0183] S450: In response to detecting a trigger operation on a custom timbre type, displays the timbre learning page.
[0184] The timbre learning page is an interactive training interface provided for the custom timbre function. It presents various typical audio types and their corresponding timbre descriptions (such as "bright tenor" and "deep cello"), guiding users to understand timbre characteristics through the association between text semantics and audio examples. This transforms the abstract selection of timbre parameters into an intuitive operation based on cognitive matching, ultimately generating a custom timbre option library that matches the user's semantic description. The timbre learning page includes at least two audio types and corresponding timbre learning text for each type. The timbre learning text refers to standardized semantic tags or explanatory text used on the timbre learning page to describe the timbre characteristics of a specific audio type.
[0185] In this embodiment, when a user triggers a custom timbre type, a structured timbre learning page is loaded. This page displays multiple audio types and their corresponding timbre feature description texts (i.e., timbre learning texts) in parallel, enabling users to understand the auditory characteristics of different timbres through semantic descriptions. This transforms the abstract timbre parameter selection into an intuitive operation based on text semantic matching, providing a cognitive foundation for the subsequent generation of custom timbres that conform to the user's intentions.
[0186] S460, In response to detecting an event that satisfies timbre learning, determine the selectable timbre corresponding to the selected audio type based on the audio type displayed on the timbre learning page and the corresponding timbre learning text.
[0187] Among them, selectable timbre refers to the set of timbre parameters that can be directly selected by the user.
[0188] In this embodiment, when it is detected that the user has completed timbre learning (such as confirming the selection or completing the interactive task), the target audio type selected by the user on the timbre learning page and the associated timbre description text are extracted. Through natural language processing and acoustic model matching algorithms, the semantics of the text are transformed into specific timbre parameter configurations. Finally, a set of preset timbre options that both conform to the technical specifications of the audio type and fit the user's text description intent are generated, which are the selectable timbres.
[0189] For example, when a user triggers the voice change function, another display content diagram on the media content creation page can be found here. Figure 11 .like Figure 11 As shown in (a), after the user triggers the "Change Timbre" function on the page, the system will display the media data editing page corresponding to the timbre change function. Figure 11 (b) This page presents several preset tone types, including "Favorites," "My," "Popular," "Commentary," "Cute," and "Foreign Language." If the user further activates the "My" preset tone type, a "+Add" control will expand (used to trigger a custom tone type), at which point the system can display the tone learning page. Figure 11 As shown in (c) or 11(e), the timbre learning page displays two audio types: "Read Aloud" and "A cappella." After the user clicks on the "Read Aloud" audio type, the page will display the corresponding timbre learning text 1; clicking the circular control at the bottom of the page will cause the system to read the text aloud and begin recording the user's personalized audio data. Figure 11 (d) Similarly, if the user clicks the "Acapella" audio type, the corresponding timbre learning text 2 will be displayed; after clicking the circular control, the user can sing the text acapella, and the system will simultaneously start recording personalized audio data. Figure 11 (f)).
[0190] In this way, by providing multiple audio types and corresponding descriptive texts through the timbre learning page, users can express their timbre preferences through intuitive semantic understanding rather than professional parameter adjustments. At the same time, the text descriptions are automatically converted into precise timbre parameters, which not only ensures the professionalism and accuracy of timbre generation, but also achieves a smooth mapping from user cognition to machine execution, effectively balancing ease of operation and customization needs.
[0191] S470, In response to detecting a completion event of editing the first audio data, determine the first media content based on the first audio data that has been edited.
[0192] The technical solution of this disclosure, when the audio editing function and the timbre-changing function are triggered, determines the first audio data as follows: In response to the triggering operation of the timbre-changing function, a media data editing page corresponding to the timbre-changing function is displayed; if the user triggers a preset control corresponding to a preset timbre type, in response to the triggering operation of any preset timbre type, multiple optional timbres associated with the timbre type are displayed; in response to the triggering operation of any optional timbre, the timbre in the second audio data is updated with the timbre corresponding to the triggered optional timbre to obtain the first audio data. If the user triggers a preset control corresponding to a custom timbre type, in response to detecting the triggering operation of the custom timbre type, a timbre learning page is displayed; in response to detecting an event that satisfies timbre learning, based on the audio type displayed on the timbre learning page and the corresponding timbre learning text, the selectable timbre corresponding to the selected audio type is determined. The technical solution provided in this embodiment offers a preset timbre path that provides standardized timbre templates through a hierarchical menu, meeting the needs for quick replacement and general use. The custom timbre path, on the other hand, transforms abstract timbre parameters into intuitive text descriptions through a semantic learning interface, lowering the professional threshold while ensuring customization accuracy. Both paths ensure the accurate implementation of operational intentions through real-time timbre rendering and visual feedback, forming a complete timbre editing solution covering everything from rapid application to refined creation.
[0193] Based on the above embodiments, this embodiment provides a more detailed explanation of the method for determining the first audio data when the triggered audio editing function, melody changing function, and copywriting assistance function are related. For specific implementation details, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. The method of this embodiment may specifically include:
[0194] S510, Display the media content creation page; wherein, the media content creation page includes a list of functions for generating media data, and the list of functions includes at least an audio editing function.
[0195] S520. In response to a trigger operation on any audio editing function, display a media data editing page corresponding to the triggered audio editing function, so as to determine the first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page.
[0196] In this embodiment, when the triggered audio editing function item is related to the melody change function item, the media data editing page corresponding to the triggered melody change function item includes a text content editing area and a melody display area.
[0197] The text content editing area is an interactive interface module within the media data editing page corresponding to the melody-changing function, specifically designed for inputting, modifying, and generating text content. This area is used to edit the second text content to be converted into the first audio data; in other words, its core function is to provide an interactive interface for text input and modification, enabling users to create or adjust a target text (i.e., the second text content). This text will serve as the raw data source for speech synthesis, and through subsequent melody matching and audio generation processes, will ultimately be converted into an audio output result (i.e., the first audio data) with specific melodic characteristics.
[0198] The melody display area is an interactive interface module within the media data editing page corresponding to the melody change function, specifically designed for presenting and selecting melody parameters. This area includes timbre selection and melody selection options. When triggered, the melody selection option displays multiple melody type items, each associated with at least one selectable melody. The timbre selection option is a parameterized control within the melody display area used to configure the timbre characteristics of the synthesized speech. The melody selection option is an interactive control specifically designed to provide melody template selection. Selectable melodies are a set of pre-coded, standardized melody templates displayed to the user after the melody selection option is triggered. For example, each selectable melody is associated with a specific musical style identifier (such as soothing or upbeat) and a set of technical parameters. Users can select these to apply the corresponding melody features to text content, thereby generating synthesized audio data with the desired musicality. It is understood that this embodiment designs a two-layer interactive architecture for the melody display area: the timbre selection item provides a configuration interface for the acoustic features of speech synthesis, while the melody selection item provides structured navigation for melody templates; when the user triggers the melody selection item, a categorized melody type menu will be expanded (such as by style or emotion). Each melody type item is further associated with a set of preset optional melody parameter packages, forming a hierarchical operation path of "type navigation → specific selection", enabling the user to efficiently locate the target melody and simultaneously configure the timbre parameters, thereby achieving precise synthesis control from text to melody audio.
[0199] Based on the above embodiments, optionally, the text content editing area also includes an intelligent writing control associated with the copywriting assistance function. The intelligent writing control is used to generate second text content displayed in the text content editing area after being triggered.
[0200] Among them, the intelligent writing control refers to the interactive function entry point integrated into the text content editing area and dedicated to assisting text generation. Its essence is to connect with the AI text generation engine through natural language processing technology, allowing users to automatically perform semantic polishing, style conversion or generate new content based on related material data by triggering operations, thereby dynamically outputting optimized text (i.e., second text content) that meets the creative needs, and providing a high-quality text input source for subsequent audio synthesis.
[0201] The second text content is derived from the text content edited in the text content editing area, after polishing, and / or generated after analyzing the source data. This can be understood as the intelligent writing control generating the target text through two paths: one is to perform grammatical optimization, sentence restructuring, or stylistic processing (polishing) on the user-inputted original text; the other is to directly parse media materials (such as visual information in images / videos) and automatically generate descriptive, narrative, or tagged text (source data analysis). The final output is standardized text data that can stand alone as new text or be integrated with the user's original content, which is the second text content.
[0202] In this embodiment, the intelligent writing control integrated into the text content editing area serves as the specific interactive entry point for the copywriting assistance function. When triggered by the user, it calls the natural language processing engine to perform two types of text generation operations—either semantic polishing and structural optimization of the original text that the user has entered in the editing area, or direct analysis of related media material data and automatic generation of descriptive text. Finally, the processing result (i.e., the second text content) is fed back to the text editing area in real time, forming a seamless collaborative process between automated text creation and manual editing.
[0203] For example, a diagram illustrating the content displayed on the media content creation page when a user triggers the melody change function can be found here. Figure 12 .like Figure 12 As shown in (a), after the user triggers the "Change Melody" function on the page, the media data editing page corresponding to the "Change Melody" function will be displayed, such as... Figure 12 (b) or Figure 12 (c) shows the page. This page presents a text editing area and a melody display area. The text editing area features a "smart writing control," which automatically generates second text content displayed in the text editing area when triggered by the user. The melody display area includes timbre selection options and melody selection options, such as... Figure 12 As shown in (b), when the user triggers the timbre selection option, multiple selectable timbres can be displayed in the melody display area; such as Figure 12 As shown in (c), when the user triggers the melody selection option, multiple selectable melodies can be displayed in the melody display area.
[0204] Optionally, based on the triggering operation on any optional melody in the media data editing page, the third audio data is determined based on the triggered selected optional melody and the second text content.
[0205] In this embodiment, when a user triggers an optional melody on the media data editing page, the preset musical parameters of that melody can be extracted. Simultaneously, combined with the second text content (user-inputted or intelligently generated text) in the text content editing area, the text content is rendered according to the acoustic characteristics of the target melody through a speech synthesis engine. This ultimately generates synthesized audio data that possesses both semantic integrity and melodic regularity—the third audio data. Upon obtaining the third audio data, the editing completion event of the first audio data is triggered. In this way, through a real-time linkage synthesis mechanism between melody and text, the efficiency and professionalism of musical audio creation are achieved. Users can trigger the automatic and precise fusion of text content and target melody parameters with a single melody selection operation. This avoids the cumbersome arrangement process in traditional music production while ensuring the melodic regularity and semantic integrity of the synthesized audio, significantly lowering the technical threshold for musical audio creation, and guaranteeing the professional sound quality and artistic expression of the output.
[0206] S530, In response to detecting a completion event of editing the first audio data, determine the first media content based on the first audio data that has been edited.
[0207] The technical solution of this disclosure embodiment, when the triggered audio editing function item is related to the melody replacement function item, the media data editing page corresponding to the triggered melody replacement function item includes a text content editing area and a melody display area. Through the dual-area collaborative editing mechanism of text and melody, the efficiency and accuracy of melody replacement operation are achieved. The text content editing area provides users with an intuitive interface for lyric / text input and modification, ensuring the freedom of content creation. The melody display area simplifies the complex selection of music parameters into a visual operation through hierarchical melody type navigation and selectable melody calls. At the same time, it integrates timbre configuration to achieve integrated adjustment of acoustic characteristics, ultimately forming a linear workflow of "text creation - melody selection - timbre adjustment", which significantly reduces the technical threshold of music editing and improves creation efficiency. The text content editing area also includes an intelligent writing control associated with the copywriting assistance function. When triggered, the intelligent writing control generates a second text content displayed in the text content editing area. In this way, on the one hand, the text polishing function automatically optimizes the accuracy and fluency of the user's input content, and on the other hand, it generates context-related adapted text through material data analysis. This not only lowers the threshold for users to create manually, but also ensures a high degree of synergy between the text content and the melody and timbre parameters, ultimately forming an integrated creative closed loop of "intelligent text generation - melody matching - audio synthesis".
[0208] Figure 13 This is a flowchart illustrating another media data editing method provided in this embodiment. Based on the above embodiments, this embodiment refines the method for determining the first audio data when the triggered audio editing function is related to the text writing assistance function. For specific implementation details, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 13 As shown, the method in this embodiment may specifically include:
[0209] S610, Display the media content creation page; wherein, the media content creation page includes a list of functions for generating media data, and the list of functions includes at least an audio editing function.
[0210] In this embodiment, when the audio editing function is triggered in relation to the copywriting function, the media data editing page includes a text content editing area, a copy length setting option, and a copy style type setting option.
[0211] The text content editing area is an interactive interface module specifically designed for inputting, modifying, and displaying text content. The text length setting option, when triggered, determines the word count range of the generated second text content. This can be understood as an interactive control within the media data editing page of the copywriting assistance function, specifically used to control the length of the generated text. The text style type setting option, when triggered, determines the text style type of the second text content. This can be understood as a classification parameter system used to define the expressive characteristics of the generated text. It can use preset style tags (such as formal, humorous, and technological) to correspond to different language model configuration parameters, allowing users to select a specific style type to guide the generation of text content with corresponding lexical features, sentence structures, and emotional tones, thereby achieving precise matching between the copy output and the target scenario or audience preferences.
[0212] In this embodiment, when the triggered audio editing function is related to copywriting assistance, the three core components—text content editing area, copy length setting option, and copy style type setting option—can be dynamically loaded in the media data editing page.
[0213] Based on this, the implementation method of the first audio data is determined as described in S620-S340, specifically including:
[0214] S620. Based on the text length setting option and text style type setting option triggered in the media data editing page, a text generation panel is displayed in the media data editing page to display a preset number of selectable texts.
[0215] The text generation panel is a dedicated interactive interface module that is dynamically loaded after the copywriting assistance function is triggered. The selectable text refers to the set of candidate texts generated in the text generation panel based on the user-set text length and style parameters, which the user can choose from. The selectable text consists of multiple versions of text output that meet preset constraints. Each selectable text carries a complete semantic expression and adapts to the target style characteristics. Users can select it to determine the final text content used, providing a standardized input source for subsequent audio conversion.
[0216] In this embodiment, when the user sets the text length through the text length setting option and the style type parameter through the text style type setting option in the media data editing page, the text generation panel can be dynamically loaded. Based on the preset parameter combination (word range, style tag), the natural language processing engine is driven to generate multiple text variations that meet the requirements. These variations are then displayed in the panel as a preset number of candidate texts (selectable texts) in the form of a structured list, forming a multi-scheme output mode guided by parameters.
[0217] For example, when a user sets the copy length to "50-100 words" and the style type to "tech style" on the copywriting page and confirms the parameters, a text generation panel can pop up on the current page, displaying three candidate texts generated based on these parameters (e.g., "Artificial intelligence is reshaping future life...", "Technological innovation drives digital transformation..."). These optional texts strictly meet the word count requirements and contain professional technical terms and concise sentence structure features.
[0218] S630. Determine the first audio data based on the selected optional text.
[0219] In this embodiment, the candidate text selected by the user from the text generation panel (i.e., selectable text) can be used as the source data for speech synthesis. By calling the audio generation engine and combining preset or user-configured acoustic parameters (such as timbre, speech rate, and intonation), the text content is converted into a digital audio output with specific acoustic characteristics. This audio data is the final generated first audio data. In this way, the efficiency and quality of audio text creation are significantly improved through parameterized guidance and a multi-scheme output mechanism. Users can obtain multiple text variations that meet the requirements by setting length and style parameters, which avoids the trial and error costs of blind generation and ensures text quality through multi-option comparison. At the same time, the text selection and audio generation are seamlessly connected, realizing an automated flow from text creation to speech synthesis, reducing operational complexity while ensuring a high degree of consistency between the final audio content and the creative intent.
[0220] For example, when a user triggers the copywriting assistance feature, a screenshot of the content displayed on the media content creation page can be found here.Figure 14 .like Figure 14 As shown in (a), after the user triggers the "Smart Writing Control" on the page, the displayed media data editing page is as follows: Figure 14 As shown in (b); the upper part of the media data editing page is the text content editing area; the middle "less than 20 characters", "20-50 characters", and "50-100 characters" controls are text length setting options; the lower "Style 1", "Style 2", and "Style 3" controls are text style type setting options. When the user selects a text length setting option and a text style type setting option, and clicks the "Generate Now" control, a text generation panel with multiple selectable text options can be displayed, such as... Figure 14 As shown in (c).
[0221] Based on the above embodiments, optionally, if the text generation panel includes optional text and new text description information is detected, the optional text is regenerated based on the optional text and the new text description information; the optional text is then displayed as a new conversation message in the text generation panel.
[0222] The newly added text description information refers to supplementary text instructions or modification suggestions entered by the user during the interaction with the text generation panel. The newly added conversation message refers to the optional text regenerated based on the user's newly added text description information, which is dynamically added to the interaction history of the text generation panel in the form of a conversation flow.
[0223] In this embodiment, when there are already generated optional texts in the text generation panel and the user inputs new descriptive information, the semantic basis of the original optional texts and the detailed requirements of the new descriptions are integrated to drive the intelligent text engine to regenerate optimized new candidate texts. These texts are then appended to the text generation panel as conversational messages. The purpose of this setup is to achieve deep collaborative creation between the intelligent text engine and the user through iterative text generation and conversational interface presentation. On the one hand, it allows users to continuously refine their needs by adding descriptive information, making the text output infinitely close to the creative intent; on the other hand, the cumulative display of conversational messages forms a traceable modification trajectory, which not only preserves contextual relevance but also enhances process controllability. Ultimately, it significantly improves the accuracy of text generation and user satisfaction while reducing trial and error costs.
[0224] S640, In response to detecting a completion event of editing the first audio data, determine the first media content based on the first audio data that has been edited.
[0225] The technical solution of this disclosure embodiment, when the audio editing function and the copywriting assistance function are triggered, includes a text content editing area, a copy length setting option, and a copy style type setting option in the media data editing page. The copy length setting option provides quantitative control over the text length, and the copy style type setting option provides classification and selection of expressive features. Together, they constitute a parameterized guidance system for the text generation process, ensuring that the output second text content simultaneously meets the dual constraints of word count specifications and style adaptation. When generating the first audio data, a text generation panel is displayed on the media data editing page based on the text length and text style settings options triggered in the media data editing page. The text generation panel displays a preset number of optional texts. Based on the selected optional texts, the first audio data is determined. Users can obtain multiple text variations that meet the requirements by setting the length and style parameters. This avoids the trial and error costs caused by blind generation and ensures text quality through comparison of multiple options. At the same time, the text selection and audio generation are seamlessly connected, realizing an automated flow from text creation to speech synthesis. This reduces the complexity of operation while ensuring a high degree of consistency between the final audio content and the creative intent.
[0226] Figure 15 This is a schematic diagram of the structure of a media data editing device provided in an embodiment of the present disclosure, as shown below. Figure 15 As shown, the device includes: a creation page display module 710, a function item editing module 720, and a media content generation module 730.
[0227] The creation page display module 710 is used to display the media content creation page; wherein the media content creation page includes a function list for generating media data, and the function list includes at least an audio editing function item;
[0228] The function item editing module 720 is used to respond to a trigger operation on any audio editing function item by displaying a media data editing page corresponding to the triggered audio editing function item, so as to determine the first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page;
[0229] The media content generation module 730 is used to determine the first media content based on the first audio data after it is edited, in response to the detection of an editing completion event of the first audio data.
[0230] The technical solution of this disclosure embodiment, in response to detecting a user's trigger operation on generating media data, displays a media content creation page. The media content creation page includes a function list for generating media data, with at least an audio editing function item. Then, in response to a trigger operation on any audio editing function item, a media data editing page corresponding to the triggered audio editing function item is displayed. Based on the editing operation on the media data editing page, first audio data corresponding to the triggered audio editing function is determined. Finally, in response to detecting an editing completion event of the first audio data, first media content is determined based on the edited first audio data. This disclosure embodiment, by deeply integrating audio editing functions into the media content creation page, effectively solves the problem of interrupted creation processes caused by fragmented functional modules. When a user triggers any audio editing function item, the corresponding media data editing page can be presented in a unified creation environment, achieving seamless connection between audio processing and the main creation process. This significantly reduces context switching costs, shortens operation paths, and enables audio data to be synchronized to the main project in real time. This not only greatly improves creation efficiency but also ensures precise collaborative editing of audio elements and visual materials, ultimately providing users with an integrated and smooth creation experience, significantly enhancing the quality of multimedia content output and user satisfaction.
[0231] Based on any optional technical solution in the embodiments of this disclosure, the media content creation page includes at least one audio track for displaying audio data, and the at least one audio track is in an empty track state before audio data is configured.
[0232] Based on any optional technical solution in the embodiments of this disclosure, the media content creation page further includes at least one material track for displaying material data, the material data including image data and / or video data, and the at least one material track is in an empty track state when no material data is configured.
[0233] Based on any optional technical solution in the embodiments of this disclosure, the media content creation page further includes a material editing function item, which is used to adjust the material data displayed in at least one material track and the audio data associated with the material data after being triggered.
[0234] Based on any optional technical solution in the embodiments of this disclosure, the audio editing function includes one or more of the following: rewriting function to adjust the text content corresponding to the audio data, changing the timbre of the audio data, changing the melody of the audio data, performing voice separation on the audio data, adjusting the playback speed of the audio data, human voice reading function, copywriting assistance function, sound effect addition function, and sound source addition function.
[0235] Based on any optional technical solution in the embodiments of this disclosure, the triggered audio editing function item is related to the rewriting function item. At least one audio track of the media content creation page displays second audio data. The function item editing module 720 is also used to respond to the triggering operation of the rewriting function item and display a media data editing page corresponding to the rewriting function item. The media data editing page displays at least one first text content after audio recognition of the second audio data and the associated audio playback time. The at least one first text content is arranged in order according to the audio playback time.
[0236] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to adjust the first text content displayed on the media data editing page in response to any editing operation on the first text content; wherein, the editing operation includes one or more of the following: word modification operation, first text content deletion operation, and segment duration adjustment operation; in response to the event that the first text content editing is completed, the first audio data is generated based on the adjusted first text content, and the user is redirected back to the media content creation page.
[0237] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to, when a trigger operation on the rewriting function item is detected, and when it is determined that the audio track includes second audio data, send the second audio data to the server, so that when the voice detection module in the server determines that the second audio data meets preset conditions, the text information of the second audio data is extracted based on the text extraction module; the text information and the second audio data are time-aligned based on the text alignment module integrated in the server, and feedback is given to obtain at least one first text content displayed in the media data editing page.
[0238] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further used to generate first audio data with the same melody as the second audio data but different text content based on the audio generation service integrated in the server and the adjusted first text content and the second audio data.
[0239] Based on any optional technical solution in the embodiments of this disclosure, the triggered audio editing function item is related to the timbre-changing function item, and at least one audio track of the media content creation page displays second audio data. The function item editing module 720 is also used to respond to the triggering operation of the timbre-changing function item and display the media data editing page corresponding to the timbre-changing function item; wherein, the media data editing page includes multiple preset timbre types and custom timbre types, and the preset timbre type item is associated with multiple selectable timbres.
[0240] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to, in response to a trigger operation on any preset timbre type, display multiple optional timbres associated with the timbre type item; in response to a trigger operation on any optional timbre, update the timbre in the second audio data with the timbre corresponding to the triggered optional timbre to obtain the first audio data; wherein the triggered timbre option is presented in a preset form.
[0241] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further used to determine the audio type corresponding to the second audio data; and to retrieve the preset timbre type corresponding to the audio type.
[0242] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to display a timbre learning page in response to detecting a trigger operation on the custom timbre type; wherein, the timbre learning page includes at least two audio types and timbre learning text corresponding to each audio type; in response to detecting an event that satisfies timbre learning, the selectable timbre corresponding to the selected audio type is determined based on the audio type and the corresponding timbre learning text displayed in the timbre learning page.
[0243] Based on any optional technical solution in the embodiments of this disclosure, the triggered audio editing function item is related to the melody changing function item. The media data editing page corresponding to the melody changing function item after it is triggered includes a text content editing area and a melody display area. The text content editing area is used to edit the second text content to be converted into the first audio data. The melody display area includes a timbre selection item and a melody selection item. The melody selection item is used to display multiple melody type items after it is triggered. Each melody selection item is associated with at least one selectable melody.
[0244] Based on any optional technical solution in the embodiments of this disclosure, the text content editing area further includes an intelligent writing control associated with the copywriting assistance function. The intelligent writing control is used to generate second text content displayed in the text content editing area after being triggered. The second text content is text content obtained after polishing the text content edited in the text content editing area, and / or text content generated after analyzing the material data.
[0245] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to determine third audio data based on the trigger operation of any optional melody in the media data editing page, based on the triggered selected optional melody and the second text content; and in response to the optional melody confirmation event, use the determined third audio data as the first audio data.
[0246] Based on any optional technical solution in the embodiments of this disclosure, the triggered audio editing function is related to the copywriting assistance function. The media data editing page includes a text content editing area, a copy length setting option, and a copy style type setting option. The copy length setting option is used to determine the word count range of the generated second text content after being triggered, and the copy style type setting option is used to determine the copy style type of the second text content after being triggered.
[0247] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to display a text generation panel in the media data editing page based on the text length setting option and text style type setting option triggered in the media data editing page, so as to display a preset number of optional texts in the text generation panel; and determine the first audio data based on the selected optional texts.
[0248] Based on any optional technical solution in the embodiments of this disclosure, the function item editing module 720 is further configured to, when the text generation panel includes optional text and new text description information is detected, regenerate optional text based on the optional text and the new text description information; and display the optional text as a new conversation message in the text generation panel.
[0249] Based on any optional technical solution in the embodiments of this disclosure, the media data editing device further includes: a material editing module;
[0250] The material editing module is used to display a material editing panel in response to a trigger operation of the material editing function item, so as to determine the first material data based on the editing operation in the material editing panel; wherein, the material editing panel includes a smart screen type option and a digital human type option, the screen type option is used to determine the screen type of the generated first material data after being triggered, and the digital human type option is used to determine the display object after being triggered.
[0251] Based on any optional technical solution in the embodiments of this disclosure, the material editing module is specifically used to respond to the trigger operation of any picture style type associated with the picture type option, and generate first material data corresponding to the selected picture style type according to the sentence segmentation result corresponding to the first audio data.
[0252] Based on any optional technical solution in the embodiments of this disclosure, the material editing module is specifically used to respond to the triggering operation of any digital human associated with the digital human type option, and drive the selected digital human according to the first audio data to obtain the first material data.
[0253] Based on any optional technical solution in the embodiments of this disclosure, the material editing module is further configured to determine the first media content based on the edited first audio data and the first material data.
[0254] Based on any optional technical solution in the embodiments of this disclosure, the media data editing device further includes: a data transmission module;
[0255] The data sending module is used to send the first media data to the first platform in response to the event of publishing the first media content; wherein the first platform is different from the platform that edits and generates the first media content.
[0256] The media data editing apparatus provided in this disclosure can execute the media data editing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the media data editing method.
[0257] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0258] The following is for reference. Figure 16 This illustration shows a structural diagram of an electronic device (e.g., a terminal device or a server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 16 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0259] like Figure 16As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0260] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 16 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0261] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0262] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0263] The electronic device provided in this disclosure and the media data editing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this disclosure can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0264] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the media data editing method provided in the above embodiments.
[0265] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0266] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0267] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0268] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: display a media content creation page; wherein the media content creation page includes a function list for generating media data, the function list including at least an audio editing function item; in response to a triggering operation of any audio editing function item, display a media data editing page corresponding to the triggered audio editing function item, to determine first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page; and in response to detecting an editing completion event of the first audio data, determine first media content based on the edited first audio data.
[0269] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0270] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0271] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not necessarily limiting in certain circumstances; for example, a page presentation module can also be described as a "module for presenting an information interaction page".
[0272] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0273] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0274] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0275] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0276] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A media data editing method, characterized in that, include: The media content creation page is displayed; wherein the media content creation page includes a list of functions for generating media data, and the list of functions includes at least an audio editing function item; In response to a trigger operation on any audio editing function, a media data editing page corresponding to the triggered audio editing function is displayed, so as to determine the first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page; In response to the detection of a completion event for editing the first audio data, the first media content is determined based on the edited first audio data.
2. The method according to claim 1, characterized in that, The media content creation page includes at least one audio track for displaying audio data, and the at least one audio track is in an empty track state before audio data is configured.
3. The method according to claim 1 or 2, characterized in that, The media content creation page also includes at least one media track for displaying media data, which includes image data and / or video data. The at least one media track is in an empty track state when no media data is configured.
4. The method according to claim 1, characterized in that, The media content creation page also includes a material editing function, which, when triggered, is used to adjust the material data displayed in at least one material track and the audio data associated with the material data.
5. The method according to claim 1, characterized in that, The audio editing functions include one or more of the following: rewriting function to adjust the text content corresponding to the audio data, changing the timbre of the audio data, changing the melody of the audio data, performing voice separation on the audio data, adjusting the playback speed of the audio data, voice reading function, copywriting assistance function, sound effect addition function, and sound source addition function.
6. The method according to claim 5, characterized in that, The triggered audio editing function item is related to the rewrite function item. At least one audio track on the media content creation page displays second audio data. The step of responding to a trigger operation on any audio editing function item by displaying a media data editing page corresponding to the triggered audio editing function item includes: In response to the triggering operation of the rewrite function item, the media data editing page corresponding to the rewrite function item is displayed; The media data editing page displays at least one first text content after audio recognition of the second audio data and the associated audio playback time, wherein the at least one first text content is arranged in order according to the audio playback time.
7. The method according to claim 6, characterized in that, The step of determining the first audio data corresponding to the triggered audio editing function based on the editing operation on the media data editing page includes: In response to any editing operation on the first text content, the first text content displayed on the media data editing page is adjusted; wherein, the editing operation includes one or more of the following: word modification operation, first text content deletion operation, and segment duration adjustment operation; In response to the event that the first text content editing is completed, the first audio data is generated based on the adjusted first text content, and the user is redirected back to the media content creation page.
8. The method according to claim 6, characterized in that, The step of displaying a media data editing page corresponding to the rewrite function item in response to a trigger operation includes: When a trigger operation on the rewrite function is detected, and it is determined that the audio track includes second audio data, the second audio data is sent to the server so that when the voice detection module in the server determines that the second audio data meets the preset conditions, the text extraction module extracts the text information of the second audio data. The text alignment module integrated into the server performs time alignment on the text information and the second audio data, and provides feedback to obtain at least one first text content displayed on the media data editing page.
9. The method according to claim 7, characterized in that, The process of generating the first audio data based on the adjusted first text content includes: Based on the audio generation service integrated in the server, the adjusted first text content and the second audio data are used to generate first audio data with the same melody as the second audio data but different text content.
10. The method according to any one of claims 5-7, characterized in that, The triggered audio editing function is related to the timbre-changing function. At least one audio track on the media content creation page displays second audio data. In response to a trigger operation on any audio editing function, the display of the media data editing page corresponding to the triggered audio editing function includes: In response to the triggering of the tone-changing function, the media data editing page corresponding to the tone-changing function is displayed; The media data editing page includes multiple preset timbre types and custom timbre types, and the preset timbre type item is associated with multiple selectable timbres.
11. The method according to claim 10, characterized in that, The step of determining the first audio data corresponding to the triggered audio editing function based on the editing operation on the media data editing page includes: In response to a trigger operation on any preset timbre type, multiple selectable timbres associated with the timbre type item are displayed; In response to a trigger operation on any selectable timbre, the timbre in the second audio data is updated with the timbre corresponding to the triggered selectable timbre, and the first audio data is obtained; Among them, the triggered timbre options are presented in a preset form.
12. The method according to claim 10, characterized in that, Before displaying the media data editing page corresponding to the color-changing function, the method further includes: Determine the audio type corresponding to the second audio data; Retrieve the preset timbre type corresponding to the audio type.
13. The method according to claim 10, characterized in that, The method further includes: In response to the detection of a trigger operation on the custom timbre type, a timbre learning page is displayed; wherein, the timbre learning page includes at least two audio types and timbre learning text corresponding to each audio type; In response to detecting an event that satisfies timbre learning, the selectable timbre corresponding to the selected audio type is determined based on the audio type displayed on the timbre learning page and the corresponding timbre learning text.
14. The method according to claim 5, characterized in that, The triggered audio editing function item is related to the melody change function item. The media data editing page corresponding to the melody change function item after it is triggered includes a text content editing area and a melody display area. The text content editing area is used to edit the second text content to be converted into the first audio data. The melody display area includes a timbre selection item and a melody selection item. The melody selection item is used to display multiple melody type items after being triggered. Each melody selection item is associated with at least one selectable melody.
15. The method according to claim 14, characterized in that, The text content editing area also includes an intelligent writing control associated with the copywriting assistance function. The intelligent writing control is used to generate second text content to be displayed in the text content editing area after being triggered. The second text content is the text content obtained after polishing the text content edited in the text content editing area, and / or the text content generated after analyzing the material data.
16. The method according to claim 14, characterized in that, The step of determining the first audio data corresponding to the triggered audio editing function based on the editing operation on the media data editing page includes: Based on the triggering operation of any selectable melody in the media data editing page, and based on the selected selectable melody and the second text content, the third audio data is determined; In response to the optional melody confirmation event, the determined third audio data is used as the first audio data.
17. The method according to claim 5 or 15, characterized in that, The triggered audio editing function is related to the copywriting assistance function. The media data editing page includes a text content editing area, a copy length setting option, and a copy style type setting option. The text length setting option is used to determine the word count range of the generated second text content when triggered, and the text style type setting option is used to determine the text style type of the second text content when triggered.
18. The method according to claim 17, characterized in that, The step of determining the first audio data corresponding to the triggered audio editing function based on the editing operation on the media data editing page includes: Based on the text length setting option and text style type setting option triggered on the media data editing page, a text generation panel is displayed on the media data editing page to display a preset number of selectable texts. The first audio data is determined based on the selected optional text.
19. The method according to claim 18, characterized in that, The method further includes: If the text generation panel includes optional text and new text description information is detected, the optional text is regenerated based on the optional text and the new text description information. The optional text is displayed as a new session message in the text generation panel.
20. The method according to claim 4, characterized in that, The media content creation page also includes material editing functions, and the method further includes: In response to a trigger operation on the material editing function item, a material editing panel is displayed to determine the first material data based on the editing operation in the material editing panel; The material editing panel includes a smart screen type option and a digital human type option. The smart screen type option is used to determine the screen type of the first material data generated after being triggered, and the digital human type option is used to determine the display object after being triggered.
21. The method according to claim 20, characterized in that, The determination of the first material data based on the editing operations in the material editing panel includes: In response to a trigger operation on any of the picture style types associated with the picture type option, first material data corresponding to the selected picture style type is generated based on the sentence segmentation result corresponding to the first audio data.
22. The method according to claim 20, characterized in that, The determination of the first material data based on the editing operations in the material editing panel includes: In response to a trigger operation on any digital human associated with the digital human type option, the selected digital human is driven according to the first audio data to obtain the first material data.
23. The method according to claim 20, characterized in that, The step of responding to the detection of a completion event for editing the first audio data, and determining the first media content based on the completed editing of the first audio data, includes: Based on the edited first audio data and the first material data, the first media content is determined.
24. The method according to claim 1, characterized in that, The method further includes: In response to the event of publishing the first media content, the first media data is sent to the first platform; The first platform is distinct from the platform used to edit and generate the first media content.
25. A media data editing device, characterized in that, include: The creation page display module is used to display the media content creation page; wherein, the media content creation page includes a function list for generating media data, and the function list includes at least an audio editing function item; The function item editing module is used to respond to the triggering operation of any audio editing function item, display the media data editing page corresponding to the triggered audio editing function item, and determine the first audio data corresponding to the triggered audio editing function based on the editing operation in the media data editing page; The media content generation module is used to respond to the detection of a completion event for editing the first audio data and determine the first media content based on the edited first audio data.
26. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the media data editing method as described in any one of claims 1-24.
27. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the media data editing method as described in any one of claims 1-24.
28. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the media data editing method as described in any one of claims 1-24.
Citation Information
Cited By
Real-time digital human video generation method and device, electronic equipment and storage medium
CN121309905A