Music data processing method and device, electronic equipment and storage medium

By displaying and analyzing the audio quality and content dimensions of music data in the music interface, combining cross-modal understanding and interactive evaluation, the problems of low quality and low distribution efficiency of music works are solved, and the quality of music works and the improvement of user experience are achieved.

CN120299476APending Publication Date: 2025-07-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510421614.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The successful distribution efficiency of music works is low and the user experience is poor, mainly due to the low yield rate problem caused by low threshold creation.

Method used

It provides a music data processing method, by displaying music data on the interface and analyzing it based on the audio quality and music content dimensions, the analysis results are obtained, and displayed on the interface for user reference, supporting the editing and quality inspection of music data, including the analysis and interactive evaluation of cross-modal comprehension capabilities.

Benefits of technology

It improves the quality and production efficiency of music works, improves user experience, and helps users improve the quality and distribution efficiency of music works through real-time feedback and editing effects guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299476A_ABST
    Figure CN120299476A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a music data processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining first music data, and displaying the first music data in a first interface; based on analysis of the first music data in a first preset dimension, obtaining a first analysis result; wherein the first preset dimension comprises an audio quality dimension and / or a music content dimension; and displaying the first analysis result on the first interface. And the quality and the production efficiency of the music works can be improved, so that the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to a method, apparatus, electronic device, and storage medium for processing music data. Background Art

[0002] With the emergence of music platforms, users can directly complete operations such as music production, uploading, and distribution based on the platforms. The reduction of the operation threshold has further stimulated the creative enthusiasm and made the music industry flourish. However, the low threshold of creation is also likely to result in a low yield rate, leading to a relatively low efficiency of successful music work distribution and a poor user experience. Summary of the Invention

[0003] Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for processing music data, which helps to improve the quality and production efficiency of music works and enhance the user experience.

[0004] In a first aspect, embodiments of the present disclosure provide a method for processing music data, including:

[0005] Obtaining first music data and displaying the first music data on a first interface;

[0006] Based on the analysis of the first music data in a first preset dimension, obtaining a first analysis result; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension;

[0007] Displaying the first analysis result on the first interface.

[0008] In a second aspect, embodiments of the present disclosure further provide a device for processing music data, including:

[0009] A display module, configured to obtain first music data and display the first music data on a first interface;

[0010] An analysis module, configured to obtain a first analysis result based on the analysis of the first music data in a first preset dimension; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension;

[0011] The display module is further configured to display the first analysis result on the first interface.

[0012] In a third aspect, embodiments of the present disclosure further provide an electronic device, which includes:

[0013] One or more processors;

[0014] A storage device, configured to store one or more programs,

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for processing music data according to any one of the embodiments of the present disclosure.

[0016] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for processing music data according to any one of the embodiments of the present disclosure when executed by a computer processor.

[0017] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, characterized in that the computer program product includes a computer program, and the computer program implements the method for processing music data according to any one of the embodiments of the present disclosure when executed by a processor.

[0018] In the technical solution of the embodiment of the present disclosure, first music data can be obtained and the first music data is displayed on a first interface; based on the analysis of the first music data in a first preset dimension, a first analysis result is obtained; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension; the first analysis result is displayed on the first interface. By displaying the analysis result of the first music data in the first preset dimension on the interface, it can provide a reference for users, help users improve the quality and production efficiency of music works, and thus improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0020] Figure 1 It is a schematic flowchart of a method for processing music data provided by an embodiment of the present disclosure;

[0021] Figure 2 It is a schematic diagram of a first interface in a method for processing music data provided by an embodiment of the present disclosure;

[0022] Figure 3 It is a schematic diagram of a second interface in a method for processing music data provided by an embodiment of the present disclosure;

[0023] Figure 4 It is a schematic diagram of a second interface in a method for processing music data provided by an embodiment of the present disclosure;

[0024] Figure 5 It is a schematic diagram of a third interface in a method for processing music data provided by an embodiment of the present disclosure;

[0025] Figure 6 A data flow schematic block diagram of a method for processing music data provided by an embodiment of the present disclosure;

[0026] Figure 7 A structural schematic diagram of a device for processing music data provided by an embodiment of the present disclosure;

[0027] Figure 8 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the protection scope of the present disclosure.

[0029] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0030] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0031] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0035] Figure 1 The figure is a schematic flowchart of a method for processing music data provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of quality inspection and analysis of music data, for example, applicable to the situation of quality inspection and analysis of music data in music application software. This method can be executed by a music data processing device, which can be implemented in the form of software and / or hardware. The device can be integrated into the application software and installed in an electronic device along with the application software, such as installed in computer devices such as mobile phones and computers.

[0036] As Figure 1 shown, the method for processing music data provided in this embodiment may include:

[0037] S110. Obtain first music data and display the first music data on a first interface.

[0038] In the embodiments of the present disclosure, interfaces with different prefixes such as "first" and "second" can be considered as user interfaces provided by the music data processing device. Among them, the first interface can be used to display music data.

[0039] Among them, obtaining the first music data may include but is not limited to at least one of the following: obtaining the first music data from a preset storage space in response to a rendering instruction of the first interface; requesting the first music data from a server in response to a rendering instruction of the first interface; importing the first music data through an import component in the first interface.

[0040] In addition, in some optional implementation manners, obtaining the first music data may include: receiving a prompt word through the first interface; where the prompt word includes lyrics and music tags; generating the first music data according to the prompt word in response to the triggering of a music generation component in the first interface; and displaying the first music data in the first interface.

[0041] Among them, the lyrics received through the first interface may include: receiving the lyrics input by the user in response to the triggering of the lyrics input component in the first interface; and / or, generating lyrics based on Artificial Intelligence Generated Content (AIGC) in response to the triggering of the lyrics generation component in the first interface; among them, the lyrics generated based on AIGC can be displayed in the lyrics input component and can be adjusted based on the editing operations input by the user. Among them, the music tags may include, but are not limited to, genre tags, theme tags, custom tags, etc. Among them, music tags can be received in response to the triggering of the tag input component in the first interface.

[0042] Among them, a music generation component may also be deployed in the first interface. In response to the triggering of the music generation component, prompt words can be generated according to the lyrics and music tags; the prompt words can be input into the AIGC model to output the first music data corresponding to the prompt words through the AIGC model; among them, the number of the generated first music data can include at least one. Among them, the AIGC model used to generate the first music data and the AIGC model used to generate the lyrics can be the same or different. Among them, the generated first music data can be displayed in the first interface.

[0043] Exemplarily, Figure 2 is a schematic diagram of the first interface in a method for processing music data provided by an embodiment of the present disclosure. Refer to Figure 2 , a tag input component 210, a lyrics input component 220, and a music generation component 230 may be deployed in the first interface.

[0044] Among them, the tag input component 210 may include a selection component 211 for tags in dimensions such as genre and theme, such as Figure 2 the selection component 211 in may include, for example, a button component. In addition, the tag input component 210 may further include an overview component 212; among them, the overview component 212 can be used to generate and display an overview according to the triggering of the selection component 211 and / or according to the input text, and can edit the overview according to the input editing operations, such as deleting music tags, etc. Such as Figure 2 the lyrics input component 220 in may include, for example, a text box input component for lyrics. Figure 2 In, in response to the triggering of the music generation component 230, multiple first music data 240 (such as Figure 2 music 1 - music N in) can be generated, and each first music data 240 can be displayed in the first interface.

[0045] In these alternative implementation manners, the automatic generation of the first music data can be achieved through AIGC, which can further accelerate the production efficiency of music works to improve the user experience.

[0046] S120. Obtain a first analysis result based on the analysis of the first music data in a first preset dimension; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension.

[0047] Among them, the audio quality dimension may include, but is not limited to, an audio parameter dimension, an abnormal audio dimension, and a Mean Opinion Score (MOS) dimension. Among them, the audio parameter dimension may include, but is not limited to, the following sub-dimensions: loudness curve, sampling frequency, encoding method, channel, bit rate, Root Mean Square (RMS) curve, frequency response curve, Clipping curve, and stutter rate, etc. Among them, the abnormal audio dimension may include, but is not limited to, the following sub-dimensions: popping sound, burst sound (which can be called POP sound), current sound, pseudo-stereo, noise, and highly similar audio, etc. Among them, the audio quality MOS dimension may include, but is not limited to, the following sub-dimensions: Difference MOS (DNS-MOS), reference-free MOS, Short-Time Objective Intelligibility (STOI), and Perceptual Evaluation of Speech Quality (PESQ), etc. Among them, the music content dimension may include, but is not limited to, dimensions such as musical score, rhythm, mode, tonality, chord, pitch, melody, tempo, beat, mixing, structural paragraph, and style genre, etc.

[0048] In the embodiments of the present disclosure, the first music data can be analyzed in terms of the audio quality dimension and / or the music content dimension through existing tools and / or self-developed tools, and the first analysis result can be obtained based on this analysis. In addition, in some implementation manners, the analysis result of the first music data in the first preset dimension can also be quantized and mapped through a preset mapping rule to obtain a quantized analysis result, and use it as the first analysis result. In some implementation manners, the first analysis result can be, for example, a score, etc.

[0049] S130. Display the first analysis result on the first interface.

[0050] In the embodiments of the present disclosure, in response to the first music data being displayed on the first interface, the first music data can be automatically analyzed in the first preset dimension. And the first analysis results corresponding to each dimension and / or sub-dimension in the first preset dimension can be displayed on the first interface. Further, the total analysis result can also be determined according to the first analysis results of each dimension, for example, the total score can be determined according to the quantized scores of each dimension, and the total analysis result can be displayed on the first interface.

[0051] In addition, analysis components corresponding to different dimensions / sub-dimensions in the first preset dimension can also be deployed in the first interface. Correspondingly, in response to the triggering of the analysis component in the first interface, a tool corresponding to the analysis component can be called to perform analysis on the first music data in the corresponding dimension. At this time, a first analysis result display corresponding to the analysis component can be presented in the first interface.

[0052] In the technical solution of the embodiment of the present disclosure, first music data can be obtained and presented in the first interface; based on the analysis of the first music data in the first preset dimension, a first analysis result is obtained; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension; the first analysis result is presented in the first interface. By presenting the analysis result of the first music data in the first preset dimension on the interface, it can provide a reference for users, help users improve the quality and production efficiency of music works, and thus enhance the user experience.

[0053] The various alternative solutions in the method for processing music data provided in the embodiment of the present disclosure and the above embodiments can be combined. The method for processing music data provided in this embodiment makes a detailed supplement to the analysis and analysis result display of music data in the editing stage.

[0054] The method for processing music data provided in this embodiment may further include: in response to the triggering of the first entry component in the first interface, presenting second music data and an editing component in the second interface; wherein, the second music data includes the data in the first music data corresponding to the first entry component; in response to the triggering of the editing component, performing corresponding editing processing on the second music data to obtain third music data; based on the analysis of the third music data in the second preset dimension, obtaining a second analysis result; wherein, the second preset dimension includes the first preset dimension; presenting the third music data and the second analysis result in the second interface.

[0055] In this embodiment, a first entry component can be deployed in the first interface. Among them, the first entry component can be regarded as the entry component of the second interface and can be deployed in the form of components such as an icon, text, or button. Among them, an editing component can be deployed in the second interface. The editing component may include, but is not limited to: a mixing component, a track adding component, a cutting component, a copy-paste component, a sound cancellation component, and an effect adding component, etc. It can be considered that the second interface can be used to edit music data.

[0056] Among them, corresponding first entry components can be deployed for all the first music data displayed in the first interface. For example, corresponding editing icon components (i.e., the first entry components) can be deployed for each piece of first music data. Thus, in response to the triggering of the first entry component, the first music data corresponding to the first entry component (i.e., the second music data) can be displayed in the second interface. Furthermore, in response to the triggering of the editing component in the second interface, corresponding editing processing can be performed on the second music data to obtain the edited second music data (i.e., the third music data).

[0057] After editing to obtain the third music data, analysis in a second preset dimension can also be performed on the third music data. Among them, the second preset dimension can include the first preset dimension and can also include a dimension for evaluating the processing effect of the editing processing, etc. Exemplarily, in the case where the editing processing includes vocal elimination processing, the second preset dimension can include, in addition to the first preset dimension, a dimension such as the vocal elimination quality dimension. Among them, the third music data can be analyzed in the second preset dimension through existing tools and / or self-developed tools, and a second analysis result can be obtained based on this analysis. In addition, in some implementation manners, through a preset mapping rule, the analysis result of the third music data in the second preset dimension can be quantitatively mapped to obtain a quantitatively analyzed result, which is used as the second analysis result. In some implementation manners, the second analysis result can be, for example, a score, etc.

[0058] In this embodiment, at least one editing component can be triggered to perform at least one stage of editing processing on the second music data to obtain the third music data at the corresponding editing stage. And, after the editing processing of each stage is completed, analysis in the second preset dimension can be performed on the obtained third music data to obtain the corresponding second analysis result. Thus, it is possible to provide real-time feedback on the editing effect for the user during the music editing stage, which can help the user improve the quality and production efficiency of the music work, so as to improve the user experience.

[0059] In addition, in some implementation manners, the second interface can also be entered through the entry component of another interface (such as the home page navigation component of the application software). And, the second music data can be imported into the second interface to edit the second music data to obtain the third music data. Furthermore, analysis in the second preset dimension can be performed on the third music data to provide real-time feedback on the editing effect of the music data, which can improve the user experience.

[0060] Exemplarily, Figure 3 is a schematic diagram of the second interface in a method for processing music data provided by an embodiment of the present disclosure. Refer to Figure 3, the second music data displayed in the second interface may include the music data corresponding to Track 1 and Track 2. Among them, in response to the triggering of the mixing component in the second interface, the music data corresponding to Track 1 and Track 2 can be mixed. Among them, the music data during the mixing process cannot be deleted; the status of the third music data during the mixing process can include "generating", "generation failed", and "generation successful". For the third music data with successful generation, analysis in a second preset dimension can be performed to obtain a second analysis result.

[0061] As Figure 3 shown, the second analysis result may include the analysis results in the audio quality dimension and the music content dimension. Figure 3 Among them, the analysis result in the audio quality dimension may include the analysis results in the audio parameter dimension, the abnormal audio dimension, and the MOS dimension; among them, the analysis result in the audio parameter dimension may include "passed" and "not passed"; the analysis result in the abnormal audio dimension may include a description of the abnormal situation, such as Figure 3 the presence of noise and popping can be described; among them, the analysis result in the MOS dimension may include a score. Figure 3 Among them, the music content dimension may include scores in dimensions such as musical scores, structural paragraphs, and rhythms. Thus, it can help users understand the quality level of the third music data after mixing processing.

[0062] In some alternative implementation manners, the editing component may include a sound conversion component; in response to the triggering of the editing component, corresponding editing processing is performed on the second music data to obtain third music data, which may include: in response to the triggering of the sound conversion component, the second music data is separated to obtain first dry voice data and accompaniment data; the first dry voice data is subjected to sound conversion processing corresponding to the sound conversion component to obtain second dry voice data; based on the mixing processing of the second dry voice data and the accompaniment data, third music data is obtained.

[0063] Among them, the sound conversion component can be used to perform sound conversion (Speech-to-Voice Conversion, SVC) processing on the second music data. Among them, SVC processing can be understood as the processing of converting the sound characteristics in the audio into another sound characteristic while retaining the content information of the audio. Among them, at least one sound conversion component can be deployed in the second interface to correspond to different sound characteristics. Exemplarily, sound conversion components corresponding to the sound characteristics of Person 1 - Person N can be deployed in the second interface.

[0064] In this embodiment, in response to the triggering of the voice conversion component, the second music data can first be separated by the existing Music Source Separation (MSS) technology to obtain the first dry voice data and the accompaniment data. Among them, the first dry voice data can be regarded as the vocal data in the second music data, and the accompaniment data can be regarded as the data in the second music data except for the vocal data. Then, through the existing SVC technology, the voice characteristics of the first dry voice data can be converted into the voice characteristics corresponding to the voice conversion component to obtain the second dry voice data. Finally, through the existing mixing technology, the second dry voice data and the accompaniment data can be mixed, and based on the mixing process, the third music data after voice conversion can be obtained.

[0065] In these alternative implementation manners, the second music data can be subjected to SVC processing to obtain the third music data after voice conversion. Among them, when at least two voice conversion components are triggered, the third music data corresponding to each voice conversion component can be obtained. By analyzing the third music data in the second preset dimension, the quality evaluation of the music work after voice conversion processing can be provided for the user. Thus, it can help the user identify the high-quality third music data to improve the quality and production efficiency of the music work.

[0066] In addition, in some further implementation manners, it may further include: obtaining a third analysis result based on the analysis of the first dry voice data in the third preset dimension. Among them, the third preset dimension includes the separation quality dimension. The first dry voice data and the third analysis result are displayed on the second interface. Among them, the separation quality dimension of the first dry voice data can be analyzed by calculating the Signal-to-Distortion Ratio (SDR) between the second music data and the separated first dry voice data, and the third analysis result can be obtained based on this analysis. Thus, the quality of the first dry voice data can be evaluated.

[0067] In some further implementation manners, it may further include: obtaining a fourth analysis result based on the analysis of the second dry voice data in the fourth preset dimension. Among them, the fourth preset dimension includes the conversion similarity dimension. The second dry voice data and the fourth analysis result are displayed on the second interface. Among them, the conversion similarity may include, but is not limited to, the timbre similarity and the style similarity. Among them, the timbre similarity and / or the style similarity can be analyzed through the existing neural network model, and the fourth analysis result can be obtained based on this analysis. Thus, the quality of the second dry voice data can be evaluated.

[0068] In these further implementation manners, during the SVC processing of the second music data, the result of the separation processing may be analyzed to obtain a third analysis result, and / or the result of the voice conversion processing may be analyzed to obtain a fourth analysis result. Moreover, the third analysis result and / or the fourth analysis result may be displayed on the second interface to help the user understand the impact of each processing step on the music quality.

[0069] Exemplarily, Figure 4 is a schematic diagram of the second interface in a method for processing music data provided by an embodiment of the present disclosure. Refer to Figure 4 , four voice conversion components (such as components 1 - 4 in the figure) are deployed on the second interface, which may respectively correspond to the voice characteristics of characters 1 - 4. In response to the triggering of components 1 and 2, the voice characteristics of the second music data may be processed into the voice characteristics of characters 1 and 2. For example, Figure 4 , the processing status of the voice conversion processing may include "generating", "generation failed", and "generation successful". In the generation successful state, the first dry voice data and the second dry voice data corresponding to component 1 may be displayed on component 5 in the second interface, and the first dry voice data and the second dry voice data corresponding to component 2 may be displayed on component 6 of the second interface. Moreover, the scores of the timbre similarity and style similarity of the two second dry voice data may be displayed on components 5 and 6.

[0070] In addition, as Figure 4 shown, the accompaniment data may also be displayed in components 5 and 6. In response to the triggering of the mixing component in component 5 or component 6, the corresponding second dry voice data and the accompaniment data may be mixed. Exemplarily, the interface of the mixing processing process may refer to Figure 3 the interface shown.

[0071] The technical solution of the embodiment of the present disclosure makes a detailed supplement to the analysis and display of the analysis results of music data in the editing stage. In this embodiment, the second interface can be entered from the first interface, and music data can be edited and processed on the second interface. Furthermore, music data can be analyzed in various preset dimensions in the editing stage, and the analysis results can be displayed on the second interface. Thus, it can be realized that in the music editing stage, a reference for the editing effect can be provided to the user in a timely manner, which can help the user improve the quality and production efficiency of music works, so as to improve the user experience. The method for processing music data provided by the embodiment of the present disclosure and the method for processing music data provided by the above embodiment belong to the same general inventive concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0072] The optional solutions in the music data processing method provided in the embodiments of the present disclosure can be combined with those in the above embodiments. The music data processing method provided in this embodiment describes in detail the analysis of music data based on cross-modal understanding ability and the display of analysis results.

[0073] In the music data processing method provided in this embodiment, the first preset dimension further includes a matching dimension with a preset hot word; the analysis process of the first music data in the first preset dimension may include: analyzing the first music data in the matching dimension with the preset hot word through a first cross-modal model.

[0074] Among them, the preset hot word may include at least one. Among them, on the premise of complying with the requirements of corresponding laws, regulations and related provisions, Internet data can be obtained through existing data acquisition technologies, and the Internet data can be processed to obtain the preset hot word.

[0075] Among them, the first cross-modal model may include a model with the processing ability of audio data and text data. For example, it may include a Contrastive Language-Image Pre-training (CLIP) model and a Convolutional Localized Audio Pre-training (CLAP) model, etc. Among them, the first music data and the preset hot word can be input into the first cross-modal model to output the matching degree (such as a matching score) between the first music data and the preset hot word through the first cross-modal model.

[0076] In this embodiment, by analyzing the first music data in the matching dimension with the preset hot word, it can help users understand the linkage between music works and current hot information and improve the user experience. In addition, in some implementation manners, the data obtained by editing the first music data (such as the third music data) can also be analyzed in the matching dimension with the preset hot word to further improve the user experience.

[0077] In some optional implementation manners, it may further include: in response to the triggering of a second entry component in the first interface, displaying a session window in the third interface; receiving first session content and fourth music data through an input component of the session window; where the first session content includes at least one round of session content; the fourth music data includes the first music data; generating second session content according to the fourth music data and the first session content through a second cross-modal model; and displaying the second session content in the session window.

[0078] Among them, a second entry component can be deployed on the first interface. Among them, the second entry component can be considered as the entry component of the third interface and can be deployed in the form of components such as icons, texts, or buttons. Among them, a session window can be included in the third interface. Among them, an input component can be deployed in the session window, and through the input component, the session content (i.e., the first session content) and the fourth music data input by the user can be received. Among them, the first session content includes at least one round of session content; among them, the fourth music data can include at least one music data. The first session content and the fourth music data can belong to the same session round or different session rounds.

[0079] Among them, the second cross-modal model can include a model with the processing capabilities of audio data and text data. For example, it can include an existing large language model (LLM). The processing device of the music data can call the second cross-modal model to understand and analyze the most recently received fourth music data and the most recently received first session content, and generate the session content for reply (i.e., the second session content). Furthermore, the second session content can be displayed in the session window to achieve interactive dialogue with the user.

[0080] Exemplarily, when the first session content includes "what instruments are mainly used in this song", the second session content can include the corresponding instruments; when the first session content includes "please describe the details of this song", the second session content can include the analysis results in the dimension of music content; when the first session content includes "what other songs do people who like this song also like", the second session content can include the retrieval results of other songs of the same style genre.

[0081] In addition, in some implementation manners, the third interface can also be entered through the entry component of other interfaces (such as the home page navigation component of the application software). And any fourth music data can be input in the third interface to conduct an interactive dialogue with the user for the fourth music data.

[0082] Exemplarily, Figure 5 is a schematic diagram of the third interface in a method for processing music data provided by an embodiment of the present disclosure. Refer to Figure 5 , a session window 510 can be displayed in the third interface, and an input component 520 corresponding to the session window can be deployed. As Figure 5 shown, the input component 520 can include, for example, a text box input component 521 and a music data import component 522. Through the text box input component 521 and the music data import component 522, the first session content 530 and the fourth music data 540 can be received respectively.

[0083] Figure 5Among them, the first session content 530 may include multiple rounds of session content. Among them, the first session content 530 and the fourth music data 540 may belong to the same session round or different session rounds. For example, Figure 5 As shown, the first session content 530 in the first round and the fourth music data 540 belong to the same session round, and the first session content 530 in the second round and the fourth music data 540 belong to different session rounds.

[0084] Among them, in response to receiving the first session content 530 and the fourth music data 540 in the first round, the corresponding second session content 550 may be generated and displayed through the second cross-modal model. In response to the first session content 530 in the second round and later, the second cross-modal model may generate and display the second session content 550 according to the most recently received fourth music data 540 (i.e., the music data 540 received in the first round) and the currently received first session content 530.

[0085] In these alternative implementations, the second cross-modal model can be used to understand and analyze music data, and the session window can be used to interactively evaluate music data based on the understanding and analysis results, which can further improve the user experience.

[0086] The technical solution of the embodiments of the present disclosure has described in detail the analysis of music data based on cross-modal understanding ability and the display of analysis results. Through a model with cross-modal analysis ability, hotspot matching analysis of music data can be realized, and interactive music content evaluation can be realized, which can further enhance the user experience. The music data processing method provided by the embodiments of the present disclosure and the music data processing method provided by the above embodiments belong to the same general concept. Technical details not described in detail in this embodiment can be referred to the above embodiments, and the same technical features have the same beneficial effects in this embodiment and the above embodiments.

[0087] The various alternative solutions in the music data processing methods provided in the embodiments of the present disclosure and the above embodiments can be combined. The music data processing method provided in this embodiment has described the application scenario in detail.

[0088] The music data processing method provided in this embodiment can be applied to music application software; the method further includes: in response to the first analysis result meeting a preset condition and the triggering of the distribution component in the first interface, submitting a distribution request for the first music data to the server.

[0089] In this embodiment, the music data processing device can be integrated into music application software. The music data processing method provided in any embodiment of the present disclosure can be executed through music application software.

[0090] Among them, the preset conditions may include conditions with default configurations, and corresponding preset conditions can be configured for each preset dimension. Among them, if the first analysis result meets the preset conditions, it can be considered that the quality of the first music data is relatively high, and the probability of the release request passing is relatively large; if the first analysis result does not meet the preset conditions, it can be considered that the quality of the first music data needs to be improved, and the probability of the release request passing is relatively low.

[0091] Among them, a release component can be deployed in the first interface. In this embodiment, the music application software can submit a release request for the first music data to the server in response to the first analysis result meeting the preset conditions and the triggering of the release component in the first interface; correspondingly, the server can perform the review of the first music data. Among them, the server can execute the release process of the first music data in response to the review passing.

[0092] In addition, the music application software can also, in response to the first analysis result not meeting the preset conditions and the triggering of the release component in the first interface, generate a prompt message according to the first analysis result that does not meet the preset conditions, and display the prompt message on the first interface. Thus, it can help the user understand the quality shortcoming part of the music work, so as to facilitate the user to make targeted editing adjustments.

[0093] In some other implementation manners, a release component can also be deployed on other interfaces of the music software. For example, a release component can be deployed on the second interface. And, in response to the second analysis result meeting the preset conditions and the triggering of the release component in the second interface, a release request for the third music data can be submitted to the server. Thus, after the first music data is edited and its analysis result meets the preset conditions, a release request for the third music data can be submitted. By increasing the deployment scenarios of the release component, a release request can be submitted conveniently and flexibly.

[0094] Exemplarily, Figure 6 is a schematic data flow diagram of a method for processing music data provided by an embodiment of the present disclosure. Refer to Figure 6 , the music data processing device may include a music quality inspection toolkit, and the music quality inspection toolkit may include tools with cross-modal understanding capabilities, analysis tools for audio quality dimensions, analysis tools for music content dimensions, tools for separating quality dimensions, and tools for converting similarity dimensions.

[0095] Figure 6 exemplarily shows a music processing link, and corresponding tools in the toolkit can be called to analyze the music data at each stage of music data processing. As Figure 6 shown, the multi-stage analysis of music data may include:

[0096] Among them, the first music data can be generated by the AIGC model and displayed on the first interface. Among them, the analysis tools in the toolkit for the audio quality dimension and the music content dimension can be called through the first interface to analyze the first music data in the corresponding dimensions and obtain the first analysis result. Among them, the first analysis result can be displayed on the first interface for users' reference.

[0097] Among them, in response to the triggering of the first entry component on the first interface, the second music data and the editing component can be displayed on the second interface; among them, the second music data includes the data corresponding to the first entry component in the first music data; among them, the editing component may include a voice conversion component.

[0098] Among them, in response to the triggering of the voice conversion component, separation processing ( Figure 6 denoted by MSS herein) can be performed on the second music data to obtain the first dry voice data and the accompaniment data. The tool for the separation quality dimension in the toolkit can be called through the second interface to analyze the first dry voice data in the corresponding dimension and obtain the third analysis result. Among them, the third analysis result can be displayed on the second interface for users' reference.

[0099] Next, voice conversion processing ( Figure 6 denoted by SVC herein) corresponding to the voice conversion component can be performed on the first dry voice data to obtain the second dry voice data. The tool for the conversion similarity dimension in the toolkit can be called through the third interface to analyze the second dry voice data in the corresponding dimension and obtain the fourth analysis result. Among them, the fourth analysis result can be displayed on the second interface for users' reference.

[0100] Finally, mixing processing can be performed on the second dry voice data and the accompaniment data to obtain the third music data. And the analysis tools in the toolkit for the audio quality dimension, the music content dimension, and the conversion similarity dimension can be called through the fourth interface to analyze the third music data in the corresponding dimensions and obtain the second analysis result. Among them, the second analysis result can be displayed on the second interface for users' reference.

[0101] Refer again to Figure 6 , the tool for the cross-modal understanding ability in the tool can be called through the fifth interface. Among them, the tool for the cross-modal understanding ability may include the first cross-modal model and the second cross-modal model. Through the first cross-modal model, matching analysis with the preset hot words can be performed on the first music data or the third music data. And the matching analysis result can be displayed on the first interface or the second interface for users' reference. In addition, through the second cross-modal model, the second conversation content can be generated according to the fourth music data (including the first music data and the third music data) and the first conversation content. And the first conversation content and the second conversation content can be displayed through the conversation window on the third interface to achieve interactive music evaluation.

[0102] In addition, a distribution request for the first music data or the third music data can be submitted to the server to request the distribution of the first music data or the third music data.

[0103] In the related art, there is no effective quality inspection scheme deployed in the music platform, and due to the long AIGC production link, users cannot locate the potential problem occurrence points based on the music finished product. However, the technical solution of the present disclosure embodiment provides a music quality inspection scheme embedded in the AIGC production process, which can perform corresponding dimension analysis at each stage of music production to provide real-time feedback on the music generation effect, helping users improve the generation quality of music works and accelerating the production process. In addition, based on the cross-modal ability of the model, hot spot matching analysis of music data can be realized, and interactive music content evaluation can be achieved, which can further enhance the user experience.

[0104] The technical solution of the present disclosure embodiment describes the application scenario in detail. By integrating the processing device for music data provided in the present disclosure embodiment in a music application software, the corresponding music processing method can be executed. Thus, it is possible to perform analysis in various preset dimensions and display the analysis results in the production process of music works at multiple stages. Thereby, real-time feedback on the music generation effect can be provided for users, which can help users improve the quality of music works, and further improve the music distribution efficiency to enhance the user experience. In addition, the music data processing method provided in the present disclosure embodiment and the music data processing method provided in the above embodiment belong to the same general concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0105] It should be noted that according to the actual logical requirements of the music application software, the interfaces with different prefixes such as "first" and "second" in the present disclosure can be designed as different regions of the same interface in the application software, or multiple interfaces can be defined as interfaces with the same prefix in the present disclosure, which is not specifically limited here. Exemplarily, in a music application software, an interface with music data display function and editing function can be referred to as both the first interface and the second interface. Another example, in a music application software, each interface with music data display function (such as the creation interface and management interface of music data, etc.) can be defined as the first interface.

[0106] Figure 7 It is a schematic structural diagram of a music data processing device provided by an embodiment of the present disclosure. The music data processing device provided in this embodiment is applicable to the situation of quality inspection and analysis of music data, for example, applicable to the situation of quality inspection and analysis of music data in a music application software.

[0107] Such asFigure 7 As shown in Figure 7 , the processing device for music data provided by an embodiment of the present disclosure may include:

[0108] A display module 710, configured to obtain first music data and display the first music data on a first interface;

[0109] An analysis module 720, configured to obtain a first analysis result based on an analysis of the first music data in a first preset dimension; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension;

[0110] The display module 710 is further configured to display the first analysis result on the first interface.

[0111] In some optional implementation manners, the display module may be configured to:

[0112] Receive a prompt word through the first interface; wherein, the prompt word includes lyrics and music tags;

[0113] In response to a trigger on a music generation component in the first interface, generate first music data according to the prompt word.

[0114] In some optional implementation manners, the display module may further be configured to, in response to a trigger on a first entry component in the first interface, display second music data and an editing component on a second interface; wherein, the second music data includes data corresponding to the first entry component in the first music data;

[0115] The processing device for music data may further include:

[0116] An editing module, configured to perform corresponding editing processing on the second music data in response to a trigger on the editing component to obtain third music data;

[0117] The analysis module may further be configured to obtain a second analysis result based on an analysis of the third music data in a second preset dimension; wherein, the second preset dimension includes the first preset dimension;

[0118] The display module is further configured to display the third music data and the second analysis result on the second interface.

[0119] In some optional implementation manners, the editing component includes a voice conversion component;

[0120] The editing module may be configured to:

[0121] In response to a trigger on the voice conversion component, perform separation processing on the second music data to obtain first dry voice data and accompaniment data;

[0122] Perform voice conversion processing corresponding to the voice conversion component on the first dry voice data to obtain second dry voice data;

[0123] Based on the mixing process of the second dry vocal data and the accompaniment data, the third music data is obtained.

[0124] In some alternative implementation manners, the analysis module may further be configured to obtain a third analysis result based on the analysis of the first dry vocal data in a third preset dimension; wherein, the third preset dimension includes a separation quality dimension.

[0125] The display module may further be configured to display the first dry vocal data and the third analysis result on a second interface.

[0126] In some alternative implementation manners, the analysis module may further be configured to obtain a fourth analysis result based on the analysis of the second dry vocal data in a fourth preset dimension; wherein, the fourth preset dimension includes a conversion similarity dimension.

[0127] The display module may further be configured to display the second dry vocal data and the fourth analysis result on a second interface.

[0128] In some alternative implementation manners, the first preset dimension further includes a matching dimension with a preset hot word.

[0129] The analysis module may be configured to analyze the first music data in terms of the matching dimension with the preset hot word through a first cross-modal model.

[0130] In some alternative implementation manners, the display module may further be configured to, in response to the triggering of a second entry component in the first interface, display a session window on a third interface.

[0131] The processing device for music data may further include:

[0132] An interaction module, which may be configured to:

[0133] Receive a first session content and a fourth music data through an input component of the session window; wherein, the first session content includes at least one round of session content; the fourth music data includes the first music data.

[0134] Generate a second session content according to the fourth music data and the first session content through a second cross-modal model.

[0135] The display module may further be configured to display the second session content in the session window.

[0136] In some alternative implementation manners, it is applied to a music application software.

[0137] The processing device for music data may further include:

[0138] A request module, configured to submit a release request for first music data to a server in response to a first analysis result satisfying a preset condition and a trigger on a release component in a first interface.

[0139] The music data processing device provided by the embodiments of the present disclosure can execute the music data processing method provided by any embodiment of the present disclosure, and has function modules and beneficial effects corresponding to the execution of the method.

[0140] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0141] The following refers to Figure 8 , which shows a schematic structural diagram of an electronic device (such as Figure 8 the terminal device or server in) 800 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0142] As Figure 8 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0143] Typically, the following devices can be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 can allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0144] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the processing method of music data of the embodiment of the present disclosure are executed.

[0145] The electronic device provided by the embodiment of the present disclosure and the processing method of music data provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0146] An embodiment of the present disclosure provides a storage medium storing computer-executable instructions, and the computer-executable instructions can be used to execute the processing method of music data provided by the above embodiment when executed by a computer processor.

[0147] It should be noted that the above storage medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores executable instructions that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable executable instructions. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit executable instructions for use by or in conjunction with an instruction execution system, apparatus, or device. The executable instructions contained on the storage medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0148] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0149] The above storage medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device.

[0150] The above storage medium carries one or more executable instructions, and when the one or more executable instructions are executed by the electronic device, the electronic device is caused to:

[0151] Obtain first music data and display the first music data on a first interface; based on the analysis of the first music data in a first preset dimension, obtain a first analysis result; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension; display the first analysis result on the first interface.

[0152] Executable instructions for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The executable instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0153] An embodiment of the present disclosure further provides a computer program product, including a computer program, which when executed by a processor can implement the music data processing method provided in any one of the embodiments of the present disclosure.

[0154] Wherein, the computer program product includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the music data processing method. Among them, the program codes may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the names of the units and modules do not, in some cases, constitute a limitation on the units and modules themselves.

[0157] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0158] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0159] According to one or more embodiments of the present disclosure, there is provided a method for processing music data, the method comprising:

[0160] Obtaining first music data and presenting the first music data on a first interface;

[0161] Based on an analysis of the first music data in a first preset dimension, obtaining a first analysis result; wherein the first preset dimension includes an audio quality dimension and / or a music content dimension;

[0162] Presenting the first analysis result on the first interface.

[0163] According to one or more embodiments of the present disclosure, there is provided a method for processing music data, further comprising:

[0164] In some alternative implementations, the obtaining of the first music data includes:

[0165] Receiving a prompt word through the first interface; wherein the prompt word includes lyrics and music tags;

[0166] In response to a trigger of a music generation component in the first interface, generating first music data according to the prompt word.

[0167] According to one or more embodiments of the present disclosure, there is provided a method for processing music data, further comprising:

[0168] In some alternative implementations, in response to a trigger of a first entry component in the first interface, presenting second music data and an editing component on a second interface; wherein the second music data includes data in the first music data corresponding to the first entry component;

[0169] In response to the triggering of the editing component, perform corresponding editing processing on the second music data to obtain third music data;

[0170] Based on the analysis of the third music data in a second preset dimension, obtain a second analysis result; wherein, the second preset dimension includes the first preset dimension;

[0171] Display the third music data and the second analysis result on the second interface.

[0172] According to one or more embodiments of the present disclosure, there is provided a method for processing music data, further including:

[0173] In some alternative implementation manners, the editing component includes a sound conversion component;

[0174] The step of in response to the triggering of the editing component, performing corresponding editing processing on the second music data to obtain third music data includes:

[0175] In response to the triggering of the sound conversion component, perform separation processing on the second music data to obtain first dry vocal data and accompaniment data;

[0176] Perform sound conversion processing corresponding to the sound conversion component on the first dry vocal data to obtain second dry vocal data;

[0177] Based on the mixing processing of the second dry vocal data and the accompaniment data, obtain the third music data.

[0178] According to one or more embodiments of the present disclosure, there is provided a method for processing music data, further including:

[0179] In some alternative implementation manners, based on the analysis of the first dry vocal data in a third preset dimension, obtain a third analysis result; wherein, the third preset dimension includes a separation quality dimension;

[0180] Display the first dry vocal data and the third analysis result on the second interface.

[0181] According to one or more embodiments of the present disclosure, there is provided a method for processing music data, further including:

[0182] In some alternative implementation manners, based on the analysis of the second dry vocal data in a fourth preset dimension, obtain a fourth analysis result; wherein, the fourth preset dimension includes a conversion similarity dimension;

[0183] Display the second dry vocal data and the fourth analysis result on the second interface.

[0184] According to one or more embodiments of the present disclosure, a method for processing music data is provided, further including:

[0185] In some alternative implementation manners, the first preset dimension further includes a matching dimension with a preset hot word;

[0186] The analysis process of the first music data in the first preset dimension includes:

[0187] Analyzing the matching dimension with the preset hot word of the first music data through a first cross-modal model.

[0188] According to one or more embodiments of the present disclosure, a method for processing music data is provided, further including:

[0189] In some alternative implementation manners, in response to the triggering of a second entry component in the first interface, a conversation window is displayed on a third interface;

[0190] Receiving first conversation content and fourth music data through an input component of the conversation window; wherein, the first conversation content includes at least one round of conversation content; the fourth music data includes the first music data;

[0191] Generating second conversation content according to the fourth music data and the first conversation content through a second cross-modal model;

[0192] Displaying the second conversation content in the conversation window.

[0193] According to one or more embodiments of the present disclosure, a method for processing music data is provided, further including:

[0194] In some alternative implementation manners, it is applied to a music application software; the method further includes:

[0195] In response to the first analysis result meeting a preset condition and the triggering of a release component in the first interface, submitting a release request for the first music data to a server.

[0196] According to one or more embodiments of the present disclosure, a device for processing music data is provided, and the device includes:

[0197] A display module, configured to obtain first music data and display the first music data on a first interface;

[0198] An analysis module, configured to obtain a first analysis result based on the analysis of the first music data in a first preset dimension; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension;

[0199] The display module is further configured to display the first analysis result on the first interface.

[0200] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0201] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0202] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A method for processing music data, characterized in that It includes: Obtain first music data and display the first music data on a first interface; Based on the analysis of the first music data in a first preset dimension, obtain a first analysis result; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension; Display the first analysis result on the first interface.

2. The method according to claim 1, wherein The obtaining of the first music data includes: Receive a prompt word through the first interface; wherein, the prompt word includes lyrics and music tags; In response to the triggering of a music generation component in the first interface, generate first music data according to the prompt word.

3. The method according to claim 1, wherein It further includes: In response to the triggering of a first entry component in the first interface, display second music data and an editing component on a second interface; wherein, the second music data includes data corresponding to the first entry component in the first music data; In response to the triggering of the editing component, perform corresponding editing processing on the second music data to obtain third music data; Based on the analysis of the third music data in a second preset dimension, obtain a second analysis result; wherein, the second preset dimension includes the first preset dimension; Display the third music data and the second analysis result on the second interface.

4. The method according to claim 3, wherein The editing component includes a voice conversion component; The responding to the triggering of the editing component and performing corresponding editing processing on the second music data to obtain third music data includes: In response to the triggering of the voice conversion component, perform separation processing on the second music data to obtain first dry voice data and accompaniment data; Perform voice conversion processing corresponding to the voice conversion component on the first dry voice data to obtain second dry voice data; Based on the mixing processing of the second dry voice data and the accompaniment data, obtain the third music data.

5. The method according to claim 4, characterized in that, It further includes: Based on the analysis of the first dry voice data in a third preset dimension, obtain a third analysis result; wherein, the third preset dimension includes a separation quality dimension; Display the first dry voice data and the third analysis result on the second interface.

6. The method according to claim 4, characterized in that, It further includes: Based on the analysis of the second dry voice data in a fourth preset dimension, obtain a fourth analysis result; wherein, the fourth preset dimension includes a conversion similarity dimension; Display the second dry voice data and the fourth analysis result on the second interface.

7. The method according to claim 1, wherein The first preset dimension further includes a matching dimension with preset hot words; The analysis process of the first music data in the first preset dimension includes: Through a first cross-modal model, analyze the matching dimension of the first music data with the preset hot words.

8. The method according to claim 1, wherein It further includes: In response to the triggering of a second entry component in the first interface, display a session window on a third interface; Receive first session content and fourth music data through an input component of the session window; wherein, the first session content includes at least one round of session content; the fourth music data includes the first music data; Through a second cross-modal model, generate second session content according to the fourth music data and the first session content; The second session content is displayed in the session window as described above.

9. The method according to any one of claims 1-8, characterized in that, Applied to a music application software; the method further includes: In response to the first analysis result meeting a preset condition and a trigger on the distribution component in the first interface, submitting a distribution request for the first music data to the server.

10. A processing device for music data, characterized in that, Comprising: A display module, configured to obtain first music data and display the first music data in a first interface; An analysis module, configured to obtain a first analysis result based on an analysis of the first music data in a first preset dimension; wherein, the first preset dimension includes an audio quality dimension and / or a music content dimension; The display module is further configured to display the first analysis result in the first interface.

11. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the music data processing method according to any one of claims 1-9.

12. A storage medium containing computer-executable instructions, the computer-executable instructions being used to execute the music data processing method according to any one of claims 1-9 when executed by a computer processor.

13. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program implements the music data processing method according to any one of claims 1-9 when executed by a processor.