Audio file generation method and device

By receiving audio generation instructions, obtaining multimodal reference information, calling large language models to identify needs, and generating songs and music score descriptions, the problem that audio files in the prior art cannot meet user needs and improve user experience.

CN120375799APending Publication Date: 2025-07-25ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510406393.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing audio file generation methods cannot meet the diverse needs of users, resulting in poor user experience.

Method used

By receiving audio generation instructions, obtaining multimodal reference information, calling the target large language model for intent recognition to determine the target audio generation requirements, and generating song descriptions and music score descriptions based on the requirements, and finally generating the target audio file.

Benefits of technology

Ensure that the generated audio files can meet user needs and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375799A_ABST
    Figure CN120375799A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio file generation method and device. According to the embodiment of the invention, after an audio generation instruction is received, multi-modal reference information is obtained, and at least one target large language model matched with the multi-modal reference information is called to carry out intention recognition on the multi-modal reference information so as to determine a target audio generation demand; and determining corresponding song description and music score description according to the target audio generation demand, and generating a target audio file according to the song description and the music score description. Wherein the target audio generation demand is used for representing a generation demand for an audio file. Therefore, by supporting the input of the multi-modal reference information and generating the target audio file according to the multi-modal reference information, the embodiment of the invention can ensure that the generated audio file can meet the user demand, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a method and device for generating audio files. Background Art

[0002] With the continuous development of artificial intelligence technology, large language models have begun to be widely used in various industries. Among them, in order to meet the diverse living needs of users, large language models are often used to provide audio file generation services for users to support them in generating corresponding audio files according to their own needs. However, when using large language models to provide audio file generation services, the audio files generated by existing audio file generation methods usually cannot meet user needs, resulting in a poor user experience. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method and device for generating audio files to ensure that the generated audio files can meet user needs, thereby improving the user experience.

[0004] In a first aspect, an embodiment of the present invention aims to provide a method for generating an audio file, the method comprising:

[0005] Receiving an audio generation instruction;

[0006] Obtaining multi-modal reference information;

[0007] Invoking at least one target large language model to perform intent recognition on the multi-modal reference information to determine a target audio generation requirement, wherein the target large language model matches the multi-modal reference information, and the target audio generation requirement is used to characterize the generation requirement for the audio file;

[0008] Determining a corresponding song description and sheet music description according to the target audio generation requirement;

[0009] Generating a target audio file according to the song description and the sheet music description.

[0010] In a second aspect, an embodiment of the present invention aims at an audio file generation device, the device comprising:

[0011] An instruction receiving unit for receiving an audio generation instruction;

[0012] A multi-modal reference information obtaining unit for obtaining multi-modal reference information;

[0013] An intent recognition unit for invoking at least one target large language model to perform intent recognition on the multi-modal reference information to determine a target audio generation requirement, wherein the target large language model matches the multi-modal reference information, and the target audio generation requirement is used to characterize the generation requirement for the audio file;

[0014] A description information determination unit, configured to generate a corresponding song description and a score description according to the target audio for demand determination;

[0015] An audio file generation unit, configured to generate a target audio file according to the song description and the score description.

[0016] In a third aspect, an embodiment of the present invention aims at a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the method described in the first aspect.

[0017] In a fourth aspect, an embodiment of the present invention aims at an electronic device, the device includes:

[0018] A memory, configured to store one or more computer program instructions;

[0019] A processor, and the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.

[0020] In a fifth aspect, an embodiment of the present invention aims at a computer program product, when the computer program product runs on a computer, the computer is caused to execute the method described in the first aspect.

[0021] After receiving an audio generation instruction, an embodiment of the present invention will obtain multimodal reference information, and call at least one target large language model that matches the multimodal reference information to perform intent recognition on the multimodal reference information to determine the target audio generation demand, and then determine the corresponding song description and score description according to the target audio generation demand, and further generate a target audio file according to the song description and the score description. Among them, the target audio generation demand is used to characterize the generation demand for the audio file. Thus, by supporting the input of multimodal reference information and generating a target audio file according to the multimodal reference information, an embodiment of the present invention can ensure that the generated audio file can meet the user's needs, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0023] Figure 1 is a flowchart of the audio file generation method according to an embodiment of the present invention;

[0024] Figure 2 is a flowchart of the target large language model determination method according to an embodiment of the present invention;

[0025] Figure 3 is a flowchart of the intent recognition method according to an embodiment of the present invention;

[0026] Figure 4 Flow chart of the description generation method according to an embodiment of the present invention;

[0027] Figure 5 Flow chart of the audio file generation method according to an embodiment of the present invention;

[0028] Figure 6 Flow chart of the human voice signal generation method according to an embodiment of the present invention;

[0029] Figure 7 Schematic diagram of the audio file generation apparatus according to an embodiment of the present invention;

[0030] Figure 8 Schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0031] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these detail parts. In order to avoid obscuring the essence of the present application, well-known methods, processes, flows, elements and circuits are not described in detail.

[0032] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.

[0033] Unless the context clearly requires otherwise, the words such as "including", "comprising" and the like in the entire application document shall be construed as the meaning of including rather than exclusive or exhaustive; that is, the meaning of "including but not limited to".

[0034] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0035] For the solutions described in this specification and the embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.

[0036] Figure 1 Flow chart of the audio file generation method according to an embodiment of the present invention. It should be understood that, Figure 1The execution subject of the audio file generation method shown can be the corresponding audio file generation device. By executing Figure 1 the audio file generation method shown, the audio file generation device can support multi-modal reference information input and generate a target audio file according to the multi-modal reference information. Thus, this embodiment can ensure that the generated audio file can meet the user's needs, thereby improving the user experience. Optionally, the audio file generation device can be a terminal device (for example, a desktop computer, a laptop computer, a smart phone, a smart speaker, a smart wearable device, a tablet computer, or a vehicle-mounted terminal, etc.), or a server (the server can be a single computer, or a cluster composed of multiple computers, or a cloud server that can elastically adjust computing resources through cloud technology), and this application does not limit this. It should be understood that when the audio file generation device is a terminal device, the audio file generation device can directly support the user to input multi-modal reference information or related instructions by interacting with the user. When the audio file generation device is a server, the audio file generation device can support the user to input multi-modal reference information or related instructions by interacting with the terminal device held by the user. As Figure 1 shown, the audio file generation method can specifically include the following steps:

[0037] Step S100, receive an audio generation instruction.

[0038] Specifically, the audio file generation device can receive an audio generation instruction. Among them, the audio generation instruction can be issued by the user and is used to instruct the audio file generation device to generate an audio file.

[0039] Optionally, the audio generation instruction can be in text form or voice form, and this application does not limit this. Schematically, the audio generation instruction in text form can be "Please help me generate a piece of music related to this video" and "Please help me generate a song according to this painting and this text", etc. The audio generation instruction in voice form can be "Please help me take a picture of the setting sun in front and then create a song for me", "My friend's birthday is today, please help me create a song to wish her a happy birthday and sing it in my voice", "This song is very nice, but I want to hear it sung in another style", "Please help expand this melody into a song", and "Please help me adjust the style of this song". It should be understood that the content of the audio generation instruction in each form given above is only for illustration, and in the actual application process, this application does not limit the content of the audio generation instruction itself.

[0040] Step S200, obtain multi-modal reference information.

[0041] Specifically, after receiving an audio generation instruction, the audio file generation device can obtain multi-modal reference information. Among them, the multi-modal reference information can specifically be reference information of multiple different information types. It should be understood that the multi-modal reference information obtained by the audio file generation device can be used as a reference basis for generating an audio file, so that it can generate an audio file that meets the user's needs in subsequent steps.

[0042] Optionally, in this embodiment, the information types of the reference information included in the multi-modal reference information obtained by the audio file generation device can be at least one or a combination of the following: text, audio, picture, video, geographical location, emotional information.

[0043] Optionally, in this embodiment, the multi-modal reference information obtained by the audio file generation device may include the audio generation instruction itself, the context information corresponding to the audio generation instruction, and / or the relevant information collected under the user's indication and settings. Specifically, in step S200, the audio file generation device may obtain the audio generation instruction itself as the multi-modal reference information. And / or, the audio file generation device may obtain the context information corresponding to the audio generation instruction as the multi-modal reference information. For example, when the audio generation instruction is "Please help me generate a piece of music related to this video", the audio file generation device may take the video input by the user as the multi-modal reference information. When the audio generation instruction is "Please help me generate a song based on this painting and this text", the audio file generation device may take the image and text description input by the user as the multi-modal reference information. And / or, the audio file generation device may call the corresponding sensor component according to the audio generation instruction to collect information to obtain multi-modal reference information. For example, when the audio generation instruction is "Please help me take a picture of the sunset in front and then create a song for me", the audio file generation device may call the image acquisition component to collect an image containing the sunset and take the collected image as the multi-modal reference information. When the audio generation instruction is "Please help me expand this melody into a song", the audio file generation device may call the audio acquisition component to collect the melody signal and take the collected melody signal as the multi-modal reference information. And / or, the audio file generation device may call the corresponding sensor component according to the preset reference information collection requirement to collect information to obtain multi-modal reference information. Among them, the reference information collection requirement may be preset by the user, and it may be used to represent the user's collection requirement for the reference information of the corresponding information type. For example, when the reference information collection requirement includes the collection requirement for emotion information, the audio file generation device may call the corresponding emotion analysis component to analyze and obtain the user's current emotion, and take the analyzed user's current emotion information as the multi-modal reference information. When the reference information collection requirement includes the collection requirement for geographical location, the audio file generation device may call the corresponding location collection component to collect the user's current location and take the collected user's current geographical location information as the multi-modal reference information.

[0044] It should be understood that the content included in the above multi-modal reference information is only for illustration. In the actual application process, the content included in the multi-modal reference information can be specifically set and adjusted by relevant personnel according to actual needs, and this application does not limit this.

[0045] Step S300, call at least one target large language model to perform intention recognition on the multi-modal reference information to determine the target audio generation requirement.

[0046] Specifically, after obtaining the multimodal reference information, the audio file generation device can call at least one target large language model to perform intent recognition on the multimodal reference information to determine the target audio generation requirements. Among them, the target large language model can be a pre-determined large language model that matches the multimodal reference information. It should be understood that different large language models support different types of information for processing. Here, the large language model that matches the multimodal reference information specifically refers to a large language model whose supported information type can cover the information types of the reference information included in the multimodal reference information. The target audio generation requirements can be used to represent the user's requirements for generating the audio file.

[0047] Figure 2 The flowchart of the method for determining the target large language model according to the embodiment of the present invention. It should be understood that before calling the target large language model to perform intent recognition on the multimodal reference information to determine the target audio generation requirements, the audio file generation device can determine the target large language model that matches the multimodal reference information by executing the method for determining the target large language model as shown in Figure 2 As shown. The method for determining the target large language model can specifically include the following steps: Figure 2 As shown, the method for determining the target large language model can specifically include the following steps:

[0048] Step S311: Determine the information types of the reference information included in the multimodal reference information.

[0049] Specifically, the audio file generation device can determine the information types of the reference information included in the obtained multimodal reference information.

[0050] Step S312: Determine at least one target large language model that matches the multimodal reference information from multiple pre-set large language models according to the information types.

[0051] Specifically, after the information types of the reference information included in the multimodal reference information, the audio file generation device can determine at least one target large language model that matches from multiple pre-set large language models. Among them, the multiple pre-set large language models can be pre-set by relevant personnel and can be called by the audio file generation device. It should be understood that the present application does not limit the set pre-set large language models.

[0052] Optionally, in this embodiment, the large language model can be deployed on the local side or on the online server. Among them, when the large language model is deployed on the local side, the audio file generation device can directly call the large language model locally. When the large language model is deployed on the online server, the audio file generation device calls the large language model by interacting with the online server.

[0053] It should be understood that in this embodiment, when the number of target large language models is one, the audio file generation device can call the target large language model according to the multimodal reference information and directly determine the target audio generation requirement based on the call result of the target large language model. When the number of target large language models is multiple, the audio file generation device can call multiple target large language models respectively according to the multimodal reference information and integrate the call results of the multiple target large language models to determine the target audio generation requirement.

[0054] Figure 3 It is a flowchart of the intention recognition method according to an embodiment of the present invention. It should be understood that by executing the intention recognition method as Figure 3 shown, the audio file generation device can call multiple target large language models respectively according to the multimodal reference information and integrate the call results of the multiple target large language models to determine the target audio generation requirement. As Figure 3 shown, the intention recognition method may specifically include the following steps:

[0055] Step S321, in response to detecting that the number of the target large language models is multiple, call each of the target large language models to perform intention recognition on the multimodal reference information to determine multiple candidate audio generation requirements.

[0056] Specifically, when detecting that the number of target large language models is multiple, the audio file generation device can call each target large language model to perform intention recognition on the multimodal reference information to determine multiple candidate audio generation requirements. Among them, each candidate audio generation requirement may be the user's generation requirement for the audio file determined by each target large language model through intention recognition of the multimodal reference information.

[0057] Optionally, in this embodiment, whether it is to call the target large language model according to the multimodal reference information to determine the target audio generation requirement or to call the target large language model according to the multimodal reference information to determine the candidate audio generation requirement, the audio file generation device can first construct an intention recognition prompt statement according to the multimodal reference information, and then input the intention recognition prompt statement of the intention recognition template into the target large language model, so as to determine the target audio generation requirement or the candidate audio generation requirement through the target large language model. Among them, the intention recognition prompt statement may be a text description, which is used to guide the target large language model to perform intention recognition on the multimodal reference information and make it output information that meets the expectations.

[0058] Further optionally, the intent recognition prompt statement can be constructed by the audio file generation device by filling the multimodal reference information into the corresponding intent recognition prompt template. Among them, the intent recognition prompt template can be a statement template prepared in advance by relevant personnel. Optionally, the intent recognition prompt template can at least include an intent recognition instruction and an intent recognition requirement. The intent recognition instruction can be used to indicate the instruction for the large language model to complete the intent recognition task. For example, the intent recognition instruction can be "Please determine the generation requirements of the user for the audio file according to the multimodal reference information input by the user". The intent recognition requirement can be used to assist the large language model in intent recognition and instruct it to output the recognition result in a corresponding format. For example, the intent recognition requirement can be "Please analyze and give the user's main or emotional requirements for the audio file to be generated, give the creation background, usage occasion and purpose of the audio file to be generated. If there is image information in the multimodal reference information, please give a description of the image content of the image information. If there is video information in the multimodal reference information, please give a description of the video content of the video information. If there is audio information in the multimodal reference information, please give a description of the audio content of the audio information".

[0059] Exemplarily, the intent recognition prompt statement constructed by the audio file generation device can be: [Please determine the generation requirements of the user for the audio file according to the multimodal reference information input by the user. Among them, the multimodal reference information is: "My friend's birthday today, please help me create a song to wish her a happy birthday and sing it to her with my voice". Please analyze and give the user's main or emotional requirements for the audio file to be generated, give the creation background, usage occasion and purpose of the audio file to be generated. If there is image information in the multimodal reference information, please give a description of the image content of the image information. If there is video information in the multimodal reference information, please give a description of the video content of the video information. If there is audio information in the multimodal reference information, please give a description of the audio content of the audio information].

[0060] It should be understood that the content included in the above-mentioned intent recognition prompt template and the intent recognition prompt statement constructed according to this template are only for illustration, and in the actual application process, the content included in the intent recognition prompt template and the intent recognition prompt statement constructed according to this template are not limited to this.

[0061] It should be understood that before filling the multimodal reference information into the corresponding intention recognition prompt template to construct an intention recognition prompt statement, in order to improve the intention recognition accuracy of the target large language model, the audio file generation device can also preprocess (i.e., supplement, modify, and adjust) the multimodal reference information so that the preprocessed multimodal reference information can meet the input requirements of the target large language model or contain richer information content. For example, for the geographical location information in the multimodal reference information, the audio file generation device can determine and supplement the point-of-interest information related to the geographical location through location retrieval, thereby enriching the information content contained in the multimodal reference information. Another example is that for the text information with errors in the multimodal reference information, the audio file generation device can modify the text information to avoid the impact of the incorrect text content on the intention recognition result of the target large language model. Still another example is that for the picture information or video information in the multimodal reference information, the audio file generation device can correspondingly adjust the format of the picture information or video information so that it meets the input requirements of the target large language model for images or videos.

[0062] Step S322: Determine the target audio generation requirement according to the model weights of each of the target large language models and the multiple candidate audio generation requirements.

[0063] Specifically, after calling multiple target large language models according to the multimodal reference information to obtain multiple candidate audio generation requirements, the audio file generation device can determine the target audio generation requirement according to the model weights of each target large language model and the multiple candidate audio generation requirements. Among them, the model weight of the target large language model can be determined by the audio file generation device, which can be used to represent the processing ability of the target large language model for the multimodal reference information. It should be noted that in this embodiment, although multiple target large language models all support processing multimodal reference information, due to the influence of the model's own performance, the processing abilities of different target large language models for the multimodal reference information may be different. By determining the target audio generation requirement according to the model weights of multiple target large language models and the multiple candidate audio generation requirements, this embodiment can ensure that the finally determined target audio generation requirement can best meet the user's original requirement.

[0064] Optionally, in step S322, the model weights of each target large language model can be pre-set by relevant personnel. Or, the model weights of each target large language model can also be determined by the audio file generation device according to the reference information ratio of each information type in the multimodal reference information and the reasoning ability of each target large language model for the corresponding information type. The present application does not limit the determination method of the model weights of each target large language model.

[0065] Optionally, in step S322, the process of determining the target audio generation requirement based on the model weights of each target large language model and multiple candidate audio generation requirements can also be implemented by the audio file generation device by invoking the corresponding large language model. Specifically, the audio file generation device can fill in the model weights of each target large language model and multiple candidate audio generation requirements into the requirement integration prompt template to obtain a requirement integration prompt statement. Furthermore, the audio file generation device can input the requirement integration prompt statement into the large language model, so as to integrate multiple candidate audio generation requirements through the large language model to determine the target audio generation requirement. It should be understood that the requirement integration prompt template can be specifically set in advance by relevant personnel, and the present application does not limit it.

[0066] Step S400: Determine the corresponding song description and sheet music description according to the target audio generation requirement.

[0067] Specifically, after determining the target audio generation requirement, the audio file generation device can determine the corresponding song description and sheet music description according to the target audio generation requirement. Among them, the song description and sheet music description can be understood as descriptions of the song and sheet music at the feature level. It should be understood that in this embodiment, the song description can be used as a reference for generating the vocal part in the finally generated audio file. The song description can specifically include relevant information describing various aspects such as the pitch, rhythm, melody, harmony, emotional atmosphere, style and genre, lyrics content, and singing style of the vocal part. The song description can be used as a reference for generating the background music part in the finally generated audio file. The sheet music description can specifically include relevant information describing various aspects such as the pitch, rhythm, melody, harmony, instrument arrangement, and musical form structure of the background music part.

[0068] Figure 4 It is a flowchart of the description generation method according to an embodiment of the present invention. It should be understood that by executing the method as Figure 4 shown, the audio file generation device can determine the corresponding song description and sheet music description according to the target audio generation requirement, that is, implement the above step S400. As Figure 4 shown, the description generation method can specifically include the following steps:

[0069] Step S410: Perform feature matching in the song feature library according to the target audio generation requirement to determine the song description.

[0070] Specifically, the audio file generation device can perform feature matching in the song feature library according to the target audio generation requirement to determine the song description.

[0071] Optionally, the song feature library may include a plurality of preset song features. Each song feature may respectively have corresponding feature application information. In step S410, the feature matching in the song feature library according to the target audio generation requirement may specifically be to match the target audio generation requirement with the feature application information of each song feature in the song feature library. It should be understood that at least one song feature matched by the audio file generation device can be used as the song description. Among them, the feature application information of the song feature can be used to characterize the applicable range of the song feature. Thus, this embodiment can determine the song description that matches the target audio generation requirement.

[0072] Further optionally, the matching of the target audio generation requirement with the feature application information of each song feature can be implemented by the audio file generation device by invoking a large language model. Specifically, the audio file generation device can construct a corresponding song description determination prompt statement according to the target audio generation requirement and the feature application information of each song feature, and input the song description determination prompt statement into the large language model to determine at least one song feature that matches the target audio generation requirement among the plurality of song features through the large language model, so as to determine the song description. Among them, the song description determination prompt statement can be a text description, which is used to guide the large language model to perform feature matching and make it output information that meets the expectations.

[0073] Further optionally, the song description determination prompt statement can be constructed by the audio file generation device by filling the target audio generation requirement and the feature application information of each song feature into the corresponding song description determination prompt template. Among them, the construction of the song description determination prompt template can be a statement template prepared in advance by relevant personnel. Optionally, the intention recognition prompt template may at least include a feature matching instruction. The feature matching instruction can be used to specify the instruction for the large language model to complete the matching task. For example, the intention recognition instruction can be "Please determine at least one song feature that matches the target audio generation requirement according to the target audio generation requirement and the feature application information of each song feature".

[0074] Exemplarily, the song description determination prompt statement constructed by the audio file generation device can be: [Please determine at least one song feature that matches the target audio generation requirement based on the target audio generation requirement and the feature applicability information of each song feature. Among them, the target audio generation requirement is: "This song is created to celebrate a friend's birthday, so the theme is the best wishes for the friend. The song should be full of warm, happy and friendly emotions, expressing the best wishes for the friend's birthday." Song features: high pitch, feature applicability information: often used to express lively and exciting emotions, or to create a tense or urgent feeling; Song feature: low pitch, feature applicability information: often used to express deep and steady emotions, such as sadness, nostalgia or solemn emotions, or to create a mysterious or majestic atmosphere; Song feature: fast rhythm, feature applicability information: often used to express energetic and exciting emotions, or can be used to create a tense or urgent feeling; Song feature: slow rhythm, feature applicability information: often used to express deep and lyrical emotions, or to help people relax.]

[0075] It should be understood that the content included in the song description determination prompt template given above and the song description determination prompt statement constructed according to this template are only for illustration. In actual application, the content included in the song description determination prompt template and the song description determination prompt statement constructed according to this template are not limited to this.

[0076] Step S420: Perform feature matching in the score feature library according to the target audio generation requirement to determine the score description.

[0077] Specifically, the audio file generation device can perform feature matching in the score feature library according to the target audio generation requirement to determine the score description.

[0078] Optionally, the score feature library may include a plurality of preset score features. Each score feature may have corresponding feature applicability information. In step S420, performing feature matching in the score feature library according to the target audio generation requirement may specifically be to match the target audio generation requirement with the feature applicability information of each score feature in the score feature library. It should be understood that at least one score feature matched by the audio file generation device can be used as the score description. Among them, the feature applicability information of the score feature can be used to characterize the applicable range of the score feature. Thus, this embodiment can determine the score description that matches the target audio generation requirement.

[0079] It should be understood that the matching of the target audio generation requirement and the feature applicability information of each score feature can also be implemented by the audio file generation device by calling a large language model. Since the implementation process is similar, it will not be elaborated here too much.

[0080] Step S500: Generate a target audio file according to the song description and the musical score description.

[0081] Specifically, after determining the song description and the musical score description, the audio file generation device can generate a target audio file according to the song description and the musical score description.

[0082] Figure 5 It is a flowchart of the audio file generation method according to an embodiment of the present invention. It should be understood that by executing the audio file generation method as Figure 5 shown, the audio file generation device can generate a target audio file according to the song description and the musical score description, that is, implement the above step S500. As Figure 5 shown, the audio file generation method may specifically include the following steps:

[0083] Step S510: Generate a vocal signal according to the song description.

[0084] Specifically, the audio file generation device can generate a vocal signal according to the song description. Among them, the vocal signal is used to constitute the vocal part in the finally generated audio file.

[0085] Figure 6 It is a flowchart of the vocal signal generation method according to an embodiment of the present invention. It should be understood that by executing the vocal signal generation method as Figure 6 shown, the audio file generation device can generate a vocal signal according to the song description.

[0086] As Figure 6 shown, the vocal signal generation method may specifically include the following steps:

[0087] Step S511: Generate lyric information according to the song description.

[0088] Specifically, the audio file generation device can generate lyric information according to the song description.

[0089] Optionally, in step S511, the process of generating lyric information according to the song description can also be implemented by the audio file generation device by calling the corresponding large language model. Specifically, the audio file generation device can fill the song description into the lyric generation prompt template to obtain a lyric generation prompt statement. Further, the audio file generation device can input the lyric generation prompt statement into the large language model, so as to generate lyric information through the large language model. It should be understood that the lyric generation prompt template can be specifically set by relevant personnel in advance, and the present application does not limit it.

[0090] Step S512: Generate the vocal signal according to the lyric information.

[0091] Specifically, after generating the lyrics information, the audio file generation device can generate a vocal signal based on the lyrics information.

[0092] Optionally, some users may hope that the vocal part in the finally generated audio file is generated based on the voiceprint information of a specific character. In this regard, when performing step S512, the audio file generation device can first detect whether the target audio generation requirement includes a voiceprint usage requirement for the target character. If it does, the audio file generation device can obtain the voiceprint information of the target character according to the voiceprint usage requirement, and then generate a vocal signal based on the lyrics information and the voiceprint information. Here, the target character can be the service acquisition object of the audio file generation service (i.e., the user himself), or a relevant public figure or virtual character selected by the user and pre-authorized, etc. The present application does not limit this. It should be understood that the audio file generation device can be set with multiple default preset characters. When the target audio generation requirement does not include a voiceprint usage requirement for the target character, the audio file generation device can select the voiceprint information of a suitable preset character according to the song description to generate a vocal signal.

[0093] Step S520: Generate a multi-track signal according to the score description.

[0094] Specifically, the audio file generation device can generate a multi-track signal according to the score description. Among them, the multi-track signal is used to constitute the background music part in the finally generated audio file. It should be understood that the multi-track signal is a signal composed of multiple tracks, and different tracks in the multi-track signal can correspond to different sound sources (for example, different musical instruments).

[0095] Optionally, in step S520, the process of generating a multi-track signal according to the score description can also be implemented by the audio file generation device by calling the corresponding large language model. Specifically, the audio file generation device can fill the score description into the background music generation prompt template to obtain a background music generation prompt statement. Furthermore, the audio file generation device can input the background music generation prompt statement into the large language model, so as to generate a multi-track signal through this large language model. It should be understood that the background music generation prompt template can be specifically set by relevant personnel in advance, and the present application does not limit it.

[0096] Step S530: Fit at least the vocal signal and the multi-track signal to generate the target audio file.

[0097] Specifically, after generating the vocal signal and the multi-track signal, the audio file generation device can fit at least the vocal signal and the multi-track signal to generate the target audio file.

[0098] Optionally, in addition to supporting the user to generate files in pure audio format, the audio file generation device can also support the user to generate files in a composite format (e.g., mp3 format or mp4 format) that contains audio information and other information. Among them, the other information can specifically include cover information and lyric information. Specifically, in step S530, the audio file generation device can also determine at least one song picture, and fit the vocal signal, multi-track signal, lyric information, and at least one song picture to generate a target audio file. It should be understood that the song picture can be used as the cover of the target audio file. The song picture can be specified by the user or generated by the audio file generation device according to relevant information (such as song description, lyric information, and score description, etc.), and this application does not limit this.

[0099] Optionally, in this embodiment, the generated vocal signal, multi-track signal, and lyric information may have corresponding timestamp information, and the audio file generation device can implement the fitting of the vocal signal, multi-track signal, and lyric information according to the timestamp information.

[0100] In the embodiment of the present invention, after receiving an audio generation instruction, multi-modal reference information is obtained, and at least one target large language model that matches the multi-modal reference information is called to perform intent recognition on the multi-modal reference information to determine the target audio generation requirement. Then, the corresponding song description and score description are determined according to the target audio generation requirement, and further, a target audio file is generated according to the song description and score description. Among them, the target audio generation requirement is used to represent the generation requirement for the audio file. Thus, by supporting the input of multi-modal reference information and generating a target audio file according to the multi-modal reference information, the embodiment of the present invention can ensure that the generated audio file can meet the user's needs, thereby improving the user experience.

[0101] Figure 7 It is a schematic diagram of the audio file generation device according to the embodiment of the present invention. As Figure 7 shown, the audio file generation device according to the embodiment of the present invention includes an instruction receiving unit 71, a multi-modal reference information obtaining unit 72, an intent recognition unit 73, a description information determining unit 74, and an audio file generating unit 75.

[0102] Specifically, the instruction receiving unit 71 is used to receive an audio generation instruction;

[0103] The multi-modal reference information obtaining unit 72 is used to obtain multi-modal reference information;

[0104] The intention recognition unit 73 is used to call at least one target large language model to perform intention recognition on the multimodal reference information to determine the target audio generation requirement, where the target large language model matches the multimodal reference information, and the target audio generation requirement is used to characterize the generation requirement for the audio file;

[0105] The description information determination unit 74 is used to determine the corresponding song description and music score description according to the target audio generation requirement;

[0106] The audio file generation unit 75 is used to generate a target audio file according to the song description and the music score description.

[0107] In the embodiment of the present invention, after receiving an audio generation instruction, multimodal reference information is obtained, and at least one target large language model that matches the multimodal reference information is called to perform intention recognition on the multimodal reference information to determine the target audio generation requirement. Then, the corresponding song description and music score description are determined according to the target audio generation requirement, and further, a target audio file is generated according to the song description and the music score description. Wherein, the target audio generation requirement is used to characterize the generation requirement for the audio file. Thus, by supporting the input of multimodal reference information and generating a target audio file according to the multimodal reference information, the embodiment of the present invention can ensure that the generated audio file can meet the user's needs, thereby improving the user experience.

[0108] Figure 8 is a schematic diagram of the electronic device in the embodiment of the present invention. The electronic device may specifically be the audio file generation device in the above embodiment. As Figure 8 shown, the electronic device includes at least one processor 81; and, a memory 82 communicatively connected to at least one processor 81; and, a communication component 83 communicatively connected to the scanning device, and the communication component 83 receives and sends data under the control of the processor 81; wherein, the memory 82 stores instructions executable by at least one processor 81, and the instructions are executed by at least one processor 81 to implement the above audio file generation method.

[0109] Specifically, the electronic device includes one or more processors 81 and a memory 82, Figure 8 taking one processor 81 as an example. The processor 81 and the memory 82 may be connected through a bus or other means, Figure 8 taking the connection through a bus as an example. As a non-volatile computer-readable storage medium, the memory 82 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor 81 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory 82, that is, implements the above audio file generation method.

[0110] The memory 82 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store an option list, etc. In addition, the memory 82 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 82 may optionally include a memory remotely provided relative to the processor 81, and these remote memories can be connected to external devices through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0111] One or more modules are stored in the memory 82 and, when executed by one or more processors 81, execute the audio file generation method in any of the above method embodiments.

[0112] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects of the executed method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0113] In an embodiment of the present invention, after receiving an audio generation instruction, multimodal reference information is obtained, and at least one target large language model matching the multimodal reference information is called to perform intent recognition on the multimodal reference information to determine the target audio generation requirement. Then, a corresponding song description and score description are determined according to the target audio generation requirement, and further a target audio file is generated according to the song description and score description. Among them, the target audio generation requirement is used to represent the generation requirement for the audio file. Thus, by supporting the input of multimodal reference information and generating a target audio file according to the multimodal reference information, the embodiment of the present invention can ensure that the generated audio file can meet the user's needs, thereby improving the user experience.

[0114] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute some or all of the above method embodiments.

[0115] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0116] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. An audio file generation method, characterized in that The method includes: Receiving an audio generation instruction; Obtaining multimodal reference information; Invoking at least one target large language model to perform intent recognition on the multimodal reference information to determine a target audio generation requirement, wherein the target large language model matches the multimodal reference information, and the target audio generation requirement is used to characterize the generation requirement for an audio file; Determining a corresponding song description and musical score description according to the target audio generation requirement; Generating a target audio file according to the song description and the musical score description.

2. The method according to claim 1, wherein The generating the target audio file according to the song description and the musical score description includes: Generating a vocal signal according to the song description; Generating a multi-track signal according to the musical score description; At least fitting the vocal signal and the multi-track signal to generate the target audio file.

3. The method according to claim 2, characterized in that, The generating the vocal signal according to the song description includes: Generating lyric information according to the song description; Generating the vocal signal according to the lyric information.

4. The method according to claim 3, wherein The generating the vocal signal according to the lyric information includes: In response to detecting that the target audio generation requirement includes a voiceprint usage requirement for a target role, obtaining the voiceprint information of the target role according to the voiceprint usage requirement, wherein the target role includes a service acquisition object of an audio file generation service; Generating the vocal signal according to the lyric information and the voiceprint information.

5. The method according to claim 3, wherein The at least fitting the vocal signal and the multi-track signal to generate the target audio file includes: Determining at least one song picture; Fitting the vocal signal, the multi-track signal, the lyric information, and the at least one song picture to generate the target audio file.

6. The method according to claim 1, characterized in that, Before invoking at least one target large language model to perform intent recognition on the multimodal reference information to determine a target audio generation requirement, the method further includes: Determining the information type of the reference information included in the multimodal reference information; Determining at least one target large language model that matches in a plurality of preset large language models according to the information type.

7. The method according to claim 1, characterized in that The invoking at least one target large language model to perform intent recognition on the multimodal reference information to determine a target audio generation requirement includes: In response to detecting that the number of the target large language models is multiple, invoking each of the target large language models to perform intent recognition on the multimodal reference information respectively to determine a plurality of candidate audio generation requirements; Determining the target audio generation requirement according to the model weights of each of the target large language models and the plurality of candidate audio generation requirements.

8. The method according to claim 1, wherein The obtaining the multimodal reference information includes: Taking the obtained audio generation instruction as the multimodal reference information; and / or Taking the context information corresponding to the audio generation instruction as the multimodal reference information; and / or Invoking a sensor component to collect information according to the audio generation instruction to obtain the multimodal reference information; and / or Invoking a sensor component to collect information according to a preset reference information collection requirement to obtain the multimodal reference information.

9. The method according to claim 1, characterized in that, The determining a corresponding song description and musical score description according to the target audio generation requirement includes: Perform feature matching in the song feature library according to the target audio generation requirement to determine the song description; Perform feature matching in the music score feature library according to the target audio generation requirement to determine the music score description.

10. The method according to claim 1, wherein The information types of the reference information included in the multimodal reference information are at least one or a combination of the following: text, audio, picture, video, geographical location, and emotion information.

11. An audio file generation device, characterized in that, The device includes: An instruction receiving unit, configured to receive an audio generation instruction; A multimodal reference information obtaining unit, configured to obtain multimodal reference information; An intention recognition unit, configured to call at least one target large language model to perform intention recognition on the multimodal reference information to determine a target audio generation requirement, where the target large language model matches the multimodal reference information, and the target audio generation requirement is used to represent the generation requirement for an audio file; A description information determination unit, configured to determine a corresponding song description and music score description according to the target audio generation requirement; An audio file generation unit, configured to generate a target audio file according to the song description and the music score description.

12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method according to any one of claims 1-10.

13. An electronic device, characterized in that, The device includes: A memory, configured to store one or more computer program instructions; A processor, where the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-10.

14. A computer program product, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1-10.