Music playing method and device, model training method and device, network equipment and medium
By obtaining the user's current status description information and using the pre-trained model to generate ambient sound effects and music synthesis, the problem of the mismatch between ambient sound effects and the user's emotional state and environmental needs in the existing technology is solved, and a personalized music experience is achieved.
Patent Information
- Application Number
- CN202510811864.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology pre-makes and adds ambient sound effects when playing music, resulting in a deviation between the produced ambient sound effects and the user's current emotional state and environmental requirements, and cannot meet the needs of personalized music experience.
By obtaining the user's current status description information, the pre-trained model is used to generate an ambient sound effect that conforms to the current status description information, and then synthesized with the playing music in the music library to generate a second playing music that is adapted to the user's current situation.
It realizes the real-time generation of ambient sound effects based on the user's current emotional state and environmental requirements, meets the needs of personalized music experience, and enhances the overall sense of atmosphere.
Smart Images

Figure CN120673728A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent audio technology, and in particular to a music playback method, model training method, device, network equipment and medium. Background Art
[0002] Currently, adding sound effects to music can enhance the sense of ambiance, allowing users to relax while enjoying music. Existing technologies use a method of adding ambient sound effects during music production, pre-fabricating and adding ambient sound effects to the music being played. However, this method cannot replace the ambient sound effects after the music is produced. Therefore, the produced ambient sound effects may deviate from the user's current emotional state and environmental needs, resulting in a mismatch between the two and failing to meet the demand for a more personalized music experience. Summary of the Invention
[0003] The purpose of the technical solution of this application is to provide a music playback method, model training method, device, network equipment and medium, which can adaptively generate matching playback sound effects according to the user's current emotional state and / or environmental requirements, so as to meet the personalized music experience needs.
[0004] One embodiment of the present application provides a music playing method, which includes:
[0005] Get the user's current status description information;
[0006] The current state description information is used as an input of a first model to obtain an ambient sound effect output by the first model; the first model is used to generate an ambient sound effect that conforms to the current state description information based on the current state description information;
[0007] The first play music in the music library is synthesized with the ambient sound effect to generate the second play music.
[0008] Optionally, the music playing method further comprises:
[0009] Get the beats per minute (BPM) of the ambient sound effect;
[0010] The first played music is selected from the music library according to the beats per minute of the ambient sound effect; wherein the beats per minute of the first played music matches the beats per minute of the ambient sound effect.
[0011] Optionally, in the music playing method, the current state description information includes one or more of emotion description information, environment description information and sound effect description information.
[0012] Optionally, the music playing method further comprises:
[0013] In a case where the acquired current state description information is in voice format, converting the current state description information in voice format into text format;
[0014] The step of using the current state description information as input to the first model includes:
[0015] The current state description information in text format is used as input of the first model.
[0016] Optionally, the music playing method further comprises:
[0017] Obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0018] encoding the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encoding the state description information to obtain state features of the state description information;
[0019] The state feature is used as input and the data feature is used as output to perform model training to obtain the first model.
[0020] Optionally, the music playing method further comprises:
[0021] Encoding the current state description information to obtain a current state feature corresponding to the current state description information;
[0022] The step of using the current state description information as input to the first model includes:
[0023] The current state feature is used as input of the first model.
[0024] Optionally, the music playing method, wherein the state feature is used as input and the data feature is used as output to perform model training, comprises:
[0025] The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
[0026] One embodiment of the present application further provides a model training method, which includes:
[0027] Obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0028] encoding the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encoding the state description information to obtain state features of the state description information;
[0029] The state feature is used as input and the data feature is used as output to perform model training to obtain a first model; wherein the first model is used to generate an ambient sound effect that conforms to the current state description information of the user based on the current state description information.
[0030] Optionally, the model training method, wherein the state feature is used as input and the data feature is used as output, performs model training, comprising:
[0031] The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
[0032] One embodiment of the present application further provides a music playing device, comprising:
[0033] Information acquisition module, used to obtain the user's current status description information;
[0034] a processing module, configured to use the current state description information as input to a first model to obtain an ambient sound effect output by the first model; the first model is configured to generate an ambient sound effect that conforms to the current state description information based on the current state description information;
[0035] The synthesis module is used to synthesize the first play music in the music library and the ambient sound effect to generate the second play music.
[0036] One embodiment of the present application further provides a model training device, which includes:
[0037] An information acquisition module, configured to obtain a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0038] an encoding module, configured to encode the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and to encode the state description information to obtain state features of the state description information;
[0039] The model training module is used to perform model training using the state feature as input and the data feature as output to obtain a first model; wherein the first model is used to generate an ambient sound effect that conforms to the current state description information of the user based on the current state description information.
[0040] One embodiment of the present application also provides a network device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, it implements the music playback method as described in any one of the above items, or implements the model training method as described in any one of the above items.
[0041] One embodiment of the present application also provides a readable storage medium, wherein a program is stored on the readable storage medium, and when the program is executed by a processor, the steps in the music playback method as described in any one of the above items are implemented, or the steps in the model training method as described in any one of the above items are implemented.
[0042] One embodiment of the present application also provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, implement the steps in the music playback method as described in any of the above items, or implement the steps in the model training method as described in any of the above items.
[0043] The above technical solutions of the embodiments of the present application have at least the following beneficial effects:
[0044] An embodiment of the present application provides a music playback method, which inputs the user's current state description information into a first model obtained through pre-training, obtains an ambient sound effect that conforms to the current state description information based on the output of the first model, and synthesizes the ambient sound effect with the first playback music in the music library to obtain a second playback music, so that the second playback music played can adapt to the user's current situational state and conform to the user's current emotional state and / or environmental requirements, so as to enhance the overall atmosphere and meet the personalized music experience needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flowchart of the music playing method described in an embodiment of the present application;
[0046] Figure 2 This is a flow chart of one embodiment of the method described in the examples of the present application;
[0047] Figure 3 This is a flow chart of another embodiment of the method described in the examples of the present application;
[0048] Figure 4 This is a schematic diagram of a first application interface using the method described in one embodiment of the present application;
[0049] Figure 5 This is a schematic diagram of a second application interface using the method described in one embodiment of the present application;
[0050] Figure 6 This is a flow chart of the model training method described in the embodiment of the present application;
[0051] Figure 7 This is a structural diagram of the music playing device according to an embodiment of the present application;
[0052] Figure 8 This is a structural diagram of the model training device described in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the technical problems, technical solutions and advantages to be solved by this application clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0056] In order to solve the problem that the existing technology adopts the method of adding ambient sound effects when playing music, the ambient sound effects are pre-made and added to the playing music. The produced ambient sound effects will deviate from the user's current emotional state and environmental requirements, and there is a mismatch between the two. The embodiment of the present application provides a music playing method and a model training method. By inputting the user's current state description information into a pre-trained first model, the ambient sound effect that meets the current state description information is obtained according to the output of the first model, and the ambient sound effect is synthesized with the first played music in the music library to obtain the second played music, so that the second played music can adapt to the user's current situational state to enhance the overall atmosphere and meet the personalized music experience needs.
[0057] like Figure 1 As shown, the music playing method described in the embodiment of the present application includes:
[0058] S110, obtaining the user's current status description information;
[0059] S120: Using the current state description information as input to a first model to obtain an ambient sound effect output by the first model; the first model is configured to generate an ambient sound effect that conforms to the current state description information based on the current state description information;
[0060] S130: synthesize the first play music in the music library and the ambient sound effect to generate second play music.
[0061] By adopting the music playback method described in this embodiment, after obtaining the user's current state description information, the first model obtained in advance can be used to automatically generate ambient sound effects that conform to the current state description information in real time. Music playback based on the ambient sound effects can meet the user's current state and / or environmental requirements and provide a more personalized and intimate music experience.
[0062] In one embodiment of the present application, optionally, the current state description information includes but is not limited to only one or more of emotion description information, environment description information and sound effect description information.
[0063] Optionally, the emotion description information is information used to describe the user's current emotional state, such as happy, excited, sad, etc.; the environment description information is information used to describe the user's current environment state, such as sunny, cloudy, dark and bright, etc.; the sound effect description information is information used to describe the sound effect required by the user, such as cheerful, gentle and sad, etc.
[0064] Optionally, step S110, obtaining the user's current status description information includes:
[0065] Get the current status description information entered by the user.
[0066] The current status description information input by the user may be in voice format or text format.
[0067] Optionally, the method further includes:
[0068] In the case where the input current state description information is in voice format, the current state description information in voice format is converted into text format; wherein, in step S120, the current state description information is used as input of the first model, including:
[0069] The current state description information in text format is used as input of the first model.
[0070] Optionally, the automatic speech recognition (ASR) technology may be used to convert the current status description information in voice format into text format.
[0071] In one embodiment of the present application, optionally, the method further includes:
[0072] The current state description information input by the user is parsed to obtain one or more of emotion description information, environment description information, and sound effect description information.
[0073] In one embodiment of the present application, optionally, the first model is obtained through model training after collecting multiple training ambient sound effects and state description information corresponding to the multiple training ambient sound effects.
[0074] By using the first model obtained by pre-training, after the user's current state description information is input into the first model, the first model can output an ambient sound effect that conforms to the current state description information.
[0075] In an embodiment of the present application, optionally, the method further includes:
[0076] Get the beats per minute (BPM) of the ambient sound effect;
[0077] The first played music is selected from the music library according to the beats per minute of the ambient sound effect; wherein the beats per minute of the first played music matches the beats per minute of the ambient sound effect.
[0078] Using this implementation, after obtaining the ambient sound effect corresponding to the current state description information, the first playback music that matches the beats per minute of the ambient sound effect is selected from the music library according to the BPM of the ambient sound effect, so that the rhythm of the selected first playback unit is consistent with the rhythm of the ambient sound effect, to ensure that when the first playback music is synthesized with the ambient sound effect, the playback effect of the generated second playback music is smoother.
[0079] Optionally, when the difference between the beats per minute of the first played music and the beats per minute of the ambient sound effect is less than or equal to a preset value, it is determined that the beats per minute of the first played music and the beats per minute of the ambient sound effect match.
[0080] In some embodiments, the corresponding beats per minute can be obtained by detecting the number of times the ambient sound effects and the accents appear per minute in the first played music. However, the method for detecting accents in music is not the focus of this application and is described in detail here.
[0081] In some embodiments of the present application, optionally, in step S130, synthesizing the first played music in the music library with the ambient sound effect includes:
[0082] The first played music and the ambient sound effect are volume-equalized and / or the tracks are merged.
[0083] It should be noted that merging the first played music with the ambient sound effect is not limited to only including volume equalization and / or track merging.
[0084] The purpose of volume equalization is to adjust the volume levels of different audio content so that the volume levels of the first played music and the ambient sound effects are consistent; the purpose of track merging is to mix the first played music and the ambient sound effects into a single audio file.
[0085] The synthesis of the first music played in the music library with the ambient sound effect is not limited to volume balancing and / or track merging. Examples of each merging method are not provided here. Furthermore, the specific methods of volume balancing and track merging are not the research focus of this application and are not described in detail here.
[0086] Using this implementation, the first playback music in the music library is synthesized with the ambient sound effect to generate the second playback music. When playing the second playback music, music with ambient sound effects that meet the user's current state and / or environmental requirements is played to meet the user's current state and / or environmental requirements.
[0087] like Figure 2 When using the music playing method described in the embodiment of the present application, the process of performing model training to obtain the first model includes:
[0088] Obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0089] encoding the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encoding the state description information to obtain state features of the state description information;
[0090] The state feature is used as input and the data feature is used as output to perform model training to obtain the first model.
[0091] After obtaining multiple training ambient sound effects, the state description information corresponding to the training ambient sound effects can be annotated manually, wherein the annotator describes the listening state of each training ambient sound effect after listening to the training ambient sound effect.
[0092] Among them, the multiple training atmosphere sound effects include sound effects applicable to various application scenarios.
[0093] In an embodiment of the present application, optionally, the training atmosphere sound effects can be input into a sound effect encoding model for sound effect encoding to obtain data features corresponding to the training atmosphere sound effects. Optionally, the sound effect encoding model adopts a variational autoencoder (VAE) architecture to represent the complex data of the training atmosphere sound effects in an effective low-dimensional manner to obtain corresponding data features, and the obtained data features are used to characterize the feature vectors of the training atmosphere sound effects. Optionally, the sound effect encoding model includes an encoder Transformer module for realizing low-dimensional feature vector representation of the training atmosphere sound effects.
[0094] Optionally, the state description information corresponding to each training ambient sound effect can be input into a state encoding model for encoding to obtain state features corresponding to the state description information. Optionally, the state encoding model uses a pre-trained large language model (such as a Llama model) to map the state description information into state encoding to obtain corresponding state features.
[0095] In an embodiment of the present application, optionally, the state feature is used as input and the data feature is used as output to perform model training, including:
[0096] The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
[0097] Among them, the loss function Loss is expressed as:
[0098]
[0099] Among them, FeatureB is the data feature corresponding to the training atmosphere sound effect, FeatureA is the state feature of the state description information, and N is the model output dimension.
[0100] Using this implementation, the state characteristics of the state description information are taken as input, and the data characteristics corresponding to the training atmosphere sound effects are used as output. During the training process, the above-mentioned loss function is used to make the input and output close. When the obtained loss function is less than or equal to the preset value, the required first model is obtained.
[0101] After obtaining the first model above, combined with Figure 1 As shown, the ambient sound effect data output by the first model is input into the sound effect decoding module to obtain an ambient sound effect that can be synthesized with the first played music. Optionally, the sound effect decoding module adopts a VAE architecture to convert the obtained ambient sound effect data into a corresponding ambient sound effect.
[0102] like Figure 3 The present invention relates to an implementation process of generating an ambient sound effect using the first model when using the music playback method described in an embodiment of the present application. In this process, after obtaining the user's current state description information, the current state description information is input into the state encoding model for encoding to obtain the current state features corresponding to the state description information.
[0103] The current state description information is used as the input of the first model, including:
[0104] The current state feature is used as input of the first model.
[0105] The first model obtained through the above training process receives the current state features, uses the above loss function as a constraint condition, obtains an ambient sound effect that satisfies the constraint condition, and outputs the ambient sound effect. The output ambient sound effect is further decoded by the audio decoding module to obtain the playable ambient sound effect. The ambient sound effect output by the audio decoding module is then synthesized with the first playable music in the music library to obtain the second playable music.
[0106] In some embodiments of the present application, optionally, in step S120, generating an ambient sound effect that conforms to the current state description information includes:
[0107] Generate multiple ambient sound effects that match the current state description information;
[0108] The step of synthesizing the first played music in the music library with the ambient sound effect includes:
[0109] Each generated ambient sound effect is synthesized with the first playback unit to generate a plurality of second playback music.
[0110] In the embodiment of the present application, the method further includes:
[0111] The plurality of second play music are played in a loop, or one of the plurality of second play music is played in a single loop.
[0112] In the embodiment of the present application, optionally, Figure 4 As shown, the music playing method of this embodiment can generate a first application interface when applied. The first application interface 10 includes a voice input button 1 for describing the current status, a content display interface 2 for converting the current status description in voice format into text format, a first touch button 3 for obtaining an ambient sound effect that matches the current status description and generating the second music to be played, and a second touch button (or confirmation button) 4 for confirming the playback of the second music. Optionally, the voice input button 1 is displayed as "Tell me about your status today."
[0113] Based on the above Figure 4 In the first application interface shown, the user long-presses the voice input button 1. While performing voice input, the content display interface 2 displays a text description of the current state, such as "I'm a little tired at work today." Then, by clicking the first touch button 3, a second playback music track is generated that matches the current state description. After the generation is complete, the second touch button 4 is clicked to play the second playback music track.
[0114] Figure 5 This is a schematic diagram of a second application interface using the method described in one of the embodiments of the present application, wherein the second application interface 10 is a playback interface for the second music playback, and the playback interface includes multiple playback control buttons for the second music playback, and the multiple playback control buttons include a music playback button, a "dissatisfied" touch button corresponding to one of the second music playback buttons, and multiple loop playback buttons for the second music playback.
[0115] Using this second application interface, you can click the play music button for each of the multiple second play music to play it. If you are not satisfied with the effect of the generated second play music, you can click the "unsatisfied" touch button to delete the corresponding second play music from the list of second play music. In addition, you can click the loop play button to loop play multiple second play music in the list of second play music.
[0116] It should be noted that the first application interface and the second application interface of the music playback method described in the embodiment of the present application are only examples and are not limited to the application interfaces listed above.
[0117] By adopting the music playback method described in the embodiment of the present application, after obtaining the user's current status description information, music that meets the current status and / or environmental requirements of the current status description information can be generated to enhance the overall atmosphere and meet personalized music experience needs. In addition, by adopting this method, there is no need for complicated means to record sound effects.
[0118] One embodiment of the present application also provides a model training method, such as Figure 6 As shown, the method includes:
[0119] S610, obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0120] S620: Encode the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encode the state description information to obtain state features of the state description information;
[0121] S630, using the state feature as input and the data feature as output, to perform model training to obtain a first model; wherein the first model is used to generate an ambient sound effect that conforms to the current state description information of the user based on the current state description information.
[0122] By adopting the model training method described in this embodiment, model training is performed using multiple training atmosphere sound effects and the state description information corresponding to each training atmosphere sound effect, and a first model can be obtained that can generate an atmosphere sound effect that conforms to the current state description information based on the current state description information.
[0123] Optionally, the model training method, wherein the state feature is used as input and the data feature is used as output, performs model training, comprising:
[0124] The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
[0125] For a detailed description of the specific implementation of the model training method described in the embodiments of this application, please refer to the above-mentioned specific implementation of the music playback method in this application, and the description will not be repeated here.
[0126] One embodiment of the present application further provides a music playing device, such as Figure 7 Shown, including:
[0127] Information acquisition module 710, used to obtain the user's current status description information;
[0128] The processing module 720 is configured to use the current state description information as input to a first model to obtain an ambient sound effect output by the first model; the first model is configured to generate an ambient sound effect that conforms to the current state description information based on the current state description information;
[0129] The synthesis module 730 is used to synthesize the first play music in the music library and the ambient sound effect to generate the second play music.
[0130] Optionally, in the music playing device, the synthesis module 730 is further configured to:
[0131] Get the beats per minute (BPM) of the ambient sound effect;
[0132] The first played music is selected from the music library according to the beats per minute of the ambient sound effect; wherein the beats per minute of the first played music matches the beats per minute of the ambient sound effect.
[0133] Optionally, in the music playing device, the current state description information includes one or more of emotion description information, environment description information and sound effect description information.
[0134] Optionally, in the music playing device, the information acquisition module 710 is further configured to:
[0135] In a case where the acquired current state description information is in voice format, converting the current state description information in voice format into text format;
[0136] The processing module 720 uses the current state description information as input to the first model, including:
[0137] The current state description information in text format is used as input of the first model.
[0138] Optionally, in the music playing device, the processing module 720 is further configured to:
[0139] Obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0140] encoding the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encoding the state description information to obtain state features of the state description information;
[0141] The state feature is used as input and the data feature is used as output to perform model training to obtain the first model.
[0142] Optionally, in the music playing device, the processing module 720 is further configured to:
[0143] Encoding the current state description information to obtain a current state feature corresponding to the current state description information;
[0144] The processing module 720 uses the current state description information as input to the first model, including:
[0145] The current state feature is used as input of the first model.
[0146] Optionally, in the music playing device, the processing module 720 uses the state feature as input and the data feature as output to perform model training, including:
[0147] The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
[0148] The music playing method and the music playing device described in the embodiment of the present application are based on the same application concept. Since the principles of solving problems by the method and the device are similar, the implementation of the device and the method can refer to each other, and the repeated parts will not be repeated.
[0149] One embodiment of the present application also provides a model training device, such as Figure 8 Shown, including:
[0150] An information acquisition module 810 is configured to obtain a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects;
[0151] The encoding module 820 is configured to encode the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and to encode the state description information to obtain state features of the state description information;
[0152] The model training module 830 is used to perform model training using the state feature as input and the data feature as output to obtain a first model; wherein the first model is used to generate an ambient sound effect that conforms to the current state description information based on the user's current state description information.
[0153] Optionally, in the model training device, the model training module 830 uses the state feature as input and the data feature as output to perform model training, including:
[0154] The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
[0155] The model training method and the model training device described in the embodiment of the present application are based on the same application concept. Since the principles of solving problems by the method and the device are similar, the implementation of the device and the method can refer to each other, and the repeated parts will not be repeated.
[0156] One embodiment of the present application also provides a network device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, it implements the music playback method as described in any one of the above items, or implements the model training method as described in any one of the above items.
[0157] Among them, the specific implementation method of executing the music playback method or model training method by the program running on the processor of the network device can refer to the detailed description of the music playback method or model training method when it is applied to the network device, and will not be repeated here.
[0158] In addition, a specific embodiment of the present application also provides a readable storage medium on which a computer program is stored, wherein when the program is executed by a processor, the steps in the music playback method or model training method as described in any one of the above items are implemented.
[0159] Specifically, the readable storage medium is applied to the above-mentioned network device. When applied to the network device, the execution steps in the corresponding music playback method or model training method are described in detail above and will not be repeated here.
[0160] Another embodiment of the present application further provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, implement the steps in any of the instant interaction methods described above, or implement the steps in any of the instant interaction methods described above.
[0161] Optionally, the embodiments of the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0162] The computer program product described in the embodiment of the present application includes computer instructions that, when executed by a processor, implement the various processes of the method embodiment shown above and can achieve the same technical effect. To avoid repetition, they will not be described here.
[0163] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection of some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0164] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may be physically included separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0165] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute some steps of the sending and receiving methods described in various embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0166] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary personnel in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A music playing method, characterized in that: include: Get the user's current status description information; Using the current state description information as input to a first model, and obtaining an ambient sound effect output by the first model; The first model is used to generate an ambient sound effect that conforms to the current state description information according to the current state description information; The first play music in the music library is synthesized with the ambient sound effect to generate the second play music.
2. The music playing method according to claim 1, wherein: The method further comprises: Get the beats per minute (BPM) of the ambient sound effect; The first played music is selected from the music library according to the beats per minute of the ambient sound effect; wherein the beats per minute of the first played music matches the beats per minute of the ambient sound effect.
3. The music playing method according to claim 1, wherein: The method further comprises: Obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects; encoding the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encoding the state description information to obtain state features of the state description information; The state feature is used as input and the data feature is used as output to perform model training to obtain the first model.
4. The music playing method according to claim 3, wherein: The state feature is used as input and the data feature is used as output to perform model training, including: The state feature is used as input and the data feature is used as output, and a calculated value of a loss function determined according to the state feature and the data feature is less than or equal to a preset value as a constraint condition for model training.
5. A model training method, characterized in that: include: Obtaining a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects; encoding the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and encoding the state description information to obtain state features of the state description information; The state feature is used as input and the data feature is used as output to perform model training to obtain a first model; wherein the first model is used to generate an ambient sound effect that conforms to the current state description information of the user based on the current state description information.
6. A music playing device, characterized in that: include: Information acquisition module, used to obtain the user's current status description information; a processing module, configured to use the current state description information as input to a first model to obtain an ambient sound effect output by the first model; The first model is used to generate an ambient sound effect that conforms to the current state description information according to the current state description information; The synthesis module is used to synthesize the first play music in the music library and the ambient sound effect to generate the second play music.
7. A model training device, characterized in that: include: An information acquisition module, configured to obtain a plurality of training ambiance sound effects and state description information corresponding to each of the training ambiance sound effects; an encoding module, configured to encode the training ambiance sound effect to obtain data features corresponding to the training ambiance sound effect; and to encode the state description information to obtain state features of the state description information; The model training module is used to perform model training using the state feature as input and the data feature as output to obtain a first model; wherein the first model is used to generate an ambient sound effect that conforms to the current state description information of the user based on the current state description information.
8. A network device, characterized in that: It includes a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the music playing method as described in any one of claims 1 to 4, or implements the model training method as described in claim 5.
9. A readable storage medium, characterized in that: The readable storage medium stores a program, which, when executed by the processor, implements the steps of the music playing method as described in any one of claims 1 to 4, or implements the steps of the model training method as described in claim 5.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps in the music playing method as described in any one of claims 1 to 4, or implement the steps in the model training method as described in claim 5.