Music generation method, apparatus, and electronic device
By acquiring initial music data and user data, extracting music features and user features using a target model, and generating personalized target music data, the problem of insufficient music universality in existing technologies is solved, and personalized music generation and real-time adaptation are realized.
Patent Information
- Application Number
- CN202211229818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-09
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-10-09
AI Technical Summary
In existing technologies, pre-uploaded music is relatively generic and cannot meet personalized needs, nor can it generate personalized music based on the user's real-time status.
By acquiring initial music data and user data, and using a target model to extract music features and user features, personalized target music data is generated. Combined with the user's physiological, facial expression, and EEG data, personalized music generation is achieved.
It enables the generation of personalized music based on the user's real-time status, improving the adaptability and comfort of the music and meeting the user's personalized needs.
Smart Images

Figure CN115662372B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of artificial intelligence, and particularly relates to a music generation method and device and electronic equipment. BACKGROUND
[0002] Many modern people are disturbed by stress, anxiety, confusion and physical sub-health, and music therapy has become a popular healing method. For example, the main principle of the singing bowl sound therapy is to use sound wave resonance therapy. The bowl body of the singing bowl sound therapy, due to the special material and knocking method, the audio vibration emitted thereby can reach the nerves of the auditory, tactile and sensory parts of the human body, achieving the effect of massaging the nerves.
[0003] In the related art, the user downloads the pre-uploaded music from a video website and plays the downloaded music to relax the mood and relieve anxiety.
[0004] However, in the above related art, the pre-uploaded music is public music and is relatively general. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide a music generation method, device and electronic equipment, which can solve the problem of general music.
[0006] In a first aspect, the embodiments of the present application provide a music generation method, which comprises:
[0007] obtaining initial music data and first user data; the first user data is obtained in the case of playing the initial music data;
[0008] determining a music feature based on the initial music data;
[0009] determining a user feature based on the first user data;
[0010] generating first target music data based on the music feature and the user feature.
[0011] In a second aspect, the embodiments of the present application provide a music generation device, which comprises:
[0012] a first obtaining module configured to obtain initial music data and first user data; the first user data is obtained in the case of playing the initial music data;
[0013] a first determining module configured to determine a music feature based on the initial music data;
[0014] a second determining module configured to determine a user feature based on the first user data;
[0015] The generating module is configured to generate first target music data based on the music features and the user features.
[0016] In a third aspect, an electronic device is provided. The electronic device includes a processor and a memory. The memory stores programs or instructions executable on the processor. The programs or instructions, when executed by the processor, implement the steps of the method according to the first aspect.
[0017] In a fourth aspect, a readable storage medium is provided. The readable storage medium stores programs or instructions. The programs or instructions, when executed by a processor, implement the steps of the method according to the first aspect.
[0018] In a fifth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to execute programs or instructions to implement the method according to the first aspect.
[0019] In a sixth aspect, a computer program product is provided. The computer program product is stored in a storage medium. The computer program product is executed by at least one processor to implement the method according to the first aspect.
[0020] In the embodiments of the present application, music features in initial music data are extracted, and user features in user data obtained in the case of playing the initial music data are extracted. First target music data is generated based on the music features and the user features. In the generation process of the first target music data, the corresponding user data when the user listens to the initial music data is considered, so that the generation of personalized music is realized. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 is one of flowcharts of the music generation method provided by the embodiments of the present application;
[0022] Figure 2 is another flowchart of the music generation method provided by the embodiments of the present application;
[0023] Figure 3 is a third flowchart of the music generation method provided by the embodiments of the present application;
[0024] Figure 4 is a schematic diagram of a network structure provided by the embodiments of the present application;
[0025] Figure 5 is a fourth flowchart of the music generation method provided by the embodiments of the present application;
[0026] Figure 6 is a fifth flowchart of the music generation method provided by the embodiments of the present application;
[0027] Figure 7 is a structural schematic diagram of a music generation apparatus provided by an embodiment of the present application.
[0028] Figure 8 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.
[0029] Figure 9 is a hardware structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0031] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in an "or" relationship.
[0032] The music generation method, music generation apparatus, electronic device and readable storage medium provided by the embodiments of the present application will be described in detail below with reference to the drawings, specific embodiments and application scenarios.
[0033] Figure 1 is one of the flow schematic diagrams of the music generation method provided by an embodiment of the present application, as shown in Figure 1 The music generation method comprises the following steps 101, 102, 103 and 104; wherein:
[0034] Step 101, obtaining initial music data and first user data; the first user data is obtained in the case of playing the initial music data.
[0035] Wherein, when playing the initial music data, the user listens to the played initial music data.
[0036] For example, when a user needs to relax by music, initial music data to be listened to can be downloaded from a video website, the initial music data is played, and the change information of the user's body when listening to the initial music data is obtained, that is, the first user data is obtained.
[0037] Step 102, determining a music feature based on the initial music data.
[0038] For example, when the initial music data is obtained, the initial music data is analyzed to obtain a music feature. The music feature can be a size feature of a musical instrument for generating the initial music, an operation means feature of the musical instrument for generating the initial music, a matched environment sound feature, and a suitable scene feature, etc. The operation means of the musical instrument can be knocking or rubbing, etc. The environment sound can be a sea wave sound, a bird call sound, etc. The scene can be summer, a fast heartbeat, etc.
[0039] Step 103, determining a user feature based on the first user data.
[0040] For example, when the first user data is obtained, the first user data is analyzed to obtain a user feature. The user feature can be a body temperature feature, a heart rate feature, an electroencephalogram feature, and an expression feature of the user when listening to the initial music data, etc.
[0041] Step 104, generating first target music data based on the music feature and the user feature.
[0042] For example, when the music feature and the user feature are obtained, the music feature and the user feature are synthesized to obtain the first target music data.
[0043] The music generation method provided by the embodiment of the present application extracts a music feature in the initial music data and extracts a user feature in the user data obtained when the initial music data is played. The first target music data is generated based on the music feature and the user feature. The corresponding user data when the user listens to the initial music data is considered in the generation process of the first target music data, so that the generation of personalized music is realized.
[0044] In an embodiment, Figure 2 is a flowchart of the music generation method provided by the embodiment of the present application, as shown in Figure 2 Step 102 can be implemented by the following steps.
[0045] Step 1021, inputting the initial music data into an encoder of a target model to obtain a music vector output by the encoder.
[0046] Step 1022, inputting the music vector into a classifier of the target model to obtain a music feature output by the classifier.
[0047] The music feature includes at least one of an instrument size feature, an instrument operation means feature, an environmental sound feature, and a scene feature; the target model is trained based on music data samples and labels corresponding to the music data samples; and the label includes at least one of an instrument size label, an instrument operation means label, an environmental sound label, and a scene label.
[0048] For example, the target model includes an encoder, a classifier, and a generator, the encoder can be a long short-term memory (LSTM) encoder, and the training process of the target model is as follows: first, each music data sample is labeled to indicate the corresponding instrument size, the corresponding instrument operation means, whether the environmental sound is matched, and the applicable scene of each music data sample, and an initial encoder, an initial classifier, and an initial generator are constructed; then, the music data sample is split into multiple music data segment samples, the music data segment samples of different lengths can be padded to the same length, each music data segment sample is encoded by using an embedding layer to obtain a corresponding encoded vector sample, for example, the encoded vector sample can be represented as batch x length x 1024, batch represents the number of data input to the LSTM each time, length represents the length of the music data segment sample, and 1024 is a hyperparameter; the encoded vector sample is input to the LSTM to obtain a music vector sample output by the LSTM, for example, the music vector sample can be represented as batch x 1024; the music vector sample is input to the initial classifier to obtain a music feature sample output by the initial classifier, the music feature sample includes an instrument size sample, an instrument operation means sample, an environmental sound sample, and a scene sample, etc., a first cross-entropy loss is determined based on the label and the music feature sample; the music vector is input to the initial generator to obtain a music sequence sample output by the initial generator, a second cross-entropy loss is determined based on the music sequence sample and a next music data segment sample of the current input music data segment sample, the sum of the first cross-entropy loss and the second cross-entropy loss is used as a total loss function, the parameters of the initial encoder, the parameters of the initial classifier, and the parameters of the initial generator are optimized based on the total loss function until a convergence condition is reached, a trained encoder, a trained classifier, and a trained generator are obtained, the target model is composed of the trained encoder, the trained classifier, and the trained generator, and self-encoding generation training and multi-task training are realized.
[0049] For example, assuming that the target model is trained based on music data samples and corresponding instrument size labels, instrument operation means labels, environmental sound labels and scene labels of the music data samples, the initial music data is input into the encoder of the target model, the initial music data is encoded by the encoder to obtain a music vector, and then the music vector is input into the classifier of the target model, and the music features are predicted based on the music vector by the classifier to obtain the instrument size features, the instrument operation means features, the environmental sound features and the scene features.
[0050] It should be noted that the generator in the target model can include an LSTM decoder and a factor classifier, the music sequence is synthesized through the LSTM decoder, and the synthesized music sequence is classified through the factor classifier to obtain the instrument size, the instrument operation means, the environmental sound and the scene, so that the generator has the ability of music synthesis and the ability of music classification.
[0051] The music generation method provided by the embodiment of the application encodes the initial music data based on the encoder of the target model to obtain a music vector, and then extracts music features based on the classifier of the target model to obtain music features, thereby realizing automatic extraction of music features.
[0052] In an embodiment, the first user data includes at least one of physiological data, expression data and brain wave data; the physiological data includes at least one of body temperature data, heart rate data and user preference data; the user preference data includes at least one of user preference instrument size, user preference instrument operation means, user preference environmental sound and user preference scene; and the above step 103 can be implemented by the following method:
[0053] Determine a target feature based on the first user data; the target feature includes at least one of physiological features, expression features and brain wave features; the physiological feature corresponding to the body temperature data includes a body temperature feature, the physiological feature corresponding to the heart rate data includes a heart rate feature, and the physiological feature corresponding to the user preference data includes a user preference feature.
[0054] In the case where the target feature includes at least two of the physiological features, the expression features and the brain wave features, at least two of the physiological features, the expression features and the brain wave features are input into a first feature fusion model to obtain the user feature output by the first feature fusion model.
[0055] The target feature is obtained by performing feature extraction on the first user data by a first feature extraction model; and the first feature extraction model is trained based on first user data samples.
[0056] Specifically, in the case that the target feature is a physiological feature, the first feature extraction model comprises a first feature extraction sub-model, and the physiological feature is obtained by performing feature extraction on the physiological data by the first feature extraction sub-model; the first feature extraction sub-model is trained based on physiological data samples; in the case that the target feature is an expression feature, the first feature extraction model comprises a second feature extraction sub-model, and the expression feature is obtained by performing feature extraction on the expression data by the second feature extraction sub-model; the second feature extraction sub-model is trained based on expression data samples; in the case that the target feature is an electroencephalogram feature, the first feature extraction model comprises a third feature extraction sub-model, and the electroencephalogram feature is obtained by performing feature extraction on the electroencephalogram data by the third feature extraction sub-model; the third feature extraction sub-model is trained based on electroencephalogram data samples.
[0057] The first feature extraction sub-model can be a three-layer fully connected network, for example, the weight size of the first fully connected network is 3x128, the weight size of the second fully connected network is 128x512, and the weight size of the third fully connected network is 512x256, and the physiological feature output by the first feature extraction sub-model has a dimension of 256; the second feature extraction sub-model can be a visual geometry group (VGG) network, and the output expression feature can have a dimension of 1024; the third feature extraction sub-model can be a combination of a one-layer fully connected layer and a transformer network, the electroencephalogram data is mapped into a 300-dimensional vector representation through the one-layer fully connected layer, and then a 6-layer transformer network is used to learn the feature representation of the electroencephalogram, and the output electroencephalogram feature can have a dimension of 512.
[0058] For example, the heart rate data and the body temperature data of the user can be collected by a wearable device such as a smart watch or a bracelet worn by the user, and the previously stored user preference data is obtained, and the user preference data is empty at the initial use; the user likes to listen to music with bird calls, so the user prefers the bird call sound; the user likes to listen to music when the heartbeat is accelerated, so the user prefers the heartbeat acceleration scene; the user likes to listen to music produced by knocking the pot body, so the user prefers the percussion operation means; the user likes the music produced by the small size pot body, so the user prefers the small size of the musical instrument, for example, the user prefers the pot body with a caliber of 5 cm, and the heart rate data, the body temperature data and the user preference data of the user when listening to the initial music data are used as the physiological data, which can better reflect the user's reaction when listening to the initial music data, and further generate target music data based on the physiological data, so that the generated target music data is more in line with the user and is more personalized.
[0059] For example, taking physiological data, expression data and brain wave data as examples of the first user data, physiological data is collected by a smart watch or a wearable device such as a bracelet worn by the user, expression pictures of the user are collected by a camera of a terminal device such as a mobile phone as expression data, for example, an expression picture is collected every 10 seconds, 6 expression pictures are collected in one minute, and the expression features corresponding to each expression picture are averaged to obtain the final expression features; the brain wave data of the user is collected by a head-mounted device, for example, the brain wave data is collected every 1 second.
[0060] The physiological data is input into the first feature extraction sub-model, the physiological data is analyzed by the first feature extraction sub-model to obtain physiological features; the expression data is input into the second feature extraction sub-model, the expression data is analyzed by the second feature extraction sub-model to obtain expression features; the brain wave data is input into the third feature extraction sub-model, the brain wave data is analyzed by the third feature extraction sub-model to obtain brain wave features, and finally the physiological features, expression features and brain wave features are fused to obtain user features.
[0061] For example, taking physiological features, expression features and brain wave features as examples of target features, the physiological features, expression features and brain wave features are input into the first feature fusion model, wherein the first feature fusion model can be a two-layer fully connected layer, for example, according to the dimensions of the physiological features, expression features and brain wave features obtained above, the physiological features have a dimension of 256, the expression features have a dimension of 1024, and the brain wave features have a dimension of 512, the weight size of the first layer fully connected layer can be 1792x2048, the weight size of the second layer fully connected layer can be 2048x1024, and finally the dimension of the user features output by the first feature fusion model is 1024.
[0062] It should be noted that in the case where the target features include any one of the physiological features, expression features and brain wave features, the any one feature is taken as the user features.
[0063] The music generation method provided by the embodiment of the present application realizes automatic extraction of user features based on the first feature extraction sub-model for extracting physiological features from physiological data, the second feature extraction sub-model for extracting expression features from expression data, and the third feature extraction sub-model for extracting brain wave features from brain wave data; in addition, at least two of the physiological features, expression features and brain wave features are spliced to obtain user features, so that the user features are more abundant and can better reflect the user's reaction when listening to the initial music data, and further, when generating target music data based on the user features, the generated target music data is more in line with the user and more personalized.
[0064] In an embodiment, Figure 3 is a third flowchart of the music generation method provided by the embodiments of the present application, as shown in Figure 3 The step 104 can be implemented by the following steps, as shown in the following.
[0065] In step 1041, the music features and the user features are input into a second feature fusion model. The second feature fusion model is used to determine the cosine similarity of each feature included in the music features corresponding to the user features, and the cosine similarity and the corresponding features are weighted and summed to obtain music fusion features output by the second feature fusion model.
[0066] The second feature fusion model can be an attention model.
[0067] For example, when the music features and the user features are obtained, the music features and the user features are input into the second feature fusion model. The second feature fusion model is used to calculate the cosine similarity of each feature included in the music features corresponding to the user features, and the cosine similarity and the corresponding features are weighted and summed to obtain the music fusion features.
[0068] In step 1042, the music fusion features are input into a generator of the target model. The generator is used to obtain the first target music data based on the music fusion features.
[0069] For example, the music fusion features are input into the generator of the trained target model. The generator is used to generate the first target music data based on the music fusion features.
[0070] In addition, the first feature extraction sub-model, the second feature extraction sub-model, the third feature extraction sub-model, the first feature fusion model, the second feature fusion model, and the generator can be trained together. The specific training process is as follows: collecting music data samples, physiological data samples, expression data samples, and brain wave data samples and music data samples, constructing a first initial feature extraction model, a second initial feature extraction model, a third initial feature extraction model, a first initial feature fusion model, and a second initial feature fusion model, marking user features in the music data samples, such as marking high temperature, heart rate, facial tension, and brain electrical activity in the music data samples; inputting the physiological data samples into the first initial feature extraction model to obtain physiological feature samples output by the first initial feature extraction model; inputting the expression data samples into the second initial feature extraction model to obtain expression feature samples output by the second initial feature extraction model; inputting the brain wave data samples into the third initial feature extraction model to obtain brain wave feature samples output by the third initial feature extraction model; then inputting the physiological feature samples, the expression feature samples, and the brain wave feature samples into the first initial feature fusion model to obtain user feature samples output by the first initial feature fusion model, inputting the user feature samples and the music feature samples into the second initial feature fusion model to obtain music fusion feature samples output by the second initial feature fusion model, and finally inputting the music fusion feature samples into the generator to obtain a predicted music vector output by the generator, constructing a cross-entropy loss function based on the predicted music vector and a music vector corresponding to the marked music data sample, optimizing the parameters of the first initial feature extraction model, the parameters of the second initial feature extraction model, the parameters of the third initial feature extraction model, the parameters of the first initial feature fusion model, the parameters of the second initial feature fusion model, and the parameters of the generator based on the cross-entropy loss function, until a convergence condition is reached, to obtain the first feature extraction sub-model, the second feature extraction sub-model, the third feature extraction sub-model, the first feature fusion model, the second feature fusion model, and the optimized generator.
[0071] Figure 4 is a network structure schematic diagram provided by the embodiment of the present application, as Figure 4As shown, the music generation method includes a target model 401, a first feature extraction sub-model 402, a second feature extraction sub-model 403, a third feature extraction sub-model 404, a first feature fusion model 405, and a second feature fusion model 406. The target model 401 includes an encoder 4011, a classifier 4012, and a generator 4013. The initial music data is input into the encoder 4011 of the target model 401, and a music vector output by the encoder 4011 is obtained. The music vector is input into the classifier 4012 of the target model 401, and a music feature output by the classifier 4012 is obtained. The physiological data is input into the first feature extraction sub-model 402, and a physiological feature output by the first feature extraction sub-model 402 is obtained. The expression data is input into the second feature extraction sub-model 403, and an expression feature output by the second feature extraction sub-model 403 is obtained. The brain wave data is input into the third feature extraction sub-model 404, and a brain wave feature output by the third feature extraction sub-model 404 is obtained. The physiological feature, the expression feature, and the brain wave feature are input into the first feature fusion model 405, and a user feature output by the first feature fusion model 405 is obtained. The user feature output by the first feature fusion model 405 and the music feature output by the classifier 4012 are input into the second feature fusion model 406, and a music fusion feature output by the second feature fusion model 406 is obtained. Finally, the music fusion feature is input into the generator 4013 of the target model 401, and first target music data is finally output.
[0072] The music generation method provided by the embodiment of the present application fuses the music feature and the user feature through the second feature fusion model to obtain the music fusion feature, and then inputs the music fusion feature into the generator of the target model to generate the first target music data conforming to the user himself, thereby realizing online generation of personalized first target music data.
[0073] In an embodiment, Figure 5 is a fourth flowchart of the music generation method provided by the embodiment of the present application, as Figure 5 shown, after the step 104, the music generation method further includes the following steps:
[0074] Step 105, playing the first target music data.
[0075] For example, when generating the first target music data conforming to the user, the first target music data is played through a loudspeaker, so that the user can listen to the first target music data conforming to himself online.
[0076] Step 106, obtaining second user data.
[0077] For example, while the user is listening to the first target music data online, the body data of the user when listening to the first target music data is continuously collected, that is, the body temperature data and the heart rate data are collected at regular time intervals through the wearable device such as a smart watch or a bracelet worn by the user, the expression data of the user is collected through a camera, and the brain wave data of the user is collected through a head-mounted device, and the newly collected body temperature data, heart rate data, expression data and brain wave data are taken as the second user data.
[0078] In step 107, if it is determined that the difference between the first user data and the second user data is less than or equal to a first preset threshold, the first user data is updated to the second user data, and the step of determining the user feature based on the first user data is returned to, to obtain second target music data.
[0079] For example, when the second user data is obtained, the second user data is compared with the first user data obtained before, for example, the second user data includes new brain wave data and new heart rate data, first, the first brain wave value corresponding to the new brain wave data and the second brain wave value corresponding to the previous brain wave data are determined, and the first permutation entropy value of the variation sequence corresponding to the new heart rate data and the second permutation entropy value of the variation sequence corresponding to the previous heart rate data are determined, the second brain wave value is compared with the first brain wave value, and the second permutation entropy value is compared with the first permutation entropy value, if it is determined that the difference between the second brain wave value and the first brain wave value is less than or equal to a first preset threshold, for example, the first preset threshold is 2 Hz, and it is determined that the difference between the second permutation entropy value and the first permutation entropy value is less than or equal to a first preset threshold, for example, the first preset threshold is 3, it is indicated that the first target music data currently played is not very suitable for the user, and the second target music data needs to be regenerated based on the user feature and the music feature corresponding to the second user data.
[0080] It should be noted that when it is determined that the difference between the second brain wave value and the first brain wave value is greater than the first preset threshold, and / or, it is determined that the difference between the second permutation entropy value and the first permutation entropy value is greater than the first preset threshold, it is indicated that the first target music data currently played is suitable for the user, the first target music data is continued to be played, and the second user data of the user when listening to the first target music data is continuously obtained.
[0081] It should be noted that the second user data can also include heart rate data, body temperature data and expression data, for example, the expression data obtained before is frowning, and the new expression data can be frowning, which indicates that the first target music data currently played is not suitable for the user, and the second target music data needs to be regenerated, which is not limited in the present application.
[0082] It should be noted that when the second user data is obtained, the second user data can also be used as a user data sample to continue relearning the model parameters of the target model, the first feature extraction sub-model, the second feature extraction sub-model, the third feature extraction sub-model, the first feature fusion model, and the second feature fusion model, so that the final obtained models are more accurate.
[0083] Step 108: Replacing the first target music data with the second target music data for playing.
[0084] For example, when the second target music data is regenerated, the second target music data is played as the music currently needed to be played, and the second user data of the user when listening to the second target music data is continuously obtained, and the played music is continuously updated to realize personalized generation of music.
[0085] The music generation method provided by the embodiments of the present application can continuously monitor the user data for feedback while playing the first target music data, and adjust the music being played in real time, so that the generated music is more personalized.
[0086] In an embodiment, Figure 6 is a fifth flowchart of the music generation method provided by the embodiments of the present application, as Figure 6 shown, after the above step 105, the music generation method further includes the following steps:
[0087] Step 109: Obtaining a stop playing instruction; the stop playing instruction includes an instruction generated based on a received closing input, or an instruction generated in a case where the brain wave data in the first user data is less than or equal to a second preset threshold.
[0088] For example, when the first target music data is played, if the user does not want to listen to the first target music data, the user can click a pre-configured closing control on a playing interface to perform a closing input, so that the electronic device generates a stop playing instruction when receiving the closing input of the user; or the electronic device compares the brain wave value corresponding to the brain wave data in the first user data with a second preset threshold, for example, the second preset threshold is 7 Hz, and in a scientific sense, the brain wave value less than 7 Hz is considered to be in a sleep state; when it is determined that the brain wave value is less than or equal to the second preset threshold, a stop playing instruction is generated.
[0089] Step 110: In response to the stop playing instruction, the first target music data is closed.
[0090] For example, when the stop playing instruction is obtained, the first target music data is stopped playing in response to the stop playing instruction.
[0091] It should be noted that when the currently played music data is the second target music data after replacement, the second target music data is also closed.
[0092] The music generation method provided in the embodiments of the present application can trigger the closing of the first target music data based on user demand, or can close the first target music data when it is detected that the user is in a sleep state, thereby realizing real-time closing of the first target music data and avoiding bringing bad experience to the user.
[0093] The music generation method provided in the embodiments of the present application can trigger the closing of the first target music data based on user demand, or can close the first target music data when it is detected that the user is in a sleep state, thereby realizing real-time closing of the first target music data and avoiding bringing bad experience to the user.
[0094] Figure 7 is a structural schematic diagram of the music generation device provided in the embodiments of the present application, as shown in Figure 7 The music generation device 700 includes a first acquisition module 701, a first determination module 702, a second determination module 703, and a generation module 704; wherein:
[0095] The first acquisition module 701 is configured to acquire initial music data and first user data; the first user data is acquired in the case of playing the initial music data;
[0096] The first determination module 702 is configured to determine a music feature based on the initial music data;
[0097] The second determination module 703 is configured to determine a user feature based on the first user data;
[0098] The generation module 704 is configured to generate first target music data based on the music feature and the user feature.
[0099] The music generation device provided in the embodiments of the present application extracts the music feature in the initial music data and extracts the user feature in the user data acquired in the case of playing the initial music data, generates the first target music data based on the music feature and the user feature, and considers the corresponding user data when the user listens to the initial music data in the generation process of the first target music data, thereby realizing the generation of personalized music.
[0100] Optionally, the first determination module 702 is further configured to:
[0101] input the initial music data into an encoder of a target model to obtain a music vector output by the encoder;
[0102] inputting the music vector into a classifier of the target model to obtain a music feature output by the classifier; the music feature comprises at least one of the following: an instrument size feature, an instrument operation means feature, an environmental sound feature, and a scene feature;
[0103] The target model is trained based on music data samples and labels corresponding to the music data samples; the labels comprise at least one of the following: an instrument size label, an instrument operation means label, an environmental sound label, and a scene label.
[0104] Optionally, the first user data comprises at least one of the following: physiological data, expression data, and brain wave data; the physiological data comprises at least one of the following: body temperature data, heart rate data, and user preference data; the user preference data comprises at least one of the following: a user preferred instrument size, a user preferred instrument operation means, a user preferred environmental sound, and a user preferred scene; and the second determining module 703 is further configured to:
[0105] determine a target feature based on the first user data; the target feature comprises at least one of the following: a physiological feature, an expression feature, and a brain wave feature; the physiological feature corresponding to the body temperature data comprises a body temperature feature, the physiological feature corresponding to the heart rate data comprises a heart rate feature, and the physiological feature corresponding to the user preference data comprises a user preference feature;
[0106] in a case where the target feature comprises at least two of the physiological feature, the expression feature, and the brain wave feature, inputting the at least two of the physiological feature, the expression feature, and the brain wave feature into a first feature fusion model to obtain the user feature output by the first feature fusion model;
[0107] The target feature is obtained by performing feature extraction on the first user data by using a first feature extraction model; and the first feature extraction model is trained based on first user data samples.
[0108] Optionally, the generating module 704 is further configured to:
[0109] inputting the music feature and the user feature into a second feature fusion model, determining, by the second feature fusion model, a cosine similarity corresponding to each feature contained in the music feature based on the user feature, and performing weighted summation on the cosine similarity and the corresponding feature to obtain a music fusion feature output by the second feature fusion model;
[0110] inputting the music fusion feature into a generator of the target model to obtain the first target music data based on the music fusion feature by using the generator.
[0111] Optionally, the music generation apparatus 700 further comprises:
[0112] a playing module, configured to play the first target music data;
[0113] a second obtaining module, configured to obtain second user data;
[0114] a third determining module, configured to, in a case where a difference between the first user data and the second user data is less than or equal to a first preset threshold, update the first user data to the second user data, and return to the step of determining the user feature based on the first user data, to obtain second target music data;
[0115] the playing module is further configured to replace the first target music data with the second target music data for playing.
[0116] Optionally, the music generation apparatus 700 further comprises:
[0117] a third obtaining module, configured to obtain a stop playing instruction; the stop playing instruction comprises an instruction generated based on a received closing input, or an instruction generated in a case where electroencephalogram data in the first user data is less than or equal to a second preset threshold;
[0118] a closing module, configured to, in response to the stop playing instruction, close the first target music data.
[0119] The music generation apparatus in the embodiments of the present applicationapplicationbe an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic deviceapplicationbe a terminal, or other devices than a terminal. For example, the electronic deviceapplicationbe a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., andapplicationbe a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., and the embodiments of the present application do not make a specific limitation.
[0120] The music generation apparatus in the embodiments of the present application can be an apparatus with an operating system. The operating system can be an Android operating system, an ios operating system, or other possible operating systems, which are not limited in the embodiments of the present application.
[0121] The music generation apparatus provided in the embodiments of the present application can realize Figures 1 to 6 The method embodiments realize various processes, which are not repeated here to avoid repetition.
[0122] Optionally, as shown in Figure 8 The embodiments of the present application also provide an electronic device 800, which includes a processor 801 and a memory 802, and the memory 802 stores programs or instructions that can run on the processor 801. When the programs or instructions are executed by the processor 801, various steps of the above music generation method embodiments are realized, and the same technical effects are achieved. To avoid repetition, these are not repeated here.
[0123] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.
[0124] Figure 9 To realize the hardware structure of an electronic device in the embodiments of the present application.
[0125] The electronic device 900 includes but is not limited to a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909, and a processor 910, and the like.
[0126] Those skilled in the art can understand that the electronic device 900 can also include a power supply (such as a battery) that supplies power to each component. The power supply can be logically connected to the processor 910 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 9 The electronic device structure shown in the embodiments of the present application does not constitute a limitation on the electronic device. The electronic device can include more or fewer components than those shown, or combine certain components, or arrange different components, which are not repeated here.
[0127] The processor 910 is configured to obtain initial music data and first user data, and the first user data is obtained in the case of playing the initial music data.
[0128] The processor 910 is further configured to determine a music feature based on the initial music data.
[0129] The processor 910 is further configured to determine a user feature based on the first user data.
[0130] The processor 910 is further configured to generate first target music data based on the music feature and the user feature.
[0131] In the above embodiment, the music feature in the initial music data is extracted, and the user feature in the user data obtained in the case of playing the initial music data is extracted, and the first target music data is generated based on the music feature and the user feature. In the process of generating the first target music data, the corresponding user data when the user listens to the initial music data is considered, so as to realize the generation of personalized music.
[0132] Optionally, the processor 910 is further configured to:
[0133] input the initial music data into an encoder of a target model to obtain a music vector output by the encoder;
[0134] input the music vector into a classifier of the target model to obtain a music feature output by the classifier; the music feature includes at least one of the following: an instrument size feature, an instrument operation means feature, an environmental sound feature, and a scene feature;
[0135] The target model is trained based on a music data sample and a label corresponding to the music data sample; the label includes at least one of the following: an instrument size label, an instrument operation means label, an environmental sound label, and a scene label.
[0136] In the above embodiment, the initial music data is encoded by the encoder of the target model to obtain the music vector, and the music vector is feature-extracted by the classifier of the target model to obtain the music feature, so as to realize the automatic extraction of the music feature.
[0137] Optionally, the first user data includes at least one of the following: physiological data, expression data, and brain wave data; the physiological data includes at least one of the following: body temperature data, heart rate data, and user preference data; the user preference data includes at least one of the following: user preferred instrument size, user preferred instrument operation means, user preferred environmental sound, and user preferred scene.
[0138] The processor 910 is further configured to:
[0139] determine a target feature based on the first user data; the target feature includes at least one of the following: physiological feature, expression feature, and brain wave feature; the physiological feature corresponding to the body temperature data includes body temperature feature, the physiological feature corresponding to the heart rate data includes heart rate feature, and the physiological feature corresponding to the user preference data includes user preference feature.
[0140] In a case where the target feature comprises at least two of the physiological feature, the expression feature, and the brain wave feature, the at least two of the physiological feature, the expression feature, and the brain wave feature are input into a first feature fusion model to obtain the user feature output by the first feature fusion model.
[0141] The target feature is obtained by performing feature extraction on the first user data by using a first feature extraction model, and the first feature extraction model is trained based on first user data samples.
[0142] In the above implementation, the physiological feature is obtained by performing feature extraction on physiological data based on a first feature extraction model, the expression feature is obtained by performing feature extraction on expression data based on a second feature extraction model, and the brain wave feature is obtained by performing feature extraction on brain wave data based on a third feature extraction model. The physiological feature, the expression feature, and the brain wave feature are determined as the user feature, so that automatic extraction of the user feature is implemented. In addition, at least two of the physiological feature, the expression feature, and the brain wave feature are spliced to obtain the user feature, so that the user feature is relatively rich and can better reflect the reaction of the user when listening to the initial music data. When the target music data is generated based on the user feature, the generated target music data is more in line with the user and is more personalized.
[0143] Optionally, the processor 910 is further configured to:
[0144] The music feature and the user feature are input into a second feature fusion model, the second feature fusion model is used to determine a cosine similarity corresponding to each feature contained in the music feature based on the user feature, and the cosine similarity and the corresponding feature are weighted and summed to obtain a music fusion feature output by the second feature fusion model.
[0145] The music fusion feature is input into a generator of the target model, and the generator is used to obtain the first target music data based on the music fusion feature.
[0146] In the above implementation, the music feature and the user feature are fused by using the second feature fusion model to obtain a music fusion feature, and the music fusion feature is input into the generator of the target model to generate the first target music data in line with the user, so that online generation of personalized first target music data is implemented.
[0147] Optionally, the processor 910 is further configured to play the first target music data.
[0148] The processor 910 is further configured to obtain second user data.
[0149] The processor 910 is further configured to, in a case where the difference between the first user data and the second user data is less than or equal to a first preset threshold, update the first user data as the second user data, and return to the step of determining the user feature based on the first user data to obtain second target music data.
[0150] The processor 910 is further configured to replace the first target music data with the second target music data for playing.
[0151] In the above embodiment, the user data is continuously monitored for feedback while the first target music data is played, and the music being played is adjusted in real time, so that the generated music is more personalized.
[0152] Optionally, the processor 910 is further configured to obtain a stop playing instruction; the stop playing instruction includes an instruction generated based on a received closing input, or an instruction generated in a case where the brain wave data in the first user data is less than or equal to a second preset threshold.
[0153] The processor 910 is further configured to, in response to the stop playing instruction, close the first target music data.
[0154] In the above embodiment, the first target music data can be closed based on user demand, or the first target music data can be closed when it is detected that the user is in a sleep state, so that the first target music data is closed in real time, and the user is prevented from having a bad experience.
[0155] It should be understood that, in the embodiments of the present application, the input unit 904 can include a graphics processor (GPU) 9041 and a microphone 9042. The graphics processor 9041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 906 can include a display panel 9061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 can include a touch detection device and a touch controller. The other input devices 9072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.
[0156] The memory 909 can be used to store software programs and various data. The memory 909 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), etc. In addition, the memory 909 can include a volatile memory or a non-volatile memory, or the memory 909 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 909 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0157] The processor 910 can include one or more processing units; optionally, the processor 910 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 910.
[0158] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned music generation method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.
[0159] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0160] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the music generation method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.
[0161] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system or a system on chip, etc.
[0162] The embodiment of the present application provides a computer program product, which is stored in a storage medium, and the program product is executed by at least one processor to realize the processes of the music generation method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.
[0163] It should be noted that in this document, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in the opposite order, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.
[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0165] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. A method for generating music, characterized in that, include: Obtain initial music data and first user data; The first user data was obtained while the initial music data was being played; Musical characteristics are determined based on the initial music data; User characteristics are determined based on the first user data, and the user characteristics are the characteristics corresponding to the user when listening to the initial music data; The music features and the user features are input into the second feature fusion model. The second feature fusion model determines the cosine similarity of each feature contained in the music features based on the user features. The cosine similarity and the corresponding features are then weighted and summed to obtain the music fusion features output by the second feature fusion model. The music fusion features are input into the generator of the target model, and the generator obtains the first target music data based on the music fusion features. The target model is trained based on music data samples and the corresponding labels of the music data samples; the labels include at least one of the following: instrument size label, instrument operation method label, ambient sound label, and scene label.
2. The music generation method according to claim 1, characterized in that, The process of determining musical features based on the initial music data includes: The initial music data is input into the encoder of the target model to obtain the music vector output by the encoder; The music vector is input into the classifier of the target model to obtain the music features output by the classifier; the music features include at least one of the following: instrument size features, instrument operation features, ambient sound features, and scene features.
3. The music generation method according to claim 1, characterized in that, The first user data includes at least one of the following: physiological data, facial expression data, and electroencephalogram (EEG) data; the physiological data includes at least one of the following: body temperature data, heart rate data, and user preference data; the user preference data includes at least one of the following: user preference for instrument size, user preference for instrument operation method, user preference for ambient sound, and user preference for scene. The step of determining user characteristics based on the first user data includes: Target features are determined based on the first user data; the target features include at least one of the following: physiological features, facial expression features, and electroencephalogram (EEG) features; the physiological features corresponding to the body temperature data include body temperature features, the physiological features corresponding to the heart rate data include heart rate features, and the physiological features corresponding to the user preference data include user preference features; When the target feature includes at least two of the physiological features, facial expression features, and brainwave features, at least two of the physiological features, facial expression features, and brainwave features are all input into the first feature fusion model to obtain the user feature output by the first feature fusion model. The target feature is obtained by extracting features from the first user data using a first feature extraction model; the first feature extraction model is trained based on the first user data sample.
4. The music generation method according to any one of claims 1-3, characterized in that, After generating the first target music data based on the music features and the user features, the method further includes: Play the first target music data; Obtain second user data; If the difference between the first user data and the second user data is less than or equal to a first preset threshold, the first user data is updated to the second user data, and the process returns to the step of determining user characteristics based on the first user data to obtain the second target music data. Replace the first target music data with the second target music data and play it.
5. A music generation device, characterized in that, include: The first acquisition module is used to acquire initial music data and first user data; The first user data was obtained while the initial music data was being played; The first determining module is used to determine music features based on the initial music data; The second determining module is used to determine user characteristics based on the first user data, wherein the user characteristics are the characteristics corresponding to the user when listening to the initial music data; The generation module is used to input the music features and the user features into the second feature fusion model, determine the cosine similarity of each feature contained in the music features based on the user features through the second feature fusion model, and perform a weighted summation of the cosine similarity and the corresponding features to obtain the music fusion features output by the second feature fusion model; The music fusion features are input into the generator of the target model, and the generator obtains the first target music data based on the music fusion features. The target model is trained based on music data samples and the corresponding labels of the music data samples; the labels include at least one of the following: instrument size label, instrument operation method label, ambient sound label, and scene label.
6. The music generation device according to claim 5, characterized in that, The first determining module is further configured to: The initial music data is input into the encoder of the target model to obtain the music vector output by the encoder; The music vector is input into the classifier of the target model to obtain the music features output by the classifier; the music features include at least one of the following: instrument size features, instrument operation features, ambient sound features, and scene features.
7. The music generation device according to claim 5, characterized in that, The first user data includes at least one of the following: physiological data, facial expression data, and electroencephalogram (EEG) data; the physiological data includes at least one of the following: body temperature data, heart rate data, and user preference data; the user preference data includes at least one of the following: user preference for instrument size, user preference for instrument operation method, user preference for ambient sound, and user preference for scene; the second determining module is further used for: Target features are determined based on the first user data; the target features include at least one of the following: physiological features, facial expression features, and electroencephalogram (EEG) features; the physiological features corresponding to the body temperature data include body temperature features. The physiological characteristics corresponding to the heart rate data include heart rate characteristics, and the physiological characteristics corresponding to the user preference data include user preference characteristics. When the target feature includes at least two of the physiological features, facial expression features, and brainwave features, at least two of the physiological features, facial expression features, and brainwave features are all input into the first feature fusion model to obtain the user feature output by the first feature fusion model. The target feature is obtained by extracting features from the first user data using a first feature extraction model; the first feature extraction model is trained based on the first user data sample.
8. The music generation apparatus according to any one of claims 5-7, characterized in that, Also includes: The playback module is used to play the first target music data; The second acquisition module is used to acquire second user data; The third determining module is used to update the first user data to the second user data and return to the step of determining user characteristics based on the first user data when the difference between the first user data and the second user data is less than or equal to a first preset threshold, so as to obtain the second target music data. The playback module is also used to replace the first target music data with the second target music data for playback.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the music generation method as described in any one of claims 1-4.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the music generation method as described in any one of claims 1-4.
Citation Information
Patent Citations
Music generation method and device, and electronic equipment
CN110853605A