Scenic spot adaptive music generation method and system based on multi-modal perception
Through multimodal perception technology and conditional control diffusion model, combined with reinforcement learning model, adaptive generation and real-time adjustment of scenic spot music are achieved, solving the problem that music in the existing technology cannot dynamically adapt to changes in scenic spots, and improving tourists' music experience.
Patent Information
- Application Number
- CN202510613412.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
AI Technical Summary
Existing scenic spot music usually relies on a preset single track or static music playlist, and cannot dynamically adapt to the changes in the scenic spot's regional, cultural background and environmental changes, and cannot make real-time adjustments based on tourists' feedback, which affects the immersion and pleasure of the tour.
Adaptive music generation method and system for scenic spots based on multimodal perception is adopted. By obtaining the geographical characteristics, cultural characteristics and music preference characteristics of the scenic spot space where tourists are located, combining the conditional control diffusion model and reinforcement learning model, music that meets the changes in the scenic spot space is generated in real time, and the condition variables generated by music are dynamically adjusted according to tourists' feedback.
It achieves a flexible fit between music and scenic spot scenes, enhances the personalization and immersion of the music experience, and ensures that different tourists get the best music experience in the scenic spot.
Smart Images

Figure CN120148447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and particularly to a method and system for generating scenic area adaptive music based on multi-modal perception. Background Art
[0002] With the continuous improvement of China's modernization level and people's living standards, people have higher pursuits in terms of the quality of services and products provided by scenic areas. As a constituent element of services, the environmental music in scenic areas has an interactive relationship with the surrounding things. When people listen to music, they can not only perceive the flow and rhythm of the music itself, but also unconsciously generate a spatial image of the music, and obtain information beyond the field of vision through hearing, greatly expanding people's experience of the scenic area space. As a carrier of regional culture, music is deeply rooted in the specific geographical, historical, and cultural backgrounds of scenic areas, and is often closely linked to local events, stories, traditional customs, and the natural environment. During the visit to the scenic area, if tourists can listen to music that highly matches the geographical location, environment, history, and cultural background of the scenic area, it is easier to arouse resonance, evoke positive emotional responses of tourists themselves, enhance their recognition and satisfaction with the scenic area, and then affect tourists' behavioral intentions, such as increasing tourists' enthusiasm for exploring the scenic area and sharing this tour experience with others. However, the existing scenic area music usually relies on a preset single track or static music playlist. Such general music solutions are mostly simple combinations of music types or styles, and fail to achieve dynamic adaptation to the geographical location, cultural background, and current environmental factors of the scenic area. At the same time, it is unable to make timely adjustments according to the feedback of tourists during the tour, resulting in an impact on the immersion and pleasure of the tour. Summary of the Invention
[0003] Based on this, it is necessary to provide a method and system for generating scenic area adaptive music based on multi-modal perception for the above technical problems.
[0004] A method for generating scenic area adaptive music based on multi-modal perception, the method includes: Obtain the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector of the scenic area space where the tourist is located.
[0005] Determine the initial values of the weights of the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector according to the tourist personal information data.
[0006] Weight the scenic area geographical feature vector, scenic area cultural feature vector, and music preference feature vector according to the weights, and then splice them to obtain a high-dimensional feature representation.
[0007] Taking the preset Gaussian noise and the high-dimensional feature representation as the input of the conditional controlled diffusion model, and using the high-dimensional feature representation as the control condition of the conditional controlled diffusion model for reverse diffusion to generate the Mel spectrogram features of the music data, and generating the music data according to the Mel spectrogram features.
[0008] Obtaining the feedback data of tourists on the music data, determining the reward data of the reinforcement learning model, taking the adjustment of the weights as the action of the reinforcement learning model to dynamically adjust the weights, and iteratively optimizing the generated music data until the preset termination condition is met.
[0009] A scenic area adaptive music generation system based on multi-modal perception, the system includes: A multi-modal information acquisition module, which is used to acquire the scenic area geographical feature vector, the scenic area cultural feature vector, the tourist music preference feature vector of the scenic area space where the tourists are located, and the tourist feedback data of the tourists on the generated music data.
[0010] A data processing and analysis module, which is used to determine the initial values of the weights of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector according to the tourist personal information data; weighting the scenic area geographical feature vector, the scenic area cultural feature vector, and the music preference feature vector according to the weights, and then splicing them to obtain a high-dimensional feature representation.
[0011] A music generation module, which is used to take the preset Gaussian noise and the high-dimensional feature representation as the input of the conditional controlled diffusion model, and use the high-dimensional feature representation as the control condition of the conditional controlled diffusion model for reverse diffusion to generate the Mel spectrogram features of the music data, and generate the music data according to the Mel spectrogram features.
[0012] A music optimization module based on tourist feedback data, which is used to determine the reward data of the reinforcement learning model according to the tourist feedback data, take the adjustment of the weights as the action of the reinforcement learning model to dynamically adjust the weights, and iteratively optimize the generated music data until the preset termination condition is met.
[0013] The above-mentioned scenic area adaptive music generation method and system based on multi-modal perception, the method generates music that fits the current scenic area space by real-time sensing the changes in the scenic area space and using a conditional controlled diffusion model. On the basis of music generation, the conditional variables of music generation are dynamically adjusted by using a reinforcement learning model according to tourist feedback, and the effect of the generated music is continuously optimized to ensure that the generated music can not only flexibly adapt to the changes in the scenic area space, but also be adjusted according to the feedback of tourists, so as to improve the fit degree of the music and the scenic area scene, and further make the music experience of different tourists in the scenic area reach the optimal state. Description of the Drawings
[0014] Figure 1Schematic flowchart of a scenic area adaptive music generation method based on multi-modal perception in an embodiment; Figure 2 Overall flowchart of a scenic area adaptive music generation method based on multi-modal perception in an embodiment; Figure 3 Block diagram of the composition of a scenic area adaptive music generation system based on multi-modal perception in an embodiment. Detailed implementation manner
[0015] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0016] In one embodiment, as Figure 1 shown, a scenic area adaptive music generation method based on multi-modal perception is provided, and this method includes the following steps: Step 100: Obtain the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector of the scenic area space where the tourist is located.
[0017] Specifically, the scenic area geographical feature vector is obtained by inputting the acquired scenic area geographical feature data into a large language model to obtain the text corresponding to the scenic area geographical feature data, and then processing the text corresponding to the scenic area geographical feature data using the CLAP pre-trained model.
[0018] The scenic area cultural feature vector is obtained by inputting the scenic area cultural feature data into a large language model to obtain the text corresponding to the scenic area cultural feature data; and then processing the text corresponding to the scenic area cultural feature data using the CLAP pre-trained model.
[0019] The music preference feature vector is obtained by processing the tourist music preference data using the CLAP pre-trained model.
[0020] When a tourist uses a preset travel APP, personal information will be input, and the personal information includes tourist music preference data.
[0021] Step 102: Determine the initial values of the weights of the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector according to the tourist personal information data.
[0022] Specifically, obtain the tourist personal information data in the preset travel APP, including gender, age, occupation, and music preference. The preset travel APP is the carrier / application program of "a scenic area adaptive music generation method based on multi-modal perception".
[0023] The weight of the scenic area geographical feature vector reflects the influence degree of the scenic area geographical feature data on the generated music; the weight of the scenic area cultural feature vector reflects the influence degree of the scenic area cultural feature data on the generated music; the weight of the tourist music preference feature vector reflects the influence degree of the tourist music preference data on the generated music.
[0024] One-hot encode the personal information of tourists to construct a tourist-weight interaction matrix in the recommendation system. In the tourist-weight interaction matrix, each tourist using the preset travel APP has the weight data of its unique scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector; for the current tourist using the preset travel APP, calculate the similarity between the current tourist and other tourists in the preset travel APP through the collaborative filtering algorithm, and take the average value of the weight data of the top five tourists with the highest similarity to the current tourist in the tourist-weight interaction matrix as the initial value of the weights of the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector of the current tourist.
[0025] Step 104: Weight the scenic area geographical feature vector, scenic area cultural feature vector, and music preference feature vector according to the weights, and then splice them to obtain a high-dimensional feature representation.
[0026] Specifically, according to the weights of the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector of the current tourist, weight the scenic area geographical feature vector 、the scenic area cultural feature vector and the tourist music preference feature vector respectively, and then splice the weighted results to obtain a high-dimensional feature representation, denoted as: ; ; ; ; Among them, is the high-dimensional feature representation, is the weight of the scenic area geographical feature vector, is the weight of the scenic area cultural feature vector, is the weight of the tourist music preference feature vector, 、 、 are the weighted scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector respectively.
[0027] High-dimensional feature representation It is a 3×768 feature matrix. Each row of the feature matrix corresponds to a feature vector. After weighting, the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector are each a row vector of the feature matrix.
[0028] Step 106: Use the preset Gaussian noise and the high-dimensional feature representation as the input of the conditional controlled diffusion model, perform reverse diffusion with the high-dimensional feature representation as the control condition of the conditional controlled diffusion model to generate the Mel spectrum features of the music data, and generate the music data according to the Mel spectrum features.
[0029] Step 108: Obtain the feedback data of tourists on the music data, determine the reward data of the reinforcement learning model, use the adjustment of the weights as the action of the reinforcement learning model to dynamically adjust the weights, and iteratively optimize the generated music data until the preset termination condition is met.
[0030] Preferably, the scenic area tourist feedback data is determined according to the satisfaction scores of tourists collected by the preset tourism APP for different scenic spots in the scenic area and the staying time of tourists at the current scenic spot.
[0031] The scenic area tourist feedback data includes: the subjective satisfaction of tourists with the music generation in the scenic area and the objective satisfaction of tourists with the music generation in the scenic area.
[0032] Determine the reward data according to the subjective satisfaction of tourists with the music generation in the scenic area and the objective satisfaction of tourists with the music generation in the scenic area. Use the reinforcement learning model according to the reward data to dynamically adjust the weights of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector, and update the weights of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector. Use the updated weights to update the control condition (i.e., the high-dimensional feature representation) of the conditional controlled diffusion model. Use the conditional controlled diffusion model to generate new music data, and perform iterative update on the generated music data until the preset termination condition is met.
[0033] In each iterative update process, the weighting process is performed on the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector obtained in step 100.
[0034] The preset termination condition is that the tourist has visited all the scenic spots in the scenic area, the tourist is physically exhausted, the closing time of the scenic area on the current day arrives, or the tourist exits the preset tourism APP and stops the background operation.
[0035] In the above-described scenic area adaptive music generation method based on multi-modal perception, the method generates music that fits the current scenic area space by real-time perceiving the changes in the scenic area space and using a conditional control diffusion model. Based on the music generation, the conditional variables for music generation are dynamically adjusted using a reinforcement learning model according to the feedback from tourists, continuously optimizing the effect of the generated music, ensuring that the generated music can not only flexibly adapt to the changes in the scenic area space but also be adjusted according to the feedback from tourists, thereby enhancing the degree of fit between the music and the scenic area scene, and further enabling the music experience of different tourists in the scenic area to reach the optimal state.
[0036] In one embodiment, before step 100, it further includes: obtaining the scenic area geographical feature data, scenic area cultural feature data, and tourist music preference data of the scenic area space where the tourist is located; after processing the scenic area geographical feature data through a large language model, obtaining the text corresponding to the scenic area geographical feature data; processing the text corresponding to the scenic area geographical feature data using a CLAP pre-trained model to obtain the scenic area geographical feature vector; after processing the scenic area cultural feature data through a large language model, obtaining the text corresponding to the scenic area cultural feature data; processing the text corresponding to the scenic area cultural feature data using a CLAP pre-trained model to obtain the scenic area cultural feature vector; processing the tourist music preference data using a CLAP pre-trained model to obtain the tourist music preference feature vector.
[0037] Specifically, the scenic area geographical feature data refers to the geographical feature data of each scenic spot in the scenic area, including: topography and landforms, such as mountains, plains, plateaus, hills, etc.; climate and meteorology, such as temperature, precipitation, sunlight, monsoon, humidity, air pressure, etc.; hydrology and water systems, such as rivers, lakes, oceans, waterfalls, glaciers, etc.; soil and vegetation, such as forests, grasslands, deserts, wetlands, etc.; geographical location, such as longitude, latitude, administrative region, relative position, etc.; time period categories, such as festivals, seasons, time, etc.
[0038] The scenic area cultural feature data refers to the cultural feature data of each scenic spot in the scenic area, including: historical and cultural categories, such as historical events, legends, etc.; traditional custom categories, such as festival customs, food customs, clothing customs, intangible cultural heritage, etc.; architectural landscape categories, such as architectural structure, material characteristics, etc.; language and cultural categories, such as dialect characteristics, poetic rhythms, proverbs and folk songs, etc.
[0039] When a tourist uses a preset travel APP, personal information will be input, and the personal information includes tourist music preference data.
[0040] After processing the scenic area geographical feature data through a large language model, the text corresponding to the scenic area geographical feature data is obtained; the text corresponding to the scenic area geographical feature data is processed using a CLAP pre-trained model to obtain the scenic area geographical feature vector; after processing the scenic area cultural feature data through a large language model, the text corresponding to the scenic area cultural feature data is obtained; the text corresponding to the scenic area cultural feature data is processed using a CLAP pre-trained model to obtain the scenic area cultural feature vector; the tourist music preference data is processed using a CLAP pre-trained model to obtain the tourist music preference feature vector.
[0041] In one embodiment, the conditional controlled diffusion model includes: a VAE pre-trained model, a UNet+CrossAttention model. Step 108 includes: mapping the Mel spectrogram features of the music data to a low-dimensional latent feature space through the encoder in the VAE pre-trained model to obtain the low-dimensional latent feature vector corresponding to the Mel spectrogram features of the music data; in the forward diffusion process of the conditional controlled diffusion model, adding Gaussian noise at each time step to the low-dimensional latent feature vector to obtain Gaussian noise; in the reverse diffusion process of the conditional controlled diffusion model, using the high-dimensional feature representation as the control condition of the conditional controlled diffusion model, and gradually predicting the Gaussian noise added to the low-dimensional latent vector in the forward diffusion process of the conditional controlled diffusion model using the UNet+Cross Attention model to generate the low-dimensional latent feature vector after removing the noise; mapping the low-dimensional latent feature vector after removing the noise through the decoder in the VAE pre-trained model to obtain the Mel spectrogram features of the reconstructed music data; inputting the Mel spectrogram features into the HiFi-GAN pre-trained model to generate music data.
[0042] Specifically, the music data is the music data obtained from an open-source music dataset.
[0043] The specific process of generating music data includes: Step 1: Select the scenic area geographical feature data from the scenic area geographical feature data - music dataset and perform feature extraction on the scenic area geographical feature data to obtain the scenic area geographical feature vector .
[0044] Step 2: Select the scenic area cultural feature data from the scenic area cultural feature data - music dataset and perform feature extraction on the scenic area cultural feature data to obtain the scenic area cultural feature vector .
[0045] Step 3: Select the tourist music preference data from the tourist music preference data - music dataset , extract features from the music preference data of tourists to obtain the music preference feature vector of tourists .
[0046] Step Four: Determine the obtained music data Mel spectrogram features , and obtain the Mel spectrogram features of the music data through short-time Fourier transform, Mel frequency scale, and Mel filter bank Mel spectrogram features .
[0047] Step Five: Build a conditional control diffusion model, which includes: a VAE pre-training model and a UNet+Cross Attention model.
[0048] Step Six: Train the conditional control diffusion model. First, use the encoder in the VAE pre-training model to map the Mel spectrogram features of the music data to a low-dimensional latent feature vector , where represents the number of channels of the Mel spectrogram features , represents the compression level of the encoder , represents the time dimension of the Mel spectrogram features , represents the frequency dimension of the Mel spectrogram features . Then, during the forward diffusion process of the conditional control diffusion model, add Gaussian noise for time steps to the latent feature vector
[0049] The forward diffusion process is defined as a Markov chain process, denoted as: ; ; where, , , represents the noise diffusion coefficient at each time step , represents the low-dimensional latent vector at the 0th time step, represents the low-dimensional latent vector after adding noise at the th time step, is the standard normal Gaussian noise.
[0050] After the forward diffusion process ends, perform the reverse diffusion process. Here, take the reverse diffusion process of the DDPM model as an example.
[0051] The purpose of the reverse diffusion process is to process the noise data Denoise step by step and represent it with high-dimensional features The weighted scenic area geographical feature vector and the weighted scenic area cultural feature vector and the weighted tourist music preference feature vector As control conditions, use the UNet+Cross Attention model to predict the noisy data .
[0052] The n noisy low-dimensional latent vector at the th time step, after passing through several ResNet blocks, the Mel spectrogram feature map of an intermediate layer of UNet is , where can be divisible by the number of channels of the Mel spectrogram feature . Divide evenly
[0053] Flatten the spatial dimension of to obtain , where .
[0054] In the Cross Attention layer, pass through the matrix to generate the Query matrix , pass the high-dimensional feature representation through the matrix to generate the Key matrix , pass the high-dimensional feature representation through the matrix to generate the Value matrix , denoted as: ; ; ; Among them, , , , , , .
[0055] Calculate the similarity matrix between the Query matrix and the Key matrix , denoted as: ; Perform Softmax normalization on the similarity matrix to obtain the normalized similarity matrix , denoted as: .
[0056] The similarity matrix and the Value matrix are calculated to obtain the context matrix , denoted as: .
[0057] Through the linear projection matrix is mapped to , denoted as: ; Among them, , .
[0058] The is mapped to the spatial dimension where the Mel spectrogram feature is located to obtain , denoted as: ; The is fused with the original Mel spectrogram feature to obtain , denoted as: .
[0059] The entire training process of the conditional control diffusion model is to use the predicted noise data to fit the noise added at each time step , and the loss function of the conditional control diffusion model is: .
[0060] The final sampling formula for the reverse diffusion process, denoted as: ; Among them, .
[0061] In one embodiment, step 102 includes: obtaining the personal information data of different tourists, the weights of the scenic area geographical feature vectors, the weights of the scenic area cultural feature vectors, and the weights of the tourists' music preference feature vectors in a preset travel APP; the personal information data of tourists includes gender, age, occupation, and music preference; constructing a tourist-weight interaction matrix in the recommendation system according to the personal information data of all tourists, the weight vectors of the scenic area geographical feature vectors, the weights of the scenic area cultural feature vectors, and the weights of the tourists' music preference feature vectors in the preset travel APP; performing one-hot encoding on the personal information data of each tourist in the tourist-weight interaction matrix, and then splicing them to obtain the corresponding personal information feature vectors; for the tourist currently using the preset travel APP, calculating the similarity between the current tourist and other tourists in the preset travel APP according to the personal information feature vectors by using the collaborative filtering algorithm; taking the average value of the weights of the scenic area geographical feature vectors, the weights of the scenic area cultural feature vectors, and the weights of the tourists' music preference feature vectors of the top five tourists with the highest similarity to the current tourist in the tourist-weight interaction matrix as the initial values of the weights of the scenic area geographical feature vectors, the scenic area cultural feature vectors, and the tourists' music preference feature vectors of the current tourist.
[0062] The initial values of the weights of the scenic area geographical feature vectors, the scenic area cultural feature vectors, and the tourists' music preference feature vectors of the current tourist are: .
[0063] Among them, represents the weights of the scenic area geographical feature vectors, the scenic area cultural feature vectors, and the tourists' music preference feature vectors of the i th tourist.
[0064] In one embodiment, the similarity between the current tourist and other tourists in the preset travel APP is: .
[0065] Among them, represents the similarity between the current tourist and other tourists in the preset travel APP, represents the current tourist, represents the personal information feature vector of the current tourist, represents other tourists in the preset travel APP, represents the personal information feature vectors of other tourists in the preset travel APP.
[0066] In one embodiment, the tourist feedback data includes: the subjective satisfaction and objective satisfaction of tourists with respect to the scenic area music generation. Step 108 includes: collecting tourist feedback through the music satisfaction scoring function in the preset tourism APP to obtain the subjective satisfaction of tourists with respect to the scenic area music generation; collecting and calculating the stay duration of tourists at this scenic spot through the positioning function of the preset tourism APP to determine the objective satisfaction of tourists with respect to the scenic area music generation; determining the reward data according to the subjective satisfaction and objective satisfaction of tourists with respect to the scenic area music generation, and the reward data is: 。
[0067] Wherein, represents the reward data, represents the subjective satisfaction of tourists with respect to the scenic area music generation, and its range is [-1, 1]; represents the objective satisfaction of tourists with respect to the scenic area music generation, and its range is [-1, 1]; represents the proportion of the subjective satisfaction of tourists with respect to the scenic area music generation in the reward data, represents the proportion of the objective satisfaction of tourists with respect to the scenic area music generation in the reward data, ; According to the reward data, the adjustment of the weights is used as the action of the reinforcement learning model to dynamically adjust the weights of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector, and iteratively optimize the generated music data until the preset termination condition is met.
[0068] In one embodiment, the subjective satisfaction of tourists with respect to the scenic area music generation is: 。
[0069] Wherein, represents the tourist satisfaction score, and its range is , and the satisfaction with the music played by the preset tourism APP is scored (such as a 1-5 or 1-10 scale), represents the maximum value of the tourist satisfaction score.
[0070] The objective satisfaction of tourists with respect to the scenic area music generation is: ; Wherein, represents the objective satisfaction of tourists with respect to the scenic area music generation, the range of , represents the stay duration of tourists, that is, the stay duration of tourists at a certain scenic spot, and its range is , represents the longest stay time of all tourists in the scenic area spots in the preset tourism APP.
[0071] In one of the embodiments, the training process of the reinforcement learning model includes: constructing a scenic area music optimization dataset; obtaining training samples from the scenic area music optimization dataset, and each training sample includes current state data, action data, reward data, next state data, and a preset termination condition ; building a reinforcement learning model; the reinforcement learning model adopts an Actor-Critic model, which includes a Critic target network, a Critic training network, an Actor target network, and an Actor training network; the Critic target network is used to approximately estimate the state-action Q value function at the next moment; the Q target value of the value function in the current state is: ; where represents the target value of the value function in the current state, Q represents the reward data, represents the discount rate, represents the state-action value function at the next moment, Q represents the action value at the next moment, represents the parameters of the Actor target network, represents the parameters of the Critic target network; The Critic training network is used to evaluate the current policy and output the state-action Q value function. Among them, the loss function during the update of the Critic training network is: ; where represents the loss function during the update of the Critic training network, represents the number of a batch of training samples sampled from the experience replay pool, represents the total number of samples in the experience replay pool, w represents the parameters of the Critic training network, represents the Q state-action value function; The Actor target network is used to provide the action value at the next moment ; The Actor training network is used to provide the policy of the current state. Combining with the Q value function of the Critic training network, the policy gradient during the parameter update of the Actor training network is: ; Among them, represents the parameters of the Actor training network, represents the state, represents the action, represents the scoring direction of the action by the Critic training network, represents the sensitivity of the Actor training network to the action; Adopt d A set of training samples is used to train the reinforcement learning model, and the model parameters are saved to obtain the trained reinforcement learning model. Among them, , represents the floor operator.
[0072] In one of the embodiments, the process of constructing the scenic area music optimization dataset includes: The current state data is: .
[0073] Among them, is the current state data, , , are the weights of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector respectively; The action data is: ; ; ; ; .
[0074] Among them, is the action data of the current state, , , are the weight floating values of the standardized scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector respectively, , , are the weight floating values of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector respectively, is the mean value of the initial action data; The initial action data obtains the action data through zero-sum projection, satisfying the constraint ; Determine the reward data according to the subjective satisfaction and objective satisfaction of the scenic area music generated by tourists in the tourist feedback data; determine the initial data of the next state according to the current state data and action data, and perform non - negativity correction on the initial data of the next state to obtain the next state data as follows: ; ; wherein, is the next state data, , , are respectively the weights of the scenic area geographical feature vector, scenic area cultural feature vector, and tourist music preference feature vector after non - negativity correction; Determine the preset termination condition according to one of the following conditions: the tourist has visited all the scenic spots in the scenic area, the tourist is physically exhausted, the closing time of the scenic area on the current day arrives, and the tourist exits the preset tourism APP and stops running in the background; construct a scenic area music optimization data set based on the experience replay pool, and the experience replay pool contains M training samples ; where represents the preset termination condition.
[0075] The overall process of scenic area adaptive music generation based on multi - modal perception is as Figure 2 shown.
[0076] It should be understood that although each step in the flowchart of Figure 1 is shown in sequence according to the arrow indication, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0077] may include multiple sub - steps or multiple stages. These sub - steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub - steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub - steps or stages of other steps. Figure 3 In one embodiment, as shown, a scenic area adaptive music generation system based on multi - modal perception is provided. The system includes: a multi - modal information acquisition module, a data processing and analysis module, a music generation module, and a music optimization module based on tourist feedback data, wherein: The multi - modal information acquisition module is used to acquire the scenic area geographical feature vector, scenic area cultural feature vector, tourist music preference feature vector of the scenic area space where the tourist is located, and the tourist feedback data of the tourist for the generated music data.
[0078] A data processing and analysis module, which is used to determine the initial values of the weights of the scenic area geographical feature vector, the scenic area cultural feature vector, and the tourist music preference feature vector according to the tourist personal information data; weight the scenic area geographical feature vector, the scenic area cultural feature vector, and the music preference feature vector according to the weights, and then splice them to obtain a high-dimensional feature representation.
[0079] Specifically, the working process of the data processing and analysis module includes: Input the scenic area geographical feature data into the large language model to obtain its corresponding text data.
[0080] Input the scenic area cultural feature data into the large language model to obtain its corresponding text data.
[0081] Obtain the tourist music preference data.
[0082] Pre-construct a scenic area geographical feature data - music dataset. The specific method is to determine the correspondence between the scenic area geographical feature data and the music data, and construct the scenic area geographical feature data - music dataset.
[0083] Pre-construct a scenic area cultural feature data - music dataset. The specific method is to determine the correspondence between the scenic area cultural feature data and the music data, and construct the scenic area cultural feature data - music dataset.
[0084] Pre-construct a tourist music preference data - music dataset. The specific method is to determine the correspondence between the tourist music preference data and the music data, and construct the tourist music preference data - music dataset.
[0085] Pre-construct a scenic area music optimization dataset for reinforcement learning training. The specific method is to determine the current state data, action data, reward data, next state data, and preset termination conditions. Among them, the reward data is determined based on the tourist feedback data, and the preset termination conditions include that the tourist has visited all the scenic spots in the scenic area, the tourist is physically exhausted, the closing time of the scenic area on the current day has arrived, and the tourist exits the preset travel APP and stops running in the background.
[0086] The construction process of the scenic area data - music dataset includes: Step P1: Input the scenic area geographical feature data and the scenic area cultural feature data into the large language model to obtain the text data corresponding to the scenic area geographical features and the text data corresponding to the scenic area cultural features respectively.
[0087] Step P2 constructs a scenic spot geographic feature data-music data set based on the correspondence between the scenic spot geographic feature data and the music data; constructs a scenic spot cultural feature data-music data set based on the correspondence between the scenic spot cultural feature data and the music data; and constructs a tourist music preference data-music data set based on the correspondence between the tourist music preference data and the music data.
[0088] The music generation module is used to use the preset Gaussian noise and high-dimensional feature representation as the input of the conditional controlled diffusion model, and the high-dimensional feature representation is used as the control condition of the conditional controlled diffusion model for reverse diffusion to generate the Mel spectrum feature of the music data, and generate the music data according to the Mel spectrum feature.
[0089] The music optimization module based on tourist feedback data is used to determine the reward data of the reinforcement learning model according to the tourist feedback data, dynamically adjust the weights by taking the adjustment of the weights as the action of the reinforcement learning model, and iteratively optimize the generated music data until the preset termination conditions are met.
[0090] Specifically, the working process of the music optimization module based on tourist feedback data includes: Step S1 collects tourists’ feedback through the music satisfaction scoring function in the preset tourism APP to obtain tourists’ subjective satisfaction with the music generated by the scenic spot. .
[0091] Step S2: Collect and calculate the length of time tourists stay at this scenic spot through the preset tourism APP positioning function, and determine the objective satisfaction of tourists with the music generated by the scenic spot .
[0092] Step S3: Fusion of objective satisfaction and subjective satisfaction data to obtain an overall satisfaction score as reward data , written as: .
[0093] Step S4: Based on the reward data ,The trained reinforcement learning model is used to adjust the weight data of the scenic spot geographical feature vector, the scenic spot cultural feature vector and the tourists’ music preference feature vector in music generation.
[0094] Step S5: Repeat steps S1 to S4 until the tourists leave the scenic area, so as to generate personalized music for different tourists in the scenic area, thereby achieving the best music experience for different tourists in the scenic area.
[0095] For the specific limitations of the scenic area adaptive music generation system based on multi-modal perception, reference may be made to the limitations of the scenic area adaptive music generation method based on multi-modal perception in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned scenic area adaptive music generation system based on multi-modal perception can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0096] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0097] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for generating scenic area adaptive music based on multimodal perception, characterized in that: The method comprises: Obtaining the scenic area geographic feature vector, scenic area cultural feature vector and tourist music preference feature vector of the scenic area space where tourists are located; Determine the initial values of the weights of the scenic spot geographic feature vector, the scenic spot cultural feature vector, and the tourist music preference feature vector according to the tourist personal information data; The scenic spot geographical feature vector, the scenic spot cultural feature vector and the music preference feature vector are weighted according to the weights, and then concatenated to obtain a high-dimensional feature representation; Using the preset Gaussian noise and the high-dimensional feature representation as inputs of a conditional controlled diffusion model, using the high-dimensional feature representation as a control condition of the conditional controlled diffusion model for reverse diffusion, generating Mel-spectrogram features of music data, and generating music data according to the Mel-spectrogram features; Acquire tourists' feedback data on the music data, determine reward data of the reinforcement learning model, dynamically adjust the weights by taking the adjustment of the weights as the action of the reinforcement learning model, and iteratively optimize the generated music data until a preset termination condition is met.
2. The method for generating scenic area adaptive music based on multimodal perception according to claim 1, characterized in that: The step of obtaining the scenic spot geographic feature vector, the scenic spot cultural feature vector and the tourist music preference data feature vector of the scenic spot space where the tourist is located also includes: Obtaining the scenic area geographical feature data, scenic area cultural feature data and tourists' music preference data of the scenic area space where tourists are located; After the scenic spot geographical feature data is processed by a large language model, a text corresponding to the scenic spot geographical feature data is obtained; The text corresponding to the scenic spot geographical feature data is processed using the CLAP pre-training model to obtain the scenic spot geographical feature vector; After the scenic spot cultural feature data is processed by a large language model, a text corresponding to the scenic spot cultural feature data is obtained; The text corresponding to the scenic spot cultural feature data is processed using the CLAP pre-training model to obtain the scenic spot cultural feature vector; The tourists' music preference data is processed using the CLAP pre-training model to obtain a tourists' music preference feature vector.
3. The method for generating scenic area adaptive music based on multimodal perception according to claim 1 is characterized in that: The conditional controlled diffusion model includes: VAE pre-training model, UNet+Cross Attention model; The method comprises: using the preset Gaussian noise and the high-dimensional feature representation as inputs of a conditional controlled diffusion model, performing reverse diffusion using the high-dimensional feature representation as a control condition of the conditional controlled diffusion model, generating a Mel-spectrum feature of music data, and generating music data according to the Mel-spectrum feature, including: The Mel-spectrogram features of the music data are mapped to a low-dimensional latent feature space through the encoder in the VAE pre-trained model, and a low-dimensional latent feature vector corresponding to the Mel-spectrogram features of the music data is obtained; In the forward diffusion process of the conditional controlled diffusion model, add Gaussian noise of time steps, get Gaussian noise; In the reverse diffusion process of the conditional controlled diffusion model, the high-dimensional feature representation is used as the control condition of the conditional controlled diffusion model, and the UNet+Cross Attention model is used to gradually predict the Gaussian noise added to the low-dimensional latent vector in the forward diffusion process of the conditional controlled diffusion model to generate a low-dimensional latent feature vector after the noise is removed; The low-dimensional latent feature vector after noise removal is mapped through the decoder in the VAE pre-trained model to obtain the Mel spectrum features of the reconstructed music data; The Mel spectrum features are input into the HiFi-GAN pre-trained model to generate music data.
4. The method for generating scenic area adaptive music based on multimodal perception according to claim 1, characterized in that: Determining the initial values of the weights of the scenic spot geographic feature vector, the scenic spot cultural feature vector, and the tourist music preference feature vector according to the tourist personal information data includes: Obtaining the personal information data of different tourists in a preset travel APP, the weight of the scenic spot geographic feature vector, the weight of the scenic spot cultural feature vector, and the weight of the tourist music preference feature vector; the tourist personal information data includes gender, age, occupation, and music preference; According to the personal information data of all tourists in the preset travel APP, the weight of the scenic spot geographical feature vector, the weight of the scenic spot cultural feature vector, and the weight of the tourist music preference feature vector, the tourist-weight interaction matrix in the recommendation system is constructed; The personal information data of each tourist in the tourist-weight interaction matrix is one-hot encoded and then concatenated to obtain a corresponding personal information feature vector; For the tourist currently using the preset travel APP, a collaborative filtering algorithm is used to calculate the similarity between the current tourist and other tourists in the preset travel APP according to the personal information feature vector; The average of the weights of the scenic spot's geographical feature vectors, the weights of the scenic spot's cultural feature vectors, and the weights of the tourist's music preference feature vectors of the top five tourists with the highest similarity to the current tourist in the tourist-weight interaction matrix is used as the initial value of the weights of the scenic spot's geographical feature vector, the scenic spot's cultural feature vector, and the tourist's music preference feature vector of the current tourist.
5. The method for generating scenic area adaptive music based on multimodal perception according to claim 4 is characterized in that: The similarity between the current tourist and other tourists in the preset travel APP is: in, Indicates the similarity between the current tourist and other tourists in the preset travel APP, Indicates the current visitor, Represents the personal information feature vector of the current tourist, Indicates other tourists in the preset travel APP, Represents the personal information feature vector of other tourists in the preset travel APP.
6. The method for generating scenic area adaptive music based on multimodal perception according to claim 1, characterized in that: Tourist feedback data include: tourists’ subjective and objective satisfaction with scenic spot music generation; Acquiring tourists' feedback data on the music data, determining reward data of the reinforcement learning model, dynamically adjusting the weights by taking the adjustment of the weights as an action of the reinforcement learning model, and iteratively optimizing the generated music data until a preset termination condition is met, including: Collect tourists’ feedback through the music satisfaction scoring function in the preset tourism APP to obtain tourists’ subjective satisfaction with the music generated by the scenic spot; The preset positioning function of the tourism APP is used to collect and calculate the length of time tourists stay at this scenic spot, and determine the objective satisfaction of tourists with the music generated by the scenic spot; According to the subjective satisfaction and objective satisfaction of tourists to the music generated by the scenic spot, the reward data is determined, and the reward data is: in, Represents reward data, It represents tourists’ subjective satisfaction with the music generated by the scenic spot, and its range is [-1,1]; It represents the objective satisfaction of tourists with the music generated by the scenic spot, and its range is [-1,1]; It represents the proportion of tourists’ subjective satisfaction with the scenic spot music generation in the reward data, It represents the proportion of tourists’ objective satisfaction with the scenic spot music generation in the reward data, ; According to the reward data, the weight adjustment is used as the action of the reinforcement learning model to dynamically adjust the weights of the scenic spot geographical feature vector, the scenic spot cultural feature vector and the tourist music preference feature vector, and the generated music data is iteratively optimized until the preset termination condition is met.
7. The method for generating scenic area adaptive music based on multimodal perception according to claim 6, characterized in that: The subjective satisfaction of tourists with the music generated by the scenic spot is: in, It indicates the tourists' satisfaction score, which ranges from , Indicates the maximum value of tourists’ satisfaction score; The objective satisfaction of tourists with the music generated by the scenic spot is: in, Indicates the length of stay of tourists, ranging from , Indicates the longest time that all tourists in the preset travel APP stay in scenic spots.
8. The method for generating scenic area adaptive music based on multimodal perception according to claim 1, characterized in that: The training process of the reinforcement learning model includes: Construct a scenic spot music optimization dataset; Obtained from the scenic spot music optimization dataset training samples, each of which contains current state data, action data, reward data, next state data, and preset termination conditions. ; Build a reinforcement learning model; the reinforcement learning model adopts an Actor-Critic model, including a Critic target network, a Critic training network, an Actor target network and an Actor training network; The critic target network is used to estimate the state-action at the next moment. Q Value function; current state Q The target value of the value function is: in, Indicates the current state Q The target value of the value function, Represents reward data, represents the discount rate, Indicates the state-action of the next moment Q Value function, represents the action value at the next moment, Represents the parameters of the Actor target network, Represents the parameters of the Critic target network; The Critic training network is used to evaluate the current strategy and output the current state-action Q Value function, where the loss function when the Critic training network is updated is: in, Represents the loss function when the Critic training network is updated, represents the number of training samples sampled from the experience replay pool, Represents the total number of samples in the experience replay pool, w Represents the parameters of the Critic training network, Indicates the current state-action Q Value function; The Actor target network is used to provide the action value at the next moment ; The Actor training network is used to provide the strategy for the current state, combined with the Critic training network Q The value function obtains the policy gradient of the Actor training network during parameter update as: in, Represents the parameters of the Actor training network, Indicates the status, Indicates action, Indicates the scoring direction of the action by the Critic training network, Indicates the sensitivity of the Actor training network to the action; use d The reinforcement learning model is trained by using a set of training samples, and the model parameters are saved to obtain a trained reinforcement learning model, wherein: , Represents the floor operator.
9. The method for generating scenic area adaptive music based on multimodal perception according to claim 8, characterized in that: The process of constructing the scenic spot music optimization dataset includes: The current status data is: in, is the current status data, , , are the weights of the scenic spot geographical feature vector, the scenic spot cultural feature vector and the tourist music preference feature vector; The action data is: in, is the action data of the current state, , , are the weighted floating values of the standardized scenic spot geographic feature vector, scenic spot cultural feature vector, and tourist music preference feature vector, respectively. , , are the weighted floating values of the scenic spot geographical feature vector, the scenic spot cultural feature vector, and the tourist music preference feature vector, respectively. is the mean of the initial action data; the initial action data Get action data through zero-sum projection , satisfying the constraints ; Determine reward data based on tourists’ subjective and objective satisfaction with scenic spot music in tourists’ feedback data; According to the current state data and action data, the initial data of the next state is determined, and the initial data of the next state is non-negatively corrected to obtain the next state data: in, is the next state data, , , are the weights of the scenic spot geographical feature vector, scenic spot cultural feature vector and tourists' music preference feature vector after non-negativity correction; The preset termination condition is determined according to one of the following conditions: the tourist has visited all the scenic spots in the scenic area, the tourist is physically exhausted, the scenic area has reached the closing time of the day, and the tourist has exited the preset travel APP and stopped the background operation; The scenic spot music optimization dataset is constructed based on the experience playback pool, which contains M training samples ;in Indicates the preset termination condition.
10. A scenic area adaptive music generation system based on multimodal perception, characterized in that: The system comprises: A multimodal information acquisition module, used to acquire the scenic spot geographic feature vector, the scenic spot cultural feature vector, the tourist music preference feature vector, and the tourist feedback data for the generated music data of the scenic spot space where the tourist is located; The data processing and analysis module is used to determine the initial values of the weights of the scenic spot geographic feature vector, the scenic spot cultural feature vector and the tourist music preference feature vector according to the tourist personal information data; weight the scenic spot geographic feature vector, the scenic spot cultural feature vector and the music preference feature vector according to the weights, and then splice them to obtain a high-dimensional feature representation; A music generation module, used to use the preset Gaussian noise and the high-dimensional feature representation as inputs of a conditional controlled diffusion model, perform reverse diffusion of the high-dimensional feature representation as a control condition of the conditional controlled diffusion model, generate Mel-spectrogram features of music data, and generate music data according to the Mel-spectrogram features; The music optimization module based on tourist feedback data is used to determine the reward data of the reinforcement learning model according to the tourist feedback data, dynamically adjust the weights by taking the adjustment of the weights as the action of the reinforcement learning model, and iteratively optimize the generated music data until the preset termination condition is met.
Citation Information
Patent Citations
Unmanned aerial vehicle ground target tracking and obstacle avoidance planning method based on deep reinforcement learning
CN117707207A
User portrait-based gift recommendation method and system
CN118917929A
Personalized recommendation method for plasticized products based on deep reinforcement learning
CN119205257A
Interactive music intelligent generation method and system adaptive to scene space
CN119558356A
Advertisement pushing method and system based on user behavior analysis
CN119887302A
Cited By
Self-adaptive acousto-optic adjustment method and system based on respiration data and environmental perception
CN120514984A
An Adaptive Acoustic-Optical Modulation Method and System Based on Respiratory Data and Environmental Perception
CN120514984B