Music playing recommendation method and device, electronic equipment, storage medium and computer program product

By collecting and quantifying consumer characteristics and using a pre-defined strategy model to determine personalized music recommendation strategies, the problem of insufficient personalization in music recommendations in shopping mall environments has been solved, thus improving the consumer experience.

CN121636744APending Publication Date: 2026-03-10CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411203614.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

When existing technology broadcasts music in shopping mall environments, it ignores consumers' personalized needs, resulting in low personalization of music recommendations.

Method used

By collecting and quantifying consumers' behavioral characteristics, location characteristics, physiological state characteristics, time characteristics, and music preference characteristics within a preset area, a personalized music recommendation strategy is determined using a preset strategy model, and target recommended music is selected from the music library based on weight parameters.

Benefits of technology

It enables personalized music recommendations based on multiple consumer characteristics, meeting the common needs of multiple users and enhancing the consumer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636744A_ABST
    Figure CN121636744A_ABST
Patent Text Reader

Abstract

The invention provides a played music recommendation method and device, electronic equipment, a storage medium and a computer program product, and relates to the technical field of music playing, and the method comprises the steps: determining a current feature corresponding to each current object in a preset region; wherein the current feature comprises at least one of the following features for the current object: a behavior feature, a position feature, a physiological status feature, a time feature and a music preference feature; inputting the current feature into a preset strategy model to determine a music recommendation strategy corresponding to the current object; and determining target recommended music based on the music recommendation strategy. The music recommended through the scheme in the embodiment of the invention has high individuation, and the recommended target music can meet public requirements of multiple users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of music playing, and in particular to a music playing recommendation method and device, electronic equipment, a storage medium and a computer program product. BACKGROUND

[0002] In a shopping mall environment, public music not only helps to create a comfortable shopping atmosphere, but also can improve the consumer experience through accurate recommendation. However, the existing technology often only plays public music according to a fixed list order and the like, often ignoring the personalized needs of consumers, resulting in low personalized degree of recommended music. SUMMARY

[0003] The music playing recommendation method, device, electronic equipment, storage medium and computer program product provided by the embodiments of the present application can recommend music with a high personalized degree.

[0004] The technical solution of the present application is implemented as follows:

[0005] The embodiments of the present application provide a music playing recommendation method, comprising:

[0006] determining a current feature corresponding to each current object in a preset area; wherein the current feature comprises at least one of the following features for the current object: behavior feature, position feature, physiological state feature, time feature and music preference feature;

[0007] inputting the current feature into a preset strategy model to determine a music recommendation strategy corresponding to the current object;

[0008] determining target recommended music based on the music recommendation strategy.

[0009] In the above solution, before the current feature is input into the preset strategy model to determine the music recommendation strategy corresponding to the current object, the method further comprises:

[0010] determining sample data respectively corresponding to a plurality of historical objects in the preset area; wherein each sample data comprises: historical features, a historical recommendation strategy and reward and punishment parameters of recommended music for the historical recommendation strategy of each historical object in an adjacent time step; wherein the historical features are used to reflect at least one of the following features of the historical object in the corresponding time step: historical behavior feature, historical position feature, historical physiological state feature, historical time feature and historical music preference feature;

[0011] training an initial strategy model based on each sample data, and stopping training when a predetermined training condition is reached, to obtain the preset strategy model.

[0012] The historical features include: first historical features and second historical features.

[0013] The first historical features of each historical object in the preset area within a first time step are determined.

[0014] The historical recommendation strategy is determined based on the first historical features by using a preset exploration method, wherein the historical recommendation strategy includes a weight parameter for reference when music is recommended.

[0015] The reward and punishment parameters are determined for the recommended music based on the historical recommendation strategy.

[0016] The second historical features of each historical object in the preset area within a second time step are determined.

[0017] The first historical features of each historical object in the preset area within a first time step are determined.

[0018] The historical behavior data, historical location data, historical physiological image data, historical time data and historical music preference data of each historical object in the preset area within the first time step are determined.

[0019] The historical behavior data, historical location data, historical physiological image data and historical music preference data are subjected to data quantization processing to determine the first historical features.

[0020] The historical behavior data is used to reflect the transaction behavior of the historical object within the first time step.

[0021] The historical location data is used to reflect the position of the historical object in the preset area and the object flow corresponding to the position within the first time step.

[0022] The historical physiological image data is used to reflect the physiological state change of the historical object within the first time step.

[0023] The historical time data is used to reflect the time and season when the historical object enters the preset area within the first time step.

[0024] The music preference data is used to reflect the music preference of the historical object in a third-party software.

[0025] The reward and punishment parameters are determined for the recommended music based on the historical recommendation strategy.

[0026] determine the reward and punishment parameter based on the copyright information of the recommended music and the feedback information of the historical object for the recommended music.

[0027] In the foregoing solution, the feedback information comprises behavior feedback information and physiological feedback information; and the determination of the reward and punishment parameter based on the copyright information of the recommended music and the feedback information of the historical object for the recommended music comprises:

[0028] determining a feedback weight and a preference degree information of the historical object for the recommended music based on the behavior feedback information and the physiological feedback information.

[0029] determining the reward and punishment parameter based on the feedback weight, the preference degree information and a copyright reward coefficient; wherein the copyright reward coefficient is related to the compliance represented by the copyright information.

[0030] In the foregoing solution, the historical recommendation strategy comprises at least one of the following weight parameters: a music type preference weight, a user interaction feedback weight, a copyright compliance weight, an environment adaptability weight, a time sensitivity weight and a user physiological response weight.

[0031] In the foregoing solution, the initial strategy model comprises an initial memory network, an initial evaluation network and an initial target network; and the training of the initial strategy model based on each sample data until the training is stopped when a predetermined training condition is reached to obtain the preset strategy model comprises:

[0032] storing each sample data in the initial memory network.

[0033] inputting a first historical feature in the first sample data, the reward and punishment parameter and a first historical recommendation strategy into the initial evaluation network to determine an expected return, and inputting the second historical feature in the first sample data into the initial target network to determine a target return.

[0034] after iteratively updating parameters of the initial evaluation network based on a difference between the expected return and the target return, training the initial strategy model by using a second sample data until the training is stopped when a predetermined training condition is reached to obtain the preset strategy model.

[0035] In the foregoing solution, the method further comprises:

[0036] updating the initial target network by using model parameters in the initial evaluation network at intervals of a predetermined duration.

[0037] In the scheme, the determining of the target recommended music based on each music recommendation strategy comprises:

[0038] Determining, based on a weight parameter in the music recommendation strategy, a set of music information for the current object in a preset music information library;

[0039] Determining N music in the set of music information whose occurrence frequency meets a set condition, and determining the target recommended music according to the N music; wherein N is an integer greater than 0.

[0040] In the scheme, the collecting of the current feature corresponding to each current object in the preset area comprises:

[0041] Determining behavior data, position data, physiological image data, time data and music preference data corresponding to each current object in the preset area;

[0042] Performing data quantization processing on the behavior data, the position data, the physiological image data, the time data and the music preference data to determine the current feature.

[0043] Embodiments of the present application also provide a music playing recommendation device, comprising:

[0044] A determining unit configured to determine a current feature corresponding to each current object in a preset area; wherein the current feature comprises at least one of the following features for the current object: behavior feature, position feature, physiological state feature, time feature and music preference feature;

[0045] A strategy determining unit configured to input the current feature into a preset strategy model to determine a music recommendation strategy corresponding to the current object;

[0046] A music determining unit configured to determine a target recommended music based on the music recommendation strategy.

[0047] Embodiments of the present application also provide an electronic device comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the processor implements the steps in the above method when executing the computer program.

[0048] Embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps in the above method.

[0049] Embodiments of the present application also provide a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps in the above method.

[0050] In the embodiments of the present application, the current characteristics corresponding to each current object in the preset area are determined, wherein the current characteristics include at least one of the following characteristics of the current object: behavior characteristics, position characteristics, physiological state characteristics, time characteristics and music preference characteristics; the current characteristics are input into a preset strategy model to determine the music recommendation strategy corresponding to the current object; and the target recommended music is determined based on the music recommendation strategy. Since the user characteristics of multiple dimensions of multiple users are considered when recommending music, the personalized difference analysis of users is more comprehensive, so the music recommended by the scheme in the embodiments of the present application has high personalization, and since the target recommended music is determined according to the music recommendation of multiple users, the preferences of most people are considered, and thus the target music recommended should meet the public needs of multiple users. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0052] Figure 2 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0053] Figure 3 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0054] Figure 4 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0055] Figure 5 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0056] Figure 6 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0057] Figure 7 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0058] Figure 8 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0059] Figure 9 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided;

[0060] Figure 10 An optional flowchart of a music playing recommendation method provided by the embodiments of the present application is provided; DETAILED DESCRIPTION

[0061] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further described in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limitations to the present application. All other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0062] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0063] If similar descriptions of "first / second" appear in the application file, the following description is added. In the following description, the terms "first\second\third" referred to only distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0065] The embodiments of the present application provide a music playing recommendation method. Please refer to Figure 1 , an optional flowchart of the music playing recommendation method provided by the embodiments of the present application will be described in conjunction with Figure 1 illustrated steps:

[0066] S101, determine a current feature corresponding to each current object in a preset area; wherein the current feature includes at least one of the following features of the current object: behavior feature, position feature, physiological state feature, time feature and music preference feature.

[0067] In the embodiments of the present application, the music playing recommendation device can collect various data of each current object in the preset area for representing the current feature through the corresponding sensor, and determine the current feature by data quantization processing on the data. Wherein, the current feature includes at least one of the following features of the current object: behavior feature, position feature, physiological state feature, time feature and music preference feature.

[0068] In the embodiments of the present application, the preset area can be a closed area, for example, a shopping mall, a restaurant, or a stadium. The preset area can also be an unclosed area, for example, a square or a football field. The music playing recommendation device can be a server with corresponding data processing function set for the preset area. The current object can be a person (consumer) in the preset area. In other embodiments, the current object can also be other entities with similar characteristics.

[0069] In the embodiments of the present application, the recommendation device collects behavior data, position data, physiological image data, time data, and music preference data corresponding to each current object in the preset area. The behavior data, the position data, the physiological image data, the time data, and the music preference data are subjected to data quantization processing to determine the current characteristics.

[0070] For example, the recommendation device can obtain the behavior data of the current object in the shopping mall, including shopping frequency (F1), shopping time (F2), shopping path (F3), and product preference (F4). The shopping frequency (F1), the shopping time (F2), and the shopping path (F3) reflect the shopping habits and activity rules of the current object in the shopping mall. For example, the current object who frequently stays for a long time can need more diversified music to maintain freshness, and the current object on a specific path can need music matching the scene atmosphere in some areas. The product preference (F4) can indirectly reflect the aesthetic tendency and lifestyle of the current object, so as to infer the music style they can like. For example, the current object who prefers high-end fashion brands can prefer to listen to jazz or light music.

[0071] The recommendation device collects the position data of the current object in the shopping mall, such as the resident area (F5) and the area passenger flow (F6).

[0072] The recommendation device obtains the music preference data of the current object, including historical playing records (F7) and music interaction (F8) as music preference characteristics.

[0073] The recommendation device obtains the time data of the current object, including recording access time (F9) and seasonal characteristics (F10) as time characteristics.

[0074] The recommendation device obtains the physiological image data of the current object, which can capture physiological characteristics such as facial expression (F11), pupil contraction state (F12), body movement (F13), and verbal expression (F14) of the consumer by using Internet of Things sensors.

[0075] In the embodiments of the present application, the recommendation device can use one-hot encoding to convert all collected discrete data (F1, F4, F5, F7, F8, F9, F10) to obtain features, and continuous data (F2, F3, F6) is directly input into the model to obtain features. The Internet of Things device quantifies the features of F11-F14 through image recognition and biological signal analysis technology, and then combines the features corresponding to each data to determine the current features.

[0076] In S102, the current features are input into a preset strategy model to determine a music recommendation strategy corresponding to the current object.

[0077] In the embodiments of the present application, the recommendation device can input the current features into the trained preset strategy model to determine the music recommendation strategy corresponding to the current object.

[0078] The music recommendation strategy can include a weight parameter referred to when music is recommended for the current object.

[0079] The trained preset strategy model can be a Double Deep Q-Network (DQN) model, and in other embodiments, it can be other models with similar functions.

[0080] In the embodiments of the present application, in order to better adapt to different scenes and current object preferences, adjustable parameter systems are adopted for music recommendation decisions. The constituent elements of the music recommendation decision include but are not limited to music type preference weight θT, consumer interaction feedback weight θI, copyright compliance weight θL, environment adaptability weight θE, time sensitivity weight θTs, and physiological response weight θP. According to the preset current features t in the mall, the recommendation device will determine a suitable music recommendation decision A_t by dynamically adjusting these weight parameters, in order to accurately predict and evaluate the recommendation suitability of each music in the current situation. Therefore, in the embodiments of the present application, the music recommendation decision can be represented as A=[θT, θI, θL, θE, θTs, θP], each weight parameter corresponds to a different key factor affecting the public music recommendation decision, and together determines which music type, song or artist's work should be played by the recommendation device to optimize the consumer experience and comply with the requirements of copyright regulations.

[0081] The role of these weight parameters is to weight the importance of different factors when recommending music to determine which music or which type of music to ultimately recommend. Exemplary:

[0082] Music type preference weight (θT):

[0083] The weight (θT) reflects the current object's preference for music or music types. A higher θT indicates that the system places more importance on the consumer's preference for a particular music type. When the system selects music, it will prioritize those music types that align with the consumer's known preferences.

[0084] The consumer interaction feedback weight (θI):

[0085] The weight θI emphasizes the importance of the current object's interaction with music. If θI is high, then the system places more importance on music that can stimulate consumer interaction, such as likes, shares, and repeat plays. During the recommendation process, the system may prioritize music that has a higher frequency of past interaction.

[0086] The copyright compliance weight (θL):

[0087] The weight θL ensures that the recommended music complies with copyright regulations. A higher θL indicates that the system places more importance on copyright issues, avoiding the recommendation of music that does not have proper authorization. If copyright compliance is very important, then even if a song is very popular, the system will not recommend it if it does not meet copyright requirements.

[0088] The environmental adaptability weight (θE):

[0089] The weight θE takes into account environmental factors, such as the suitability of music for different areas of a store or time periods. A higher θE means that the system will place more importance on recommending music that matches the current environment. For example, recommending soft background music in a quiet reading area or energetic music in a sports equipment area.

[0090] The time sensitivity weight (θTs):

[0091] The weight θTs takes into account time factors, such as the time of day or seasonal changes. A higher θTs means that the system will place more importance on recommending music that is relevant to the current time. For example, playing energetic music in the morning or more relaxing music in the evening.

[0092] The physiological response weight (θP):

[0093] The weight θP takes into account the current object's physiological responses, such as assessing the consumer's emotional state through facial expressions, pupil dilation, and other indicators. A higher θP indicates that the system places more importance on these physiological signals. The system may recommend music based on the consumer's physiological responses, such as playing relaxing music when the consumer appears to be relaxed.

[0094] How to influence the recommendation process: In the recommendation process, the recommendation device will select a music recommendation strategy At according to the current characteristics St of the current object, that is, recommend one or more pieces of music. This selection is based on the combination of the above weight parameters, and the importance of each parameter is determined by its corresponding weight value. For example, if θT is high, the music type will be the main consideration in selecting music; if θI is high, the system may select those with a higher frequency of past interactions. These weights achieve dynamic optimization by adjusting the recommendation strategy, ensuring that the recommended music not only meets the personal preferences of the previous object, but also adapts to changes in the market environment, while ensuring copyright compliance. The system will adjust these weights based on the feedback (positive or negative) of the current object, making better choices in future recommendations.

[0095] S103, determining the target recommended music based on the music recommendation strategy.

[0096] In the embodiments of the present application, the recommendation device can determine a music set or a music type set corresponding to the current object in the preset music library for the weight parameters in the music recommendation strategy. The target recommended music is determined based on the music or music type with the highest frequency of occurrence in the multiple music sets or music type sets.

[0097] In the embodiments of the present application, the recommendation device can determine a music information set for the current object based on the weight parameters in the music recommendation strategy in the preset music information library. N pieces of music that meet the set condition in terms of frequency of occurrence are determined in the multiple music information sets, and the target recommended music is determined according to the N pieces of music. N is an integer greater than 0. The set condition can be that the frequency of occurrence ranks in the top N.

[0098] In the embodiments of the present application, the current characteristics corresponding to each current object in the preset area are determined, wherein the current characteristics include at least one of the following characteristics for the current object: behavior characteristics, location characteristics, physiological state characteristics, time characteristics, and music preference characteristics; the current characteristics are input into a preset strategy model to determine the music recommendation strategy corresponding to the current object; and the target recommended music is determined based on the music recommendation strategy. Since multiple dimensions of user characteristics of multiple users are considered when recommending music, the analysis of individual differences of users is more comprehensive, so the music recommended by the scheme in the embodiments of the present application has high personalization, and since the target recommended music is determined according to the music recommendations corresponding to multiple users, the preferences of most people are considered, and therefore the target music recommended should meet the public needs of multiple users.

[0099] Please refer to Figure 2 An optional flowchart of the music playing recommendation method provided by the embodiments of the present application will be described in combination with the steps:

[0100] S201, determine sample data corresponding to a plurality of historical objects in the preset area respectively; wherein each sample data comprises: historical characteristics, historical recommendation strategy and reward and punishment parameters of recommended music for the historical recommendation strategy of each historical object in adjacent time steps.

[0101] In the embodiment of the application, the recommendation device can collect and determine sample data corresponding to a plurality of historical objects in the preset area in adjacent time steps. Each sample data comprises: historical characteristics, historical recommendation strategy and reward and punishment parameters of recommended music for the historical recommendation strategy of each historical object in adjacent time steps.

[0102] The historical recommendation strategy comprises a weight parameter referred to when music is recommended; and the historical characteristics are used to reflect at least one of the following characteristics of the historical object in the corresponding time step: historical behavior characteristics, historical location characteristics, historical physiological state characteristics, historical time characteristics and historical music preference characteristics. The time step can be a predetermined time length, and the specific value of the time step is not limited in the embodiment of the application. The historical object can be a person appearing in the historical time in the preset area.

[0103] In the embodiment of the application, the recommendation device can construct a Markov decision process model (MDP) according to the historical characteristics, and convert the music recommendation problem into an MDP problem. Define historical characteristics S, S represents the degree of love of consumers for public music, wherein S contains all combinations of historical characteristics. Define historical recommendation strategy A, which can determine the type of music played for the historical object or the song ID based on the historical recommendation strategy, according to the music feature coding. Design reward and punishment parameters r, give positive feedback according to the positive response (such as love degree, interaction frequency) of the current object to the music and the copyright permission state, and give negative feedback if the copyright requirement is not met or the consumer response is negative.

[0104] S202, train an initial strategy model based on each sample data, stop training when a predetermined training condition is reached, and obtain the preset strategy model.

[0105] In the embodiment of the application, the recommendation strategy can train an initial strategy model based on each sample data, stop training when the initial strategy model reaches a predetermined training number, or the difference between the expected return and the target return output by the initial strategy model reaches a predetermined threshold, and obtain a preset strategy model.

[0106] In the embodiment of the present application, the initial strategy model uses a double network DQN algorithm, and initializes the memory unit, the evaluation network and the target network. In each iteration, the historical recommendation strategy At is selected according to the historical feature St, and the new historical feature St+1 and the reward and punishment parameter r corresponding to the historical object are collected. The value function Q(St, At) is updated as E[G(St, At)]. Wherein, G is the reward function, and E is the expectation function. The loss function L(θ) of the DQN is optimized by gradient descent to adjust the model parameters, so that the target Q function approximates the true Q function, and the training is stopped until the difference between the target Q function and the true Q function is less than a predetermined threshold, and the preset strategy model is obtained.

[0107] In the embodiment of the present application, sample data corresponding to a plurality of historical objects in the preset area is collected and determined; wherein each sample data includes: historical features, historical recommendation strategies and reward and punishment parameters of the music recommended by the historical recommendation strategies of each historical object in adjacent time steps. The historical features are used to reflect at least one of the following features of the historical object in the corresponding time step: historical behavior features, historical location features, historical physiological state features, historical time features and historical music preference features. The initial strategy model is trained based on each sample data, and the training is stopped until a predetermined training condition is reached, and the preset strategy model is obtained. Since the historical user features in multiple dimensions are considered during model training, the individual difference analysis of the historical user is more comprehensive, so the model trained by the scheme in the embodiment of the present application can consider more comprehensive diversity when determining the music recommendation strategy corresponding to the user, and the recommended music has higher individualization.

[0108] Please refer to Figure 3 An optional flowchart of the music playing recommendation method provided by the embodiment of the present application is shown in Figure 2 S201 in the embodiment of the present application can also be implemented by S301 to S304, which will be described in combination with the steps:

[0109] S301, determining the first historical features of each historical object in the preset area in the first time step.

[0110] In the embodiment of the present application, the recommendation device can collect various historical data of each historical object in the preset area in the first time step for reflecting the first historical features, and determine the first historical features by quantizing the various historical data.

[0111] Wherein, the first historical features are used to reflect at least one of the following features of the historical object in the corresponding first time step: historical behavior features, historical location features, historical physiological state features, historical time features and historical music preference features.

[0112] S302, determining the historical recommendation strategy based on the first historical feature by using a preset exploration method.

[0113] In the embodiments of the present application, the recommendation device can determine the corresponding historical recommendation strategy by using a preset exploration method for the first historical feature. The historical recommendation strategy includes a weight parameter referred to when music is recommended.

[0114] In the embodiments of the present application, the recommendation device can select a historical recommendation strategy according to the first historical feature by using an ε-greedy strategy. And based on the historical recommendation strategy, the corresponding music is selected for playing, and the recommendation device collects new historical features in the second time step corresponding to the historical object. In other embodiments, the recommendation device can determine the historical recommendation strategy by using other exploration methods.

[0115] The historical recommendation strategy includes at least one of the following weight parameters: music type preference weight, user interaction feedback weight, copyright compliance weight, environment adaptability weight, time sensitivity weight, and user physiological response weight.

[0116] S303, determining the reward and punishment parameters based on the music recommended based on the historical recommendation strategy.

[0117] In the embodiments of the present application, the recommendation device can determine the recommended music based on the historical recommendation strategy. And based on the copyright information of the recommended music and the feedback information of the historical object to the recommended music, the reward and punishment parameters are determined.

[0118] S304, determining the second historical feature of each historical object in the preset area in the second time step.

[0119] In the embodiments of the present application, the recommendation device can collect various historical data of each historical object in the preset area in the second time step for reflecting the second historical feature, and determine the second historical feature by performing data quantization processing on the various historical data.

[0120] Since the determined first historical feature and second historical feature both consider historical user features in multiple dimensions, the personalized difference analysis of historical users is more comprehensive, so the model trained by the first historical feature and the second historical feature can consider more comprehensive diversity when determining the music recommendation strategy corresponding to the user, and the recommended music has higher personalization.

[0121] Please refer to Figure 4 , an optional flowchart of the music playing recommendation method provided by the embodiments of the present application, Figure 3 S301 in the above-mentioned flowchart can also be implemented by S401 to S402, which will be described in conjunction with the steps:

[0122] S401, determine historical behavior data, historical position data, historical physiological image data, historical time data and historical music preference data of each historical object in the preset area within the first time step.

[0123] In the embodiment of the application, the recommendation device can collect historical behavior data, historical position data, historical physiological image data, historical time data and historical music preference data of each historical object in the preset area within the first time step.

[0124] The historical music preference data can be obtained through a third-party music software.

[0125] The historical position data is used to reflect the position of the historical object in the preset area and the object flow corresponding to the position within the first time step.

[0126] The historical physiological image data is used to reflect the physiological state change of the historical object within the first time step.

[0127] The historical time data is used to reflect the time and season when the historical object enters the preset area within the first time step.

[0128] The music preference data is used to reflect the music preference of the historical object in the third-party software.

[0129] S402, data quantization processing is performed on the historical behavior data, the historical position data, the historical physiological image data and the historical music preference data to determine the first historical feature.

[0130] In the embodiment of the application, the recommendation device can perform data quantization processing on the historical behavior data, the historical position data, the historical physiological image data and the historical music preference data respectively, determine the features corresponding to each data, and combine each feature to determine the first historical feature.

[0131] Since the determined first historical feature considers historical user features in multiple dimensions, the analysis of the individual differences of the historical user is more comprehensive, so the model trained through the first historical feature can consider more comprehensive diversity when determining the music recommendation strategy corresponding to the user, and the recommended music has higher individuality.

[0132] Please refer to Figure 5 , an optional flowchart of a music playing recommendation method provided by the embodiment of the application, Figure 3 S303 in the embodiment of the application can also be implemented through S501, which will be described in combination with the steps:

[0133] S501, determine the reward and punishment parameter based on the copyright information of the recommended music and the feedback information of the historical object to the recommended music; wherein the feedback information is used to reflect the preference degree of the historical object to the recommended music.

[0134] In the embodiment of the application, the recommendation device plays the recommended music after determining the recommended music, and collects the feedback information of the historical object to the recommended music in the process of playing the recommended music. The reward and punishment parameter is determined based on the feedback information and the copyright information of the recommended music.

[0135] The copyright information is used to represent whether the copyright of the recommended music is compliant. The feedback information is used to reflect the preference degree of the historical object to the recommended music.

[0136] The feedback information can include physiological, expression, action and language information feedback by the historical object to the recommended music.

[0137] In the embodiment of the application, the reward and punishment parameter is determined based on the copyright information of the recommended music and the feedback information of the historical object to the recommended music. Since the copyright information of the music and the feedback information of the object are considered in the process of determining the reward and punishment parameter, the determined reward and punishment parameter can accurately reflect the preference degree of the user to the recommended music.

[0138] Please refer to Figure 6 , an optional flowchart of the music playing recommendation method provided by the embodiment of the application, Figure 5 S501 in the embodiment of the application can also be implemented by S601 to S602, which will be described in combination with the steps:

[0139] S601, determine the feedback weight and preference degree information of the historical object to the recommended music based on the behavior feedback information and the physiological feedback information.

[0140] In the embodiment of the application, the feedback information includes behavior feedback information and physiological feedback information. The recommendation device can collect the behavior feedback information and the physiological feedback information of the user to the recommended music through the sensor in the process of playing the recommended music. The behavior feedback information and the physiological feedback information are quantitatively processed to determine the feedback weight and the preference degree information.

[0141] Among them, the behavior feedback information includes whether the user actively selects to repeat playing the recommended music, whether to listen completely, whether to skip the song, whether to like or share the song, and other interactive behavior information. These behaviors reflect the user's direct preference degree for the song. The physiological feedback information includes the data collected through the Internet of Things sensors, such as facial expression recognition (smiling or frowning), pupil dilation (excited or relaxed state), body movement (whether to sway with the rhythm of the music), and verbal expression (whether to hum or comment), and other non-verbal reactions can indirectly reflect the user's emotional state and the user's inner feelings for the music.

[0142] S602, determine the reward and punishment parameters based on the feedback weight, the preference degree information and the copyright reward coefficient; wherein the size of the copyright reward coefficient is related to the compliance represented by the copyright information.

[0143] In the embodiments of the application, the recommendation device can add the product of the feedback weight and the preference degree information to the copyright reward coefficient to determine the reward and punishment parameters. The size of the copyright reward coefficient is related to the compliance represented by the copyright information.

[0144] In the embodiments of the application, the reward parameter r is used to evaluate the reward size obtained by the system when moving from the first historical feature St to the second historical feature St+1 after playing the recommended music. The specific design is as follows: the historical object feedback and the copyright compliance are taken as function evaluation indicators, that is, when the recommended music is liked by the consumer, actively interacted, and the copyright compliance is ensured, the system will obtain a positive reward, and the feedback weight is equal to the product of the Positive Feedback Weight (denoted as Wp) and the Feedback Score (denoted as F, the preference degree information), plus a copyright compliance reward coefficient Copy rightCompliance Weight (denoted as Wc). Among them, the reward coefficient when the copyright is compliant is greater than the reward coefficient when the copyright is not compliant. That is, the reward parameter r(St,At)=Wp×F+Wc.

[0145] Among them, "Feedback Score" (F) refers to an index used to quantify the positive or negative reaction of consumers to the played music in the music recommendation system. This score is usually calculated based on a variety of consumer behaviors and physiological indicators. In short, the Feedback Score is an instant feedback mechanism in the system optimization of the music recommendation process, which helps the system to learn and understand the actual acceptance of the historical object to the music. The Positive Feedback Weight is a parameter used to quantify and adjust the influence of positive feedback of consumers in the music recommendation algorithm, and by adjusting it, the recommendation strategy can be optimized to be closer to the actual preferences of users, thereby improving the overall performance of the recommendation system and user satisfaction.

[0146] wherein Wp, F, Wc can be defined according to actual training requirements. Conversely, if the recommended music fails to meet the copyright requirements or the consumer reacts negatively to it, the system will suffer a negative penalty, the value of which is equal to the Penalty Value (denoted as Pv): reward parameter r(St, At) = -pv, the reward function aims to guide the system to consider not only the user's preferences and interaction conditions when making music recommendations, but also to ensure that the recommended music is copyright compliant, thereby optimizing the overall recommendation strategy through reinforcement learning.

[0147] In the embodiments of the present application, the feedback weight and the preference degree information of the historical object for the recommended music are determined based on the behavior feedback information and the physiological feedback information. The reward and punishment parameters are determined based on the feedback weight, the preference degree information and the copyright reward coefficient; wherein the size of the copyright reward coefficient is related to the compliance represented by the copyright information. Since the copyright information of the music and the feedback information of the object are considered in the process of determining the reward and punishment parameters, the dimensions considered are more comprehensive, so the reward and punishment parameters determined can accurately reflect the preference degree of the user for the recommended music. Further, through more accurate reward and punishment parameters, a pre-set strategy model with better performance can also be trained.

[0148] Please refer to Figure 7 , an optional flowchart of the music playing recommendation method provided by the embodiments of the present application, Figure 3 S201 in the above-mentioned S201 can also be realized by S701 to S703, which will be described in conjunction with the steps:

[0149] S701, store each sample data in the initial memory network.

[0150] In the embodiments of the present application, the recommendation device can initialize the memory network, the initial evaluation network and the initial target network. The recommendation device stores each sample data in the initial memory network.

[0151] In the embodiments of the present application, a double network DQN algorithm can be used: first, initialize the memory unit (ReplayMemory) to store sample data (first historical features, historical recommendation strategy, reward and punishment parameters, second historical features), and initialize the initial evaluation network (Current Q Network) and the initial target network (Target Q Network) at the same time, and the weight parameters of the two are denoted as θ and θ'.

[0152] 1. Initializing Replay Memory: Purpose: Replay memory stores historical experiences, including a four-tuple of state, action, reward, and new state. This is crucial for training the DQN model. Operation: Create a sufficiently large data structure (such as a queue or list) to store experience samples. These samples will be randomly sampled during training to update the model's weights. Logic: Initialize an empty replay memory. Typically, a maximum capacity limit is set; once this limit is exceeded, new experiences will overwrite the oldest experiences.

[0153] 2. Initializing the Current Q Network: Purpose: The current Q network is used to estimate the value of taking an action in the current state. Operation: Construct a neural network model whose input layer receives information from the state space, and whose output layer predicts the value of each possible action. Logic: Initialize the network's weight parameters, which will be updated during training to approximate the true Q-value.

[0154] 3. Initializing the Target Q Network: Purpose: The target network is used to stabilize the training process and reduce fluctuations during value network updates. Operation: Construct a neural network model with the same structure as the value network. Logic: Initialize the weight parameters of the target network to be the same as those of the value network. During training, the weight parameters of the target network are periodically updated to match the weight parameters of the value network, but the update frequency is low to maintain stability.

[0155] 4. Initialize the memory unit, valuation network, and target network. Operation: Before the algorithm begins, initialize these three components. Logic: Memory Unit: Create an empty data structure and set its maximum capacity. Valuation Network: Build a neural network and initialize its weight parameters. Target Network: Build a neural network with the same structure as the valuation network and initialize its weight parameters to be the same as the valuation network.

[0156] S702. Input the first historical feature in the first sample data, the reward / penalty parameter, and the first historical recommendation strategy into the initial valuation network to determine the expected return, and input the second historical feature in the first sample data into the initial target network to determine the target return.

[0157] In this embodiment of the application, the first historical feature in the first sample data, the reward and punishment parameters, and the first historical recommendation strategy are input into the initial valuation network to determine the expected return, and the second historical feature in the first sample data is input into the initial target network to determine the target return.

[0158] S703, after iteratively updating the parameters of the initial evaluation network based on the difference between the expected return and the target return, training the initial strategy model using a second sample data until a predetermined training condition is reached to stop training, obtaining the preset strategy model.

[0159] In the embodiments of the present application, after iteratively updating the parameters of the initial evaluation network based on the difference between the expected return and the target return, the initial strategy model is trained using a second sample data until a predetermined training condition is reached to stop training, obtaining the preset strategy model.

[0160] In the embodiments of the present application, the first historical feature, the historical recommendation strategy, the reward and punishment parameter and the second historical feature (St, At, r, St+1) are stored into the initial memory network. St, At and r are input into the initial evaluation network to calculate the value function Q(St, At), that is, the expected return of performing action At in state St, which is expressed as: Q(St, At) = E[G(St, At)]

[0161] Wherein, G(St, At) is the discount return, which can be expressed by Bellman equation as:

[0162] G(St, At) = r + γ*max_{At+1}Q(St+1, At+1; θ); γ is the discount factor, which reflects the degree of discounting future rewards.

[0163] A loss function L(θ) is constructed for the initial strategy model, which is used to measure the gap between the target Q function (from the initial target network) and the current Q function (from the initial evaluation network), and its expression is:

[0164] L(θ) = E[(TargetQ(St, At; θ') - Q(St, At; θ))^2]

[0165] The loss function L(θ) is optimized by gradient descent method, and the weight parameter θ of the initial evaluation network is updated, so that the evaluation network is more accurate to approach the actual Q value, thereby gradually improving the preset strategy model.

[0166] In some other embodiments, the recommendation device can update the initial target network using the model parameters in the initial evaluation network at intervals of a predetermined time length.

[0167] In the embodiments of the present application, the deep reinforcement learning DQN algorithm model is optimized. The historical features and the historical recommendation strategy are defined, the reward parameter is designed, the positive and negative feedbacks are given based on the real-time feedback of the consumers and the copyright compliance, the dynamic optimization of the music recommendation strategy is realized, and the personalized music can be accurately recommended in the face of different objects and different scenes.

[0168] Please see Figure 8 The following is an optional flowchart illustrating the music recommendation method provided in this application embodiment, which will be described in conjunction with the steps:

[0169] S11, Data Collection.

[0170] In this embodiment of the application, consumer consumption behavior data, location data, music preference data, time data, physiological data, etc. are obtained.

[0171] S12, Feature processing.

[0172] In this embodiment, all discrete data are converted using one-hot encoding, while continuous data is directly input into the model. IoT devices quantify physiological data using image recognition and biosignal analysis techniques to obtain multiple features.

[0173] S13. Strategy model construction: transforming the music recommendation problem into an MDP problem.

[0174] In this embodiment, a Markov Decision Process (MDP) model is constructed based on the above features, transforming the music recommendation problem into an MDP problem. A feature S is defined, representing the consumer's preference for publicly broadcast music, where S contains the combined states of all features, S = (F1, F2, ..., F14). A strategy A is defined, representing the weight parameters for recommending music. A reward function r is designed, providing positive feedback based on the consumer's positive response to the music and the copyright licensing status; negative feedback is given if copyright requirements are not met or the consumer's response is negative. A dual-network DQN algorithm is used to initialize the memory unit, the evaluation network, and the target network.

[0175] S14. Deep reinforcement learning strategy to optimize recommendation strategy.

[0176] In this embodiment, the dual-network DQN algorithm is trained until a preset training condition is met to obtain a trained DQN model.

[0177] S15. Generate a recommendation list.

[0178] In this embodiment, based on a trained DQN model and the data of objects within the current shopping mall, the optimal music recommendation strategy is determined. The model outputs the optimal music sequence, which corresponds to the music list most likely to be accepted by consumers.

[0179] Based on the trained DQN model, the optimal music recommendation strategy is determined for the current shopping mall environment. After training, the system can generate the optimal music recommendation strategy in real time based on the mall's current environmental conditions (such as real-time consumer behavior, time, season, etc.). This strategy is manifested as a series of actions (i.e., music recommendation decisions), each aimed at maximizing consumer satisfaction and copyright compliance. The recommendation strategy is not a static playlist, but a dynamic adjustment process that continuously optimizes music selection based on the mall environment and consumer feedback. Therefore, the optimal strategy is reflected in the model's ability to select the most suitable music genre or song at any given moment to adapt to current consumer needs and the mall atmosphere.

[0180] The model outputs the optimal action sequence, which corresponds to a music list most likely to be accepted by consumers and is copyright compliant. The optimal action sequence is the system's music recommendation decision sequence under a specific state; each action corresponds to a specific music recommendation, such as playing a certain type of music or the works of a certain artist. This sequence is generated by the model through multiple iterations and learning, based on feedback from the reward parameters, by continuously optimizing the value function Q(St,At). The selection of each action comprehensively considers factors such as consumer music preferences, interaction, copyright compliance, and environmental adaptability to ensure that each piece of music in the sequence is highly personalized, timely, and copyright compliant. The optimal action sequence is ultimately presented to the mall operator in the form of a music playlist, reflecting the set of music most likely to be accepted by consumers and whose copyright is not compromised under the current environmental conditions.

[0181] In this embodiment, a multi-feature fusion-based consumer behavior analysis and prediction model is employed. The key is the construction of a model that comprehensively considers multiple features, including consumer shopping frequency, shopping duration, shopping path, product preferences, frequented areas, regional foot traffic, historical playback records, music interaction, access time, and seasonal characteristics. Machine learning algorithms are used to fuse these features to predict consumers' music preferences. Furthermore, a Markov decision process model and deep reinforcement learning algorithms are used: the music recommendation problem is constructed as a Markov decision process and optimized using the Deep Reinforcement Learning (DQN) algorithm. Features and recommendation strategies are defined, reward parameters are designed, and positive and negative feedback are provided based on real-time consumer feedback and copyright compliance, achieving dynamic optimization of the music recommendation strategy.

[0182] Please see Figure 9 This is a schematic diagram of the music recommendation device provided in the embodiments of this application.

[0183] This application embodiment also provides a music recommendation device 800, including:

[0184] The determining unit 801 is used to determine the current features corresponding to each current object within a preset area; wherein, the current features include at least one of the following features for the current object: behavioral features, location features, physiological state features, time features, and music preference features;

[0185] The strategy determination unit 802 is used to input the current features into a preset strategy model to determine the music recommendation strategy corresponding to the current object;

[0186] The music determination unit 803 is used to determine the target recommended music based on the music recommendation strategy.

[0187] In this embodiment of the application, the determining unit 801 in the music recommendation device 800 is used to determine sample data corresponding to multiple historical objects in the preset area; wherein, each sample data includes: historical features of each historical object in adjacent time steps, historical recommendation strategy, and reward and punishment parameters for the music recommended by the historical recommendation strategy;

[0188] The historical recommendation strategy includes weight parameters referenced when making music recommendations; the historical features are used to reflect at least one of the following features of the historical object within the corresponding time step: historical behavioral features, historical location features, historical physiological state features, historical time features, and historical music preference features.

[0189] The initial policy model is trained based on each sample data until the predetermined training conditions are met, at which point training stops, and the preset policy model is obtained.

[0190] In this embodiment of the application, the determining unit 801 in the music recommendation device 800 is used to determine the first historical feature of each historical object in the preset area within a first time step.

[0191] The historical recommendation strategy is determined based on the first historical feature using a preset exploration method;

[0192] The reward and penalty parameters are determined for the music recommended based on the historical recommendation strategy.

[0193] Determine the second historical feature of each historical object within the preset area within a second time step.

[0194] In this embodiment of the application, the music recommendation device 800 determines the historical behavior data, historical location data, historical physiological image data, historical time data, and historical music preference data of each historical object within the preset area within the first time step;

[0195] The historical behavioral data, historical location data, historical physiological image data, and historical music preference data are subjected to data quantification processing to determine the first historical feature.

[0196] In this embodiment of the application, the historical behavior data is used to reflect the transaction behavior of the historical object within the first time step;

[0197] The historical location data is used to reflect the location of the historical object within the preset area and the object traffic corresponding to the location within the first time step;

[0198] The historical physiological image data is used to reflect the changes in the physiological state of the historical object within the first time step;

[0199] The historical time data is used to reflect the time and season when the historical object entered the preset area within the first time step;

[0200] The music preference data is used to reflect the music preferences of the historical object in third-party software.

[0201] In this embodiment of the application, the music recommendation device 800 determines the reward and punishment parameters based on the copyright information of the recommended music and the feedback information of the historical object regarding the recommended music; wherein, the feedback information is used to reflect the degree of preference of the historical object for the recommended music.

[0202] In this embodiment of the application, the feedback information includes: behavioral feedback information and physiological feedback information; the music recommendation device 800 determines the feedback weight and preference information of the historical object for the recommended music based on the behavioral feedback information and the physiological feedback information;

[0203] The reward and punishment parameters are determined based on the feedback weight, the preference information, and the copyright reward coefficient; wherein the magnitude of the copyright reward coefficient is related to the compliance represented by the copyright information.

[0204] In this embodiment of the application, the historical recommendation strategy includes at least one of the following weight parameters: music genre preference weight, user interaction feedback weight, copyright compliance weight, environmental adaptability weight, time sensitivity weight, and user physiological response weight.

[0205] In this embodiment of the application, the initial strategy model includes: an initial memory network, an initial evaluation network, and an initial target network; the music recommendation device 800 is used to store each of the sample data in the initial memory network;

[0206] The first historical feature from the first sample data, the reward / penalty parameter, and the first historical recommendation strategy are input into the initial valuation network to determine the expected return, and the second historical feature from the first sample data is input into the initial target network to determine the target return;

[0207] After iteratively updating the parameters of the initial valuation network based on the difference between the expected return and the target return, the initial policy model is then trained using the second sample data until the predetermined training conditions are met, at which point training stops, and the preset policy model is obtained.

[0208] In this embodiment of the application, the music recommendation device 800 is used to update the initial target network using the model parameters in the initial valuation network at predetermined intervals.

[0209] In this embodiment of the application, the music determination unit 803 in the music recommendation device 800 is used to determine the music information set for the current object in a preset music information database based on the weight parameters in the music recommendation strategy;

[0210] N music tracks whose frequency of occurrence meets a set condition are identified from multiple music information sets, and the target recommended music is determined based on the N music tracks; where N is an integer greater than 0.

[0211] In this embodiment of the application, the determining unit 803 in the music recommendation device 800 is used to collect and determine the behavioral data, location data, physiological image data, time data and music preference data corresponding to each current object in the preset area;

[0212] The behavioral data, location data, physiological image data, time data, and music preference data are subjected to data quantization processing to determine the current feature.

[0213] It should be noted that, in the embodiments of this application, if the above-described item information processing method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an item information processing device (which may be a personal computer, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0214] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the method of the music recommendation device.

[0215] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0216] It should be noted that, Figure 10 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 10 As shown, this application embodiment provides an electronic device 900, including a memory 902 and a processor 901. The memory 902 stores a computer program that can run on the processor 901. When the processor 901 executes the program, it implements the steps in the above-described method, wherein;

[0217] Processor 901 typically controls the overall operation of electronic device 900.

[0218] The memory 902 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) in the processor 901 and various modules in the electronic device 900. It can be implemented by flash memory or random access memory (RAM).

[0219] Correspondingly, this application embodiment also provides a computer program product, including a computer program that can be executed by the processor 901 of the electronic device 900 to complete the steps in the method of the music recommendation device 800.

[0220] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0221] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0222] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the apparatus or units can be electrical, mechanical, or other forms.

[0223] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0224] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0225] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0226] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0227] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of playing music recommendation, characterized by, The method comprises: determining a current feature corresponding to each current object in a preset area; wherein the current feature comprises at least one of the following features of the current object: behavior feature, position feature, physiological state feature, time feature, and music preference feature; inputting the current feature into a preset strategy model to determine a music recommendation strategy corresponding to the current object; determining target recommended music based on the music recommendation strategy.

2. The music recommendation method of claim 1, wherein, Before the current feature is input into the preset strategy model to determine the music recommendation strategy corresponding to the current object, the method further comprises: determining sample data corresponding to a plurality of historical objects in the preset area respectively; wherein each sample data comprises: historical features of each historical object in adjacent time steps, a historical recommendation strategy, and reward and punishment parameters of recommended music for the historical recommendation strategy; wherein the historical recommendation strategy comprises a weight parameter referred to when music is recommended; the historical features are used to reflect at least one of the following features of the historical object in the corresponding time step: historical behavior feature, historical position feature, historical physiological state feature, historical time feature, and historical music preference feature; training an initial strategy model based on each sample data until a predetermined training condition is met to stop training, thereby obtaining the preset strategy model.

3. The method of claim 2, wherein, The historical features comprise first historical features and second historical features; and the determination of the sample data corresponding to the plurality of historical objects in the preset area respectively comprises: determining the first historical features of each historical object in the preset area in a first time step; determining the historical recommendation strategy based on the first historical features using a preset exploration method; determining the reward and punishment parameters for the recommended music based on the historical recommendation strategy; determining the second historical features of each historical object in the preset area in a second time step.

4. The method of claim 3, wherein, The determination of the first historical features of each historical object in the preset area in the first time step comprises: determining historical behavior data, historical position data, historical physiological image data, historical time data, and historical music preference data of each historical object in the preset area in the first time step; performing data quantization processing on the historical behavior data, the historical position data, the historical physiological image data, and the historical music preference data to determine the first historical features.

5. The method of claim 4, wherein, The historical behavior data are used to reflect transaction behavior of the historical object in the first time step; The historical position data are used to reflect the position of the historical object in the preset area and the object flow corresponding to the position in the first time step; The historical physiological image data are used to reflect physiological state changes of the historical object in the first time step; The historical time data are used to reflect the time and season when the historical object enters the preset area in the first time step; The music preference data are used to reflect music preference of the historical object in a third-party software.

6. The method of claim 4, wherein, The determination of the reward and punishment parameters for the recommended music based on the historical recommendation strategy comprises: determine the reward and punishment parameter based on the copyright information of the recommended music and the feedback information of the historical object for the recommended music.

7. The method of claim 6, wherein the music recommendation is played. The feedback information includes behavior feedback information and physiological feedback information. Determining the reward and punishment parameter based on the copyright information of the recommended music and the feedback information of the historical object for the recommended music includes: determining the feedback weight and preference degree information of the historical object for the recommended music based on the behavior feedback information and the physiological feedback information; determining the reward and punishment parameter based on the feedback weight, the preference degree information and the copyright reward coefficient, wherein the size of the copyright reward coefficient is related to the compliance represented by the copyright information.

8. The method of claim 3, wherein the music is played in response to a user's request. The historical recommendation strategy includes at least one of the following weight parameters: music type preference weight, user interaction feedback weight, copyright compliance weight, environment adaptability weight, time sensitivity weight and user physiological response weight.

9. The method of claim 3 to 8, wherein, The initial strategy model includes an initial memory network, an initial evaluation network and an initial target network. Training the initial strategy model based on each sample data until the predetermined training condition is met to stop training to obtain the preset strategy model includes: storing each sample data in the initial memory network; inputting the first historical feature, the reward and punishment parameter and the first historical recommendation strategy in the first sample data into the initial evaluation network to determine the expected return, and inputting the second historical feature in the first sample data into the initial target network to determine the target return; iteratively updating the parameters of the initial evaluation network based on the difference between the expected return and the target return, and then training the initial strategy model using the second sample data until the predetermined training condition is met to stop training to obtain the preset strategy model.

10. The method of claim 9, wherein, The method further includes: updating the initial target network using the model parameters in the initial evaluation network at intervals of a predetermined duration.

11. The method of claim 3 to 8, wherein, Determining the target recommended music based on the music recommendation strategy includes: determining a set of music information for the current object in the preset music information library based on the weight parameters in the music recommendation strategy; determining N music whose occurrence frequency meets the set condition in a plurality of music information sets, and determining the target recommended music according to the N music; wherein N is an integer greater than 0.

12. The method of claim 3 to 8, wherein, Determining the current feature corresponding to each current object in the preset area includes: determining the behavior data, location data, physiological image data, time data and music preference data corresponding to each current object in the preset area; quantitatively processing the behavior data, location data, physiological image data, time data and music preference data to determine the current feature.

13. A music recommendation device for playing music, characterized by includes: A determining unit is configured to determine a current feature corresponding to each current object in a preset area, wherein the current feature comprises at least one of the following features of the current object: a behavior feature, a position feature, a physiological state feature, a time feature, and a music preference feature. A strategy determining unit is configured to input the current feature into a preset strategy model to determine a music recommendation strategy corresponding to the current object. A music determining unit is configured to determine target recommended music based on the music recommendation strategy.

14. An electronic device, comprising: A computer program product, comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the processor implements the steps in the method according to any one of claims 1 to 12 when running the computer program.

15. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps in the method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps in the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Music recommendation method, device and equipment based on wearable equipment and storage medium

    CN115795085A

  • Music recommendation method and system and intelligent cabin

    CN117216315A

  • Music recommendation for influencing physiological state

    US20220233807A1