A music recommendation method and device

By acquiring and integrating the multimodal features of the target user, generating the target audio sequence and determining the music content based on its characteristics, the problem of inaccurate music recommendation in the prior art is solved, and more efficient music recommendation is achieved.

CN118916512BActive Publication Date: 2025-06-20AVATR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410999685.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-06-20
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

The existing car music recommendation system is difficult to achieve accurate matching and recall of music content, resulting in poor recommendation results.

Method used

By obtaining the multimodal features of the target user, performing feature fusion, generating a target audio sequence, determining the music content based on the target audio features, and forming a music recommendation list.

Benefits of technology

Effectively make up for the differences in expression between multimodal characteristics and music content, improve the accuracy and comprehensiveness of music recommendations, and thus improve the recommendation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118916512B_ABST
    Figure CN118916512B_ABST
Patent Text Reader

Abstract

The present application discloses a music recommendation method and apparatus. The method includes: obtaining at least one type of modal feature for music recommendation for a target user; each type of modal feature in the at least one type of modal feature is obtained by performing feature extraction on a type of modal data; performing feature fusion on the at least one type of modal feature to obtain a fused modal feature; generating a target audio sequence for characterizing the music recommendation requirement of the target user based on the fused modal feature; determining at least one target music content at least based on the target audio feature of the target audio sequence to form a music recommendation list; and outputting the music recommendation list. The solution of the present application can effectively make up for the differences in the expression modes between multi-modal features and music content, improve the accuracy and comprehensiveness of music recommendation, and further improve the music recommendation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and relates to, but is not limited to, a music recommendation method and apparatus. Background Art

[0002] With the continuous improvement of the intelligent level of automobiles, the vast majority of automobiles are equipped with in-vehicle music recommendation systems. The in-vehicle music recommendation system aims to provide users with personalized music content to meet their music listening needs during travel. Currently, the in-vehicle music recommendation system is still in an immature stage and it is difficult to provide users with accurate and satisfactory music content.

[0003] In related technologies, when making music recommendations, a music recall strategy based on text features such as keywords, tags, and themes is usually adopted to match and recall music content. However, due to the objective differences in the expression methods between text and music, the above solutions cannot achieve accurate matching and recall of music content, resulting in poor music recommendation effects. Summary of the Invention

[0004] This application provides a music recommendation method, apparatus, device, storage medium, and computer program product, which can effectively make up for the differences in the expression methods between multi-modal features and music content, improve the accuracy and comprehensiveness of music recommendations, and thus improve the music recommendation effect.

[0005] The technical solution of this application is implemented as follows:

[0006] In a first aspect, this application provides a music recommendation method, and the music recommendation method includes:

[0007] Obtain at least one type of modal feature for music recommendation for a target user; each type of modal feature in the at least one type of modal feature is obtained by extracting features from a type of modal data;

[0008] Perform feature fusion on the at least one type of modal feature to obtain a fused modal feature;

[0009] Generate a target audio sequence for characterizing the music recommendation needs of the target user based on the fused modal feature;

[0010] Determine at least one target music content based at least on the target audio features of the target audio sequence to form a music recommendation list;

[0011] Output the music recommendation list.

[0012] In a second aspect, this application provides a music recommendation apparatus, and the apparatus includes:

[0013] An acquisition unit, configured to acquire at least one type of modal feature for music recommendation for a target user; each type of the at least one type of modal feature is obtained by performing feature extraction on a type of modal data;

[0014] A fusion unit, configured to perform feature fusion on the at least one type of modal feature to obtain a fused modal feature;

[0015] A generation unit, configured to generate a target audio sequence for characterizing the music recommendation requirement of the target user based on the fused modal feature;

[0016] A determination unit, configured to determine at least one target music content based at least on the target audio feature of the target audio sequence to form a music recommendation list;

[0017] An output unit, configured to output the music recommendation list.

[0018] In a third aspect, the present application further provides an electronic device, including at least a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, the above music recommendation method is implemented.

[0019] In a fourth aspect, the present application further provides a storage medium, where a computer program is stored on the storage medium, and when the computer program on the storage medium is executed, the above music recommendation method is implemented.

[0020] In a fifth aspect, the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the above music recommendation method is implemented.

[0021] The present application provides a music recommendation method, apparatus, device, storage medium, and computer program product. The music recommendation method includes: acquiring at least one type of modal feature for music recommendation for a target user; each type of the at least one type of modal feature is obtained by performing feature extraction on a type of modal data; performing feature fusion on the at least one type of modal feature to obtain a fused modal feature; generating a target audio sequence for characterizing the music recommendation requirement of the target user based on the fused modal feature; determining at least one target music content based at least on the target audio feature of the target audio sequence to form a music recommendation list; and outputting the music recommendation list.

[0022] In the solution of the present application, by generating a target audio sequence for characterizing the user's music recommendation requirements, it is possible to determine at least one target music content based on the target audio features of the target audio sequence, and then form and output a music recommendation list based on the at least one target music content. It can be seen that the solution of the present application fully considers that audio features pay more attention to the expression in aspects such as rhythm, melody, harmony, and timbre. With the help of audio features, it can be accurately mapped to music content, and music content can also be completely and accurately characterized by audio features. Therefore, by generating a target audio sequence for characterizing the user's music recommendation requirements and determining at least one target music content based on the target audio features of the target audio sequence, accurate matching and recall of music content can be achieved. In this way, the difference in the expression mode between multi-modal features and music content can be effectively compensated, the accuracy and comprehensiveness of music recommendation can be improved, and thus the music recommendation effect can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 FIG. 6 is a first alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0024] Figure 2 FIG. 7 is a second alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0025] Figure 3 FIG. 8 is a third alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0026] Figure 4 FIG. 9 is a fourth alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0027] Figure 5 FIG. 10 is a fifth alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0028] Figure 6 FIG. 11 is a sixth alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0029] Figure 7 FIG. 12 is a seventh alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0030] Figure 8 FIG. 13 is an eighth alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0031] Figure 9 FIG. 14 is a ninth alternative flowchart of the music recommendation method provided by an embodiment of the present application;

[0032] Figure 10An optional schematic diagram of the overall business module of the music recommendation method provided by the embodiments of this application;

[0033] Figure 11 An optional schematic diagram of the system architecture of the music recommendation method provided by the embodiments of this application;

[0034] Figure 12 An optional schematic diagram of the recall content balancing scheme of the music recommendation method provided by the embodiments of this application;

[0035] Figure 13 An optional schematic diagram of the time penalty function curve of the music recommendation method provided by the embodiments of this application;

[0036] Figure 14 An optional schematic diagram of the large model algorithm flowchart of the music recommendation method provided by the embodiments of this application;

[0037] Figure 15 An optional structural schematic diagram of the music recommendation device provided by the embodiments of this application;

[0038] Figure 16 An optional structural schematic diagram of the electronic device provided by the embodiments of this application. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will further describe the specific technical solutions of the application in detail with reference to the accompanying drawings in the embodiments of this application. The following embodiments are used to illustrate this application, but are not used to limit the scope of this application.

[0040] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0041] In the following description, the terms "first / second / third" are only used to distinguish different objects, and do not represent a specific order for the objects, and do not have a limitation on the order of precedence. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of this application described here can be implemented in an order other than the order illustrated or described here.

[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0043] The embodiments of the present application provide a music recommendation method, apparatus, device, storage medium, and computer program product. In practical applications, the music recommendation method can be implemented by a music recommendation apparatus, and each functional entity in the music recommendation apparatus can be jointly implemented by the hardware resources of an electronic device, such as computing resources (e.g., a processor) and communication resources (e.g., for supporting various communication methods such as optical cables and cellular).

[0044] Next, the embodiments of the music recommendation method, apparatus, device, storage medium, and computer program product provided by the present application will be described.

[0045] In a first aspect, the embodiments of the present application provide a music recommendation method. The functions implemented by this method can be achieved by a processor in an electronic device (e.g., a terminal device or a controller) calling program code. Of course, the program code can be stored in a computer storage medium. It can be seen that this electronic device includes at least a processor and a storage medium.

[0046] Next, taking the electronic device as the execution subject as an example, the music recommendation method provided by the embodiments of the present application will be described.

[0047] Figure 1 is a schematic flowchart of the music recommendation method according to the embodiments of the present application. As Figure 1 shown, this process may include but is not limited to the following S101 to S105.

[0048] S101. The electronic device obtains at least one type of modal feature for music recommendation for the target user.

[0049] Each type of modal feature in the at least one type of modal feature is obtained by performing feature extraction on a type of modal data.

[0050] The target user can be any user with a music recommendation requirement. The embodiments of the present application do not limit the type and number of target users, which can be set according to actual situations.

[0051] Modal data refers to the data used for music recommendation for the target user and can be any data related to the target user. The embodiments of the present application do not limit the type and number of modal data, which can be set according to actual situations. Exemplarily, the modal data can be of text type, audio type, image type, video type, etc.

[0052] In a possible implementation, the modal data at least includes user data, environmental data, vehicle data, etc. Among them: User data can be any parameter related to the target user himself. For example, the basic data of the user (such as age, gender, education level, occupation, driving style, riding duration, etc.), the behavior data of the user (such as actions, expressions, voices, activity levels, interaction methods, etc.), the music preference data of the user (such as listening habits, music preferences, playlists, collection preferences, etc.), the historical music data of the user (such as historical listening records, historical music evaluation and feedback records, historical music sharing records, historical music search records, historical music browsing records, etc.); Environmental data can be any data related to the current environment where the target user is located. For example, time, weather, temperature, humidity, light, air quality, ambient sound, location, road type, number of pedestrians, traffic congestion, other vehicle behaviors, etc.; Vehicle data can be any data related to the vehicle the target user is riding in. For example, the driving data of the vehicle (such as speed, location, battery level, mileage, gear, driving mode, etc.), the status data of the vehicle (such as the status of each component, ambient light, whether it enters a geofence, etc.), the configuration data of the vehicle (such as vehicle model, vehicle layout, drive method, etc.).

[0053] In practice, the modal data can also include other data related to the target user, and these data can all be used as the basis for music recommendation for the target user to achieve music recommendation for the target user.

[0054] Modal features refer to the features obtained after feature extraction of modal data. The types and quantities of modal features in the embodiments of the present application are not limited and can be set according to actual situations. It can be understood that modal features correspond to modal data.

[0055] Exemplarily, at least one type of modal data for music recommendation for the target user can be collected or gathered in advance through specific sensors and specific application programming interfaces and stored in an electronic device. Correspondingly, S101 can be implemented as: the electronic device obtains at least one type of modal data from local or the cloud; traverses at least one type of modal data; for each type of modal data, uses a feature extraction algorithm or a feature extraction model to perform feature extraction on this type of modal data to obtain a type of modal feature; after the traversal is completed, at least one type of modal feature is obtained.

[0056] The embodiments of the present application do not limit the specific content or type of the feature extraction algorithm or the feature extraction model, and can be set according to actual situations.

[0057] S102. The electronic device performs feature fusion on the at least one type of modal features to obtain fused modal features.

[0058] The fused modality features refer to the features obtained after fusing at least one type of modality features. The embodiments of the present application do not limit the specific content of the fused modality features, which can be set according to actual situations.

[0059] Exemplarily, S102 can be implemented as follows: The electronic device uses a feature fusion algorithm or a feature fusion model to perform feature fusion on at least one type of modality features; after the feature fusion is completed, the fused modality features are obtained.

[0060] The embodiments of the present application do not limit the specific content or type of the feature fusion algorithm or the feature fusion model, which can be set according to actual situations. Exemplarily, the feature fusion model can be a Transformer model. The Transformer model is a deep learning model based on the attention mechanism, with powerful language understanding and generation capabilities. It can focus on and correlate different parts of the input features, effectively integrate various different types of features, and thus achieve a more comprehensive and accurate information expression.

[0061] By performing feature fusion on at least one type of modality features, the deficiencies and limitations of single modality features can be made up for, different sources and different types of modality features can be effectively integrated, and the potential associations and synergistic effects hidden between different modality features can be mined, so as to obtain a more comprehensive and integrated information expression.

[0062] S103. The electronic device generates a target audio sequence for characterizing the music recommendation requirements of the target user based on the fused modality features.

[0063] An audio sequence refers to a discrete element set or an audio information stream composed of multiple audio data units arranged in a specific order. Among them, the audio data unit can be an audio sampling point, an audio frame, an audio segment, or a complete audio file, etc.

[0064] The embodiments of the present application do not limit the specific type and content of the audio sequence, which can be set according to actual situations. Exemplarily, the audio sequence can be a discrete digital sequence formed by digitizing and encoding an audio signal, such as an MP3 file, or an analog audio sequence with continuity and logic containing the original sound wave signal.

[0065] A target audio sequence refers to an audio sequence for characterizing the music recommendation requirements of a target user. The embodiments of the present application do not limit the specific content of the target audio sequence, which can be set according to actual situations.

[0066] Exemplarily, S103 can be implemented as follows: The electronic device predicts the music recommendation requirements of the target user based on the fused modality features; and generates a target audio sequence that conforms to the music recommendation requirements of the target user based on the music recommendation requirements of the target user.

[0067] The embodiments of the present application do not limit the specific manner in which the electronic device predicts the music recommendation requirements of the target user based on the fused modal features, and can be set according to the actual situation.

[0068] In a possible implementation manner, the electronic device may adopt a rule-based requirement prediction method to determine the music recommendation requirements of the target user. For example, when Rule 1 includes that the user likes light music, if the fused modal features represent that the target user likes light music, then the electronic device determines that the fused modal features meet Rule 1. At this time, the music recommendation requirements of the target user may be for music with beautiful melodies, slow rhythms, and ethereal timbres; when Rule 2 includes that the user is driving while fatigued, if the fused modal features represent that the target user is driving while fatigued, then the electronic device determines that the fused modal features meet Rule 2. At this time, the music recommendation requirements of the target user may be for music with exciting melodies, compact rhythms, and wild timbres.

[0069] In another possible implementation manner, the electronic device may adopt a model-based requirement prediction method to determine the music recommendation requirements of the target user. For example, when the rules include that the user likes light music and the user is driving while fatigued, if the fused modal features represent that the target user has no specific music preference, the target user is not in a fatigued driving state, and the target user opens the music application, then the electronic device determines that the fused modal features do not meet the rules. At this time, the model can be used to process the fused modal features, and the music recommendation requirements output by the model are used as the music recommendation requirements of the target user.

[0070] Simply put, the rule-based requirement prediction method is applicable to the situation where the fused modal features meet the predefined clear rules. In this case, the target user has clear music recommendation requirements. At this time, using the rule-based requirement prediction method to determine the music recommendation requirements of the target user, the implementation solution is simple and convenient and not prone to errors; the model-based requirement prediction method is applicable to the situation where the fused modal features do not meet the predefined clear rules. In this case, the target user does not have clear music recommendation requirements. At this time, using the model-based requirement prediction method to determine the music recommendation requirements of the target user, the implementation solution is accurate, efficient, and has a wide application range.

[0071] S104. The electronic device determines at least one target music content based on at least the target audio features of the target audio sequence to form a music recommendation list.

[0072] The audio feature refers to a quantization index or attribute used to describe the essential characteristics of an audio signal.

[0073] The target audio feature refers to the audio feature of the target audio sequence. The embodiments of the present application do not limit the specific type and quantity of the target audio feature, which can be set according to the actual situation. Exemplarily, the target audio feature may include rhythm feature, melody feature, timbre feature, pitch feature, harmony feature, music emotion feature, music style feature, time-domain feature, frequency-domain feature, energy feature, and so on.

[0074] The target music content refers to the music content recommended to the target user. The embodiments of the present application do not limit the specific type and quantity of the target music content, which can be set according to the actual situation.

[0075] The music recommendation list refers to the list formed by all the target music content. The embodiments of the present application do not limit the specific manifestation form of the music recommendation list, which can be set according to the actual situation. Exemplarily, the manifestation form of the music recommendation list can be in the form of a table, including music name, music type, music introduction, author introduction, music duration, detailed content, and so on.

[0076] In a possible implementation manner, S104 may be implemented as: The electronic device matches the target audio sequence with each music content included in the music library based on the target audio feature of the target audio sequence; determines at least one target music content based on the matching degree between the target audio sequence and each music content included in the music library; and sorts the at least one target music content to form a music recommendation list.

[0077] In another possible implementation manner, S104 may be implemented as: The electronic device matches the target audio sequence with each music content included in the music library based on the target audio feature and the target audio change feature of the target audio sequence; determines at least one target music content based on the matching degree between the target audio sequence and each music content included in the music library; and sorts the at least one target music content to form a music recommendation list. Wherein, the target audio change feature refers to the change trend or direction of the target audio sequence in time.

[0078] S105. The electronic device outputs the music recommendation list.

[0079] Exemplarily, S105 may be implemented as: The electronic device displays the music recommendation list in the form of text, table, chart, etc. on the display interface of the display screen, or the electronic device calls the voice assistant to display and broadcast the music recommendation list.

[0080] The present application provides a music recommendation method, the method comprising: obtaining at least one type of modal feature for music recommendation for a target user; each type of modal feature in the at least one type of modal feature is obtained by performing feature extraction on a type of modal data; performing feature fusion on the at least one type of modal feature to obtain a fused modal feature; generating a target audio sequence for characterizing the music recommendation requirements of the target user based on the fused modal feature; determining at least one target music content based at least on the target audio features of the target audio sequence to form a music recommendation list; and outputting the music recommendation list.

[0081] In the solution of the present application, by generating a target audio sequence for characterizing the music recommendation requirements of the user, it is possible to determine at least one target music content based on the target audio features of the target audio sequence, and further form and output a music recommendation list based on the at least one target music content. It can be seen that the solution of the present application fully takes into account that audio features pay more attention to the expression of aspects such as rhythm, melody, harmony, and timbre, and with the help of audio features, they can be accurately mapped to music content, and music content can also be completely and accurately characterized by audio features. Therefore, by generating a target audio sequence for characterizing the music recommendation requirements of the user and determining at least one target music content based on the target audio features of the target audio sequence, accurate matching and recall of music content can be achieved. In this way, the difference in the expression modes between multi-modal features and music content can be effectively compensated, the accuracy and comprehensiveness of music recommendation can be improved, and thus the music recommendation effect can be improved.

[0082] Next, the process in which the electronic device generates a target audio sequence for characterizing the music recommendation requirements of the target user based on the fused modal feature in S103 will be described.

[0083] Referring to Figure 2 the content shown, this process may include but is not limited to the following S201 and S202.

[0084] S201. The electronic device processes the fused modal feature using a first audio model to predict the music recommendation type of the target user.

[0085] The music recommendation type includes at least one of the following: melody type, rhythm type, emotion type, and music style type.

[0086] The first audio model belongs to an inference-based large model.

[0087] The music recommendation type refers to the music type for characterizing the music recommendation requirements of the target user, and at least includes melody type, rhythm type, emotion type, music style type, etc.

[0088] In the embodiments of the present application, the specific contents of the melody type, rhythm type, emotion type, and music style type are not limited and can be set according to actual situations. Exemplarily, the melody type may include: beautiful and smooth, fluctuating, unique and novel, etc.; the rhythm type may include: lively and brisk, steady and powerful, changeable and interesting, etc.; the emotion type may include: happy and pleasant, warm and touching, exciting and inspiring, calm and soothing, etc.; the music style type may include: classical and elegant, popular and fashionable, rock and passionate, jazz and free, folk and simple, electronic and dynamic, world and diverse, etc.

[0089] The first audio model is used to determine and output the music recommendation type based on the fused modal features. In the embodiments of the present application, the specific type of the first audio model is not limited and can be set according to actual situations.

[0090] It can be understood that the first audio model can be any pre-trained large model with reasoning ability. The first audio model having reasoning ability means that the first audio model can analyze and understand the input information, refer to technologies such as deep learning, intelligent reasoning, and knowledge fusion, combine the language rules, logical rules, and massive knowledge learned in the pre-training stage, and analyze and deduce through semantic analysis and pattern recognition, and finally obtain and output a conclusion.

[0091] Previously, multi-modal data related to different users can be collected or gathered, such as user data of different populations, environmental data of different environments, and vehicle data of different vehicles; feature extraction is performed on the multi-modal data to obtain multi-modal features; feature fusion is performed on the multi-modal features to obtain fused modal features; the initial model is trained based on the fused modal features and related prompt words, and the first audio model is obtained after the training is completed.

[0092] In the embodiments of the present application, the specific training process of the first audio model is not limited and can be set according to actual situations.

[0093] Exemplarily, S201 can be implemented as: the electronic device inputs the fused modal features into the pre-trained first audio model; the first audio model uses its reasoning ability to perform computational processing on the input fused modal features, predicts and outputs the music recommendation type corresponding to the fused modal features; the electronic device determines the music recommendation type as the music recommendation type of the target user.

[0094] S202. The electronic device processes the music recommendation type using a second audio model to generate a target audio sequence that conforms to the music recommendation type.

[0095] The second audio model belongs to a generative large model.

[0096] A second audio model, which is used to generate and output an audio sequence that conforms to the music recommendation type based on the music recommendation type. The specific type of the second audio model in the embodiments of the present application is not limited and can be set according to actual situations.

[0097] It can be understood that the second audio model can be any pre-trained large model with generation capabilities. The second audio model having generation capabilities means that the second audio model can analyze and understand the input information, and combine the language rules, logical rules, and vast knowledge learned during the pre-training stage to independently create and generate content with a certain degree of logic and rationality.

[0098] Previously, multiple groups of music recommendation types can be obtained; the initial model is trained based on the multiple groups of music recommendation types and related prompt words, and the second audio model is obtained after the training is completed.

[0099] The specific training process of the second audio model in the embodiments of the present application is not limited and can be set according to actual situations.

[0100] Exemplarily, S202 can be implemented as: the electronic device inputs the music recommendation type into the pre-trained second audio model; the second audio model uses its generation capabilities to perform computational processing on the input music recommendation type, generates and outputs an audio sequence that conforms to the music recommendation type; the electronic device determines this audio sequence as the target audio sequence.

[0101] It should be noted that the first audio model and the second audio model can be the same or different.

[0102] In the case where the first audio model and the second audio model are the same, the first audio model and the second audio model can be collectively referred to as the audio model, and the audio model has both inference capabilities and generation capabilities. Correspondingly, S201 and S202 can be jointly implemented as: the electronic device inputs the fused modal features into the pre-trained audio model; the audio model uses its inference capabilities and generation capabilities to perform computational processing on the input fused modal features, generates and outputs an audio sequence for characterizing the music recommendation needs of the target user; the electronic device determines this audio sequence as the target audio sequence.

[0103] Briefly speaking, by introducing a pre-trained first audio model with reasoning ability to calculate and process the fused modal features, the music recommendation type of the target user can be predicted, enabling accurate reasoning of the music recommendation type and ensuring that the predicted music recommendation type better conforms to the potential interests of the target user; by introducing a pre-trained second audio model with generation ability to calculate and process the music recommendation type and generate a target audio sequence that conforms to the music recommendation type, the accuracy and comprehensiveness of the target audio sequence can be improved, making the generated target audio sequence more in line with the music recommendation type of the target user. Its implementation solution is simple, efficient, and highly scalable.

[0104] Next, the process in S104 where the electronic device determines at least one target music content based on at least the target audio features of the target audio sequence to form a music recommendation list will be described.

[0105] Refer to Figure 3 the content shown. This process may include but is not limited to the following S301 to S304.

[0106] S301. The electronic device determines at least one candidate music content in the music library based on the target audio features; or, based on the target audio features and the target audio change features of the target audio sequence, determines the at least one candidate music content in the music library.

[0107] Candidate music content refers to the music content that is initially determined to be recommended to the target user and has not been screened and sorted yet. In the embodiments of the present application, there are no limitations on the specific type and quantity of the candidate music content, which can be set according to the actual situation.

[0108] Audio change features refer to the changing trend or direction of the audio signal over time.

[0109] Target audio change features refer to the audio change features of the target audio sequence. In the embodiments of the present application, there are no limitations on the specific type and quantity of the target audio change features, which can be set according to the actual situation. Exemplarily, the target audio change features may include rhythm change features, melody change features, timbre change features, pitch change features, harmony change features, music emotion change features, music style change features, time-domain change features, frequency-domain change features, energy change features, and so on.

[0110] In a possible implementation manner, S301 may be implemented as follows: The electronic device uses an audio feature extraction tool to extract audio features from the target audio sequence to obtain target audio features; matches the target audio features with the audio features of each music content included in the music library; and determines at least one candidate music content based on the matching degree between the target audio features and the audio features of each music content included in the music library.

[0111] In another possible implementation, S301 may be implemented as follows: The electronic device uses an audio feature extraction tool to extract audio features from the target audio sequence to obtain target audio features, and uses an audio change feature analysis tool to process and analyze the target audio sequence to obtain target audio change features; matches the target audio features with the audio features of each music content included in the music library, and matches the target audio change features with the audio change features of each music content included in the music library; determines at least one candidate music content based on the matching degree between the target audio features and the audio features of each music content included in the music library, and the matching degree between the target audio change features and the audio change features of each music content included in the music library.

[0112] The embodiments of the present application do not limit the specific types of the audio feature extraction tool and the audio change feature analysis tool, which can be set according to actual situations. For example, the audio feature extraction tool may include software such as Praat, Audacity, MATLAB, etc.; the audio change feature analysis tool may include software such as Sonic Visualiser, FL Studio, etc., and may also include a pre-trained artificial intelligence model with audio change feature analysis capabilities.

[0113] S302. The electronic device determines the at least one target music content from the at least one candidate music content.

[0114] Exemplarily, S302 may be implemented as follows: The electronic device filters the at least one candidate music content based on a pre-set screening rule to obtain at least one target music content.

[0115] The embodiments of the present application do not limit the specific content and quantity of the screening rule, which can be set according to actual situations. Exemplarily, the screening rule may include at least one of the following: representative content screening, duplicate removal of similar content, filtering of listened content, sensitive word filtering, label filtering, and emotional score filtering. Among them:

[0116] Representative content screening refers to screening out high-value music content with representativeness to improve the quality of music recommendations.

[0117] Duplicate removal of similar content means that for multiple music contents with a similarity score greater than a threshold, one of the music contents is retained and the others are deleted to avoid redundancy.

[0118] Filtering of listened content means deleting the music content that the user has already listened to avoid repeated recommendations.

[0119] Sensitive word filtering means deleting the music content containing specific sensitive words to avoid the spread of negative music content.

[0120] Label filtering refers to deleting music content containing specific labels to avoid spreading negative music content.

[0121] Emotional score filtering refers to deleting music content with a negative emotional score to avoid spreading negative music content.

[0122] S303. The electronic device sorts the at least one target music content to obtain the music recommendation list.

[0123] Exemplarily, S303 can be implemented as: The electronic device sorts the at least one target music content based on a preset sorting rule to obtain the music recommendation list.

[0124] The embodiments of the present application do not limit the specific content and quantity of the sorting rule, which can be set according to the actual situation. Exemplarily, the sorting rule can include at least one of the following: matching degree sorting, timeliness sorting, popularity sorting. Among them:

[0125] Matching degree sorting refers to a sorting method according to the level of matching degree.

[0126] Timeliness sorting refers to a sorting method according to the level of closeness to the current hot event.

[0127] Popularity sorting refers to a sorting method according to the level of popularity.

[0128] Briefly speaking, by determining the target audio feature and the target audio change feature of the target audio sequence, the accurate matching between the target audio sequence and the music content included in the music library can be realized, and the accuracy and comprehensiveness of the candidate music content can be improved. Moreover, by screening the candidate music content and sorting the target music content, more accurate and comprehensive high-value music content can be provided for users, and the accuracy and rationality of music recommendation can be enhanced.

[0129] Next, the process in which the electronic device determines the at least one candidate music content in the music library based on the target audio feature and the target audio change feature of the target audio sequence in S301 will be described.

[0130] Refer to Figure 4 the content shown. This method may include but is not limited to the following S401 and S402.

[0131] For each of the at least one first music content included in the music library, perform the following processing to obtain the at least one candidate music content:

[0132] S401. The electronic device determines the audio feature matching degree between the target audio feature and the first audio feature of the first music content, and the audio change feature matching degree between the target audio change feature and the first audio change feature of the first music content.

[0133] The first music content refers to the music content included in the music library. The embodiments of the present application do not limit the type and quantity of the first music content, which can be determined according to the actual situation.

[0134] The first audio feature refers to the audio feature of the first music content. Exemplarily, similar to the target audio feature, the first audio feature can also include rhythm feature, melody feature, timbre feature, pitch feature, harmony feature, music emotion feature, music style feature, time domain feature, frequency domain feature, energy feature, etc.

[0135] The audio feature matching degree refers to the matching degree between the target audio feature and the first audio feature. The embodiments of the present application do not limit the specific manifestation form of the audio feature matching degree, which can be determined according to the actual situation. Exemplarily, the manifestation form of the audio feature matching degree can be in the form of a score or a percentage.

[0136] The first audio change feature refers to the audio change feature of the first music content. Exemplarily, similar to the target audio change feature, the first audio change feature can also include rhythm change feature, melody change feature, timbre change feature, pitch change feature, harmony change feature, music emotion change feature, music style change feature, time domain change feature, frequency domain change feature, energy change feature, etc.

[0137] The audio change feature matching degree refers to the matching degree between the target audio change feature and the first audio change feature. The embodiments of the present application do not limit the specific manifestation form of the audio change feature matching degree, which can be determined according to the actual situation. Exemplarily, the manifestation form of the audio change feature matching degree can be in the form of a score or a percentage.

[0138] Exemplarily, S401 can be implemented as follows: The electronic device calculates the distance value or cosine similarity value between the target audio feature and the first audio feature to obtain the audio feature matching degree between the target audio feature and the first audio feature; calculates the distance value or cosine similarity value between the target audio change feature and the first audio change feature to obtain the audio change feature matching degree between the target audio change feature and the first audio change feature.

[0139] S402. When the audio feature matching degree is greater than the first threshold and the audio change feature matching degree is greater than the second threshold, the electronic device determines the first music content as one of the candidate music contents.

[0140] The first threshold and the second threshold are used to assist in determining candidate music content. The embodiments of the present application do not limit the specific values of the first threshold and the second threshold, which can be determined according to the actual situation.

[0141] Exemplarily, S402 can be implemented as: the electronic device determines the magnitude relationship between the audio feature matching degree and the first threshold, and the magnitude relationship between the audio change feature matching degree and the second threshold; in the case where the audio feature matching degree is greater than the first threshold and the audio change feature matching degree is greater than the second threshold, the first music content is determined as a candidate music content.

[0142] Briefly speaking, by determining the audio feature matching degree between the target audio feature and the first audio feature of each first music content in the music library, and the audio change feature matching degree between the target audio change feature and the first audio change feature of each first music content in the music library, accurate matching between the target audio feature and the first audio feature, and between the target audio change feature and the first audio change feature can be achieved. Furthermore, accurate matching between the target audio sequence and the first music content can be achieved, improving the accuracy and comprehensiveness of the candidate music content, and the solution is simple and efficient.

[0143] It should be noted that in S301, the electronic device determines at least one candidate music content based on the target audio feature in the music library, including: for each of the at least one first music content included in the music library, performing the following processing to obtain the at least one candidate music content: determining the audio feature matching degree between the target audio feature and the first audio feature of the first music content; in the case where the audio feature matching degree is greater than the first threshold, determining the first music content as one of the candidate music contents.

[0144] Next, the process of the electronic device determining the audio feature matching degree between the target audio feature and the first audio feature of the first music content in S401 will be described.

[0145] In the case where the target audio feature includes a target rhythm feature, a target melody feature, and a target timbre feature, and the first audio feature includes a first rhythm feature, a first melody feature, and a first timbre feature, referring to Figure 5 the content shown, this process may include but is not limited to the following S501 to S503.

[0146] The target rhythm feature, the target melody feature, and the target timbre feature refer to the rhythm feature, the melody feature, and the timbre feature of the target audio feature; the first rhythm feature, the first melody feature, and the first timbre feature refer to the rhythm feature, the melody feature, and the timbre feature of the first audio feature.

[0147] In practice, the target audio features and the first audio features may also include other features, which will not be enumerated one by one here.

[0148] S501. The electronic device respectively determines the rhythm matching degree between the target rhythm feature and the first rhythm feature, the melody matching degree between the target melody feature and the first melody feature, and the timbre matching degree between the target timbre feature and the first timbre feature.

[0149] The rhythm matching degree refers to the matching degree between the target rhythm feature and the first rhythm feature.

[0150] The melody matching degree refers to the matching degree between the target melody feature and the first melody feature.

[0151] The timbre matching degree refers to the matching degree between the target timbre feature and the first timbre feature.

[0152] The embodiments of the present application do not limit the specific forms of the rhythm matching degree, the melody matching degree, and the timbre matching degree, which can be set according to actual situations.

[0153] Exemplarily, S501 can be implemented as follows: The electronic device calculates the distance value or cosine similarity value between the target rhythm feature and the first rhythm feature to obtain the rhythm matching degree; calculates the distance value or cosine similarity value between the target melody feature and the first melody feature to obtain the melody matching degree; calculates the distance value or cosine similarity value between the target timbre feature and the first timbre feature to obtain the timbre matching degree.

[0154] S502. The electronic device respectively determines the first weight corresponding to the rhythm matching degree, the second weight corresponding to the melody matching degree, and the third weight corresponding to the timbre matching degree.

[0155] The first weight, the second weight, and the third weight respectively refer to the weights corresponding to the rhythm matching degree, the melody matching degree, and the timbre matching degree. The embodiments of the present application do not limit the specific values of the first weight, the second weight, and the third weight, which can be set according to actual situations.

[0156] Exemplarily, the first weight corresponding to the rhythm matching degree, the second weight corresponding to the melody matching degree, and the third weight corresponding to the timbre matching degree can be determined in advance based on human experience or a specific algorithm or a specific model and stored in the electronic device. Correspondingly, S502 can be implemented as follows: The electronic device obtains the first weight, the second weight, and the third weight from the local or the cloud.

[0157] S503. The electronic device performs weighted averaging on the rhythm matching degree, the melody matching degree, and the timbre matching degree based on the first weight, the second weight, and the third weight to obtain the audio feature matching degree.

[0158] Exemplarily, S503 may be implemented as: calculating the product of the rhythm matching degree and the first weight, the product of the melody matching degree and the second weight, and the product of the timbre matching degree and the third weight by the electronic device; and determining the sum of these three products as the audio feature matching degree.

[0159] Briefly, by determining the rhythm matching degree between the target rhythm feature and the first rhythm feature, the melody matching degree between the target melody feature and the first melody feature, the timbre matching degree between the target timbre feature and the first timbre feature, as well as the first weight corresponding to the rhythm matching degree, the second weight corresponding to the melody matching degree, and the third weight corresponding to the timbre matching degree, the audio feature matching degree can be accurately determined, taking into account the matching of the target audio features and the first audio features in terms of rhythm, melody, timbre, etc., and achieving precise matching between the target audio features and the first audio features.

[0160] It should be noted that in the case where the target audio change feature includes a target rhythm change feature, a target melody change feature, and a target timbre change feature, and the first audio change feature includes a first rhythm change feature, a first melody change feature, and a first timbre change feature, the electronic device determines the audio change feature matching degree between the target audio change feature and the first audio change feature of the first music content in S401, including: respectively determining the rhythm change feature matching degree between the target rhythm change feature and the first rhythm change feature, the melody change feature matching degree between the target melody change feature and the first melody change feature, and the timbre change feature matching degree between the target timbre change feature and the first timbre change feature; respectively determining the fourth weight corresponding to the rhythm change feature matching degree, the fifth weight corresponding to the melody change feature matching degree, and the sixth weight corresponding to the timbre change feature matching degree; and performing weighted averaging on the rhythm change feature matching degree, the melody change feature matching degree, and the timbre change feature matching degree based on the fourth weight, the fifth weight, and the sixth weight to obtain the audio change feature matching degree.

[0161] Next, the process of the electronic device determining the at least one target music content from the at least one candidate music content in S302 will be described.

[0162] Refer to Figure 6 the content shown, this process may include but is not limited to the following S601 and S602.

[0163] S601. The electronic device determines the representative identifier of each candidate music content in the at least one candidate music content.

[0164] The representative identifier is used to characterize whether the candidate music content is representative in the music genre to which it belongs.

[0165] The embodiments of the present application do not limit the specific content of the music genre, which can be determined according to the actual situation. Exemplarily, in the case of dividing music content by style, the music genre can include pop, rock, classical, jazz, folk, electronic, hip-hop, metal, blues, etc.; in the case of dividing music content by regional culture, the music genre can include Chinese-style music, European and American music, African music, Latin music, etc.; in the case of dividing music content by historical period, the music genre can include Baroque music, classical music, romantic music, modern music, etc.; in the case of dividing music content by performance form, the music genre can include solo singing, chorus, symphony, chamber music, etc.; in the case of dividing music content by usage scenario, the music genre can include dance music, film and television soundtracks, game music, etc.; in the case of dividing music content by emotional expression, the music genre can include cheerful music, sad music, exciting music, etc.

[0166] The embodiments of the present application do not limit the specific content of the representative identifier, which can be determined according to the actual situation. Exemplarily, the representative identifier can be "0" or "1", where "0" indicates non-representativeness and "1" indicates representativeness.

[0167] For example, in the case where the music genre to which it belongs is "pop music", if the candidate music content is song A by singer 1, then the representative identifier of song A can be "1", indicating that song A is representative in the music genre "pop music" to which it belongs; if the candidate music content is classic song B, then the representative identifier of song B can be "0", indicating that song B is not representative in the music genre "pop music" to which it belongs.

[0168] Exemplarily, the representative identifier of each candidate music content in at least one candidate music content can be determined in advance based on human experience or a specific algorithm or a specific model and stored in an electronic device. Correspondingly, S601 can be implemented as: the electronic device obtains the representative identifier of each candidate music content in at least one candidate music content from local or the cloud.

[0169] S602. The electronic device determines, as the at least one target music content, the candidate music content in the at least one candidate music content whose representative identifier takes a first value.

[0170] The first value characterizes that the candidate music content is representative in the music genre to which it belongs.

[0171] The first value is used to assist in determining the target music content. The embodiments of the present application do not limit the specific content of the first value, which can be set according to actual situations. Exemplarily, the first value can be "1", and "1" indicates representativeness.

[0172] Exemplarily, S602 can be implemented as follows: The electronic device traverses at least one candidate music content; for each candidate music content, it determines whether the representative identifier of the candidate music content is the first value. If so, it determines the candidate music content as a target music content. If not, it continues to determine whether the representative identifier of the next candidate music content is the first value until the traversal is completed; after the traversal is completed, at least one target music content is obtained.

[0173] Briefly speaking, by determining the representative identifier of each candidate music content, at least one target music content with representativeness can be screened out from at least one candidate music content. In this way, high-value music content can be provided for users, effectively improving the quality and accuracy of music recommendations.

[0174] Next, the process in which the electronic device sorts the at least one target music content in S303 to obtain the music recommendation list will be described.

[0175] Refer to Figure 7 the content shown. This process may include but is not limited to the following S701 and S702.

[0176] S701. The electronic device determines the timeliness score of each target music content among the at least one target music content.

[0177] The timeliness score is used to represent the closeness between the target music content and the current hot event.

[0178] The embodiments of the present application do not limit the specific content and quantity of the current hot events, which can be set according to actual situations. Exemplarily, the current hot events may include Hot Event 1, Hot Event 2, Hot Event 3, etc.

[0179] Exemplarily, the electronic device can obtain the current hot events in real time based on a specific application programming interface such as a news application programming interface in advance. Correspondingly, S701 can be implemented as follows: The electronic device traverses at least one target music content; for each target music content, it determines the closeness between the target music content and the current hot events based on a specific algorithm or a specific model to obtain the timeliness score of the target music content; after the traversal is completed, the timeliness scores of each target music content among the at least one target music content are obtained.

[0180] For example, when the target music content is song C, if the current hot event is hot event 3, then the timeliness score of song C can be 100; if the current hot event is hot event 2, then the timeliness score of song C can be 15.

[0181] S702: The electronic device sorts the at least one target music content in descending order of timeliness score to obtain the music recommendation list.

[0182] Exemplarily, S702 can be implemented as follows: the electronic device traverses at least one target music content; for any two target music contents among the at least one target music content, performs pairwise comparisons according to the timeliness scores, determines the sorting results of the two target music contents, and adjusts the arrangement order of the two target music contents; after the traversal is completed, a music recommendation list is obtained.

[0183] Simply put, by determining the timeliness score of each target music content, at least one target music content can be sorted in descending order according to the timeliness score, so that music content with a high timeliness score is displayed first. In this way, the timeliness and pertinence of music recommendations can be effectively improved, allowing users to more quickly access music content that is closely related to current hot events, making the music recommendation system more intelligent and in line with actual conditions.

[0184] Next, the process of the electronic device determining the timeliness score of each of the at least one target music content in S701 is described.

[0185] refer to Figure 8 As shown in the content, the process may include but is not limited to the following S801 to S803.

[0186] For each of the at least one target music content, the following processing is performed:

[0187] S801: The electronic device obtains the release time of the target music content and the relevance between the target music content and the current hot event.

[0188] The relevance is used to characterize the degree of association or similarity between the target music content and the current hot event. The embodiment of the present application does not limit the specific form of the relevance, and it can be set according to the actual situation.

[0189] Exemplarily, S801 may be implemented as follows: the electronic device obtains the release time of the target music content from the music library; and determines the relevance of the target music content to the current hot event based on a specific algorithm or a specific model.

[0190] S802. The electronic device determines a target timeliness function of the target music content based on the target music type of the target music content.

[0191] The target music type refers to the music type of the target music content. The embodiments of the present application do not limit the specific content and classification method of the target music type, which can be determined according to the actual situation.

[0192] The timeliness function refers to a mathematical expression or algorithm for determining the timeliness score of music content.

[0193] The target timeliness function refers to the timeliness function corresponding to the target music content. The embodiments of the present application do not limit the specific content of the target timeliness function, which can be set by itself according to the actual situation.

[0194] Exemplarily, the corresponding relationship between the music type and the timeliness function can be determined in advance based on human experience or a specific algorithm or a specific model and stored in the electronic device. Correspondingly, S802 can be implemented as: the electronic device determines the target music type of the target music content; searches for the target music type in this corresponding relationship; and determines the timeliness function corresponding to the target music type as the target timeliness function.

[0195] S803. The electronic device determines the timeliness score of the target music content based on the release time, the relevance, and the target timeliness function.

[0196] Exemplarily, S803 can be implemented as: the electronic device calculates the difference between the release time and the current time, substitutes this difference and the relevance into the target timeliness function, and calculates the timeliness score of the target music content.

[0197] Briefly, considering that the timeliness change trends of music content corresponding to different music types are different. For example, the timeliness of pop music changes relatively quickly and its timeliness will decay rapidly within a few days, while the timeliness of classical music and jazz music changes relatively slowly and may not start to decay until several months or even several years. Therefore, by determining the timeliness functions corresponding to different music types, and obtaining the release time of the target music content and the relevance between the target music content and the current hot events, the timeliness score of the target music content can be determined more accurately, and then a more comprehensive and higher-value music recommendation service can be provided for users.

[0198] The music recommendation method provided by the embodiments of the present application may further include an adjustment process of the target audio sequence.

[0199] After generating the target audio sequence for characterizing the music recommendation needs of the target user based on the fused modal features, refer to Figure 8 the content shown, this process may include but is not limited to the following S901 to S903.

[0200] S901. The electronic device obtains environmental data related to acoustic characteristics in the environment where the target user is currently located.

[0201] The environmental data can be any data related to acoustic characteristics in the environment where the target user is currently located. The embodiments of the present application do not limit the specific content of the environmental data, which can be set according to actual situations. Exemplarily, the environmental data can include: weather, ambient sound, location, congestion situation, etc.

[0202] Exemplarily, S901 can be implemented as: The electronic device obtains the environmental data of the environment where the target user is currently located through a specific sensor or a specific application programming interface.

[0203] S902. The electronic device generates environmental audio corresponding to the environmental data based on the environmental data.

[0204] The environmental audio refers to the audio used to create a specific environmental atmosphere. The embodiments of the present application do not limit the specific content of the environmental audio, which can be set according to actual situations.

[0205] In a possible implementation manner, S902 can be implemented as: When the environmental data includes ambient sound, the electronic device performs enhancement or weakening processing on the ambient sound to obtain environmental audio corresponding to the environmental data.

[0206] In another possible implementation manner, S902 can be implemented as: When the environmental data does not include ambient sound, the electronic device uses an audio generation model to analyze and process the environmental data to generate environmental audio corresponding to the environmental data.

[0207] For example, when the environmental data includes weather and the weather is thunderstorm, the environmental audio generated by using the audio generation model can include the sounds of wind, rain, thunder and lightning.

[0208] S903. The electronic device adds the environmental audio to the target audio sequence to obtain the adjusted target audio sequence.

[0209] Exemplarily, S903 can be implemented as: The electronic device adds the environmental audio to the target audio sequence through audio editing software to obtain the adjusted target audio sequence.

[0210] Simply put, by generating ambient audio corresponding to the ambient data based on the ambient data related to the acoustic characteristics of the target user's current environment, and adding the ambient audio to the target audio sequence, the realism of the target audio sequence can be increased, and the dimension and level of the target audio sequence can be enriched, thereby making the target music content determined based on the target audio sequence more realistic and rich, thereby providing users with music recommendation services that are adapted to the real environment.

[0211] Below, the music recommendation method provided in the embodiment of the present application is explained through a complete embodiment.

[0212] With the rapid development of artificial intelligence technology, personalized recommendation systems have been widely used in various fields. In the field of in-car music, how to provide accurate music recommendation services based on the personalized needs and preferences of users has become an urgent problem to be solved. In-car music comes from different media sources. When a user plays a piece of music, it can recommend music that is close to or has a certain semantic relevance to the music based on the contextual semantics of the listening, providing users with a richer extended reading experience, thereby improving product stickiness and user activity. At the same time, in music recommendation, combined with the actual characteristics of large amount of music data, diverse content and more repetition, it is necessary to solve the problems of music duplication and content diversity in music-related recommendations, and provide objective and neutral music recommendation content. Most traditional music recommendation systems are based on rules or simple machine learning models, which cannot handle large-scale data and complex user behaviors. At the same time, the various links of the traditional music recommendation process are relatively scattered, resulting in poor recommendation results.

[0213] The big model can couple data feature learning, type prediction, recommendation calculation, etc. into an end-to-end form, learn better results during deep learning, and the big model can achieve tasks in an end-to-end efficient manner while obtaining better recommendation accuracy. It can be seen that the recommendation method based on the big model is of great application value. Therefore, in order to solve the above problems, the embodiment of the present application discusses the two aspects of the end-to-end efficient fusion of different link features of the big model and the more powerful high-dimensional user information learning ability, so as to help achieve better in-car music recommendations. The music recommendation method provided in the embodiment of the present application relates to the field of artificial intelligence and music recommendation technology, and in particular to a personalized in-car music recommendation design scheme based on a big model. The scheme uses big model technology to deeply mine and analyze massive data, build a user personalized model, and achieve more accurate music recommendations.

[0214] The overall business module of the personalized in-car music recommendation design solution provided in the embodiment of the present application is as follows: Figure 10 As shown, including:

[0215] Data Collection Module 1001: This module is responsible for collecting personalized data of users, including but not limited to information such as users' music preferences, listening habits, driving habits, and vehicle status. These data can be obtained through in-vehicle sensors, user inputs, and external data sources, etc.

[0216] Data Preprocessing Module 1002: This module performs preprocessing operations on the collected data, such as cleaning, deduplication, and classification, to ensure the quality and accuracy of the data. At the same time, this module is also responsible for feature extraction and representation learning of the data to better utilize large models for training and inference.

[0217] Data Feature Extraction Module 1003: This module is used to perform data analysis on the input data and extract the features of the input data using algorithm models.

[0218] Scene Recognition Module 1004: This module is used to accurately identify and distinguish various different music scenes. This module can perceive and analyze user information, the environment information where the user is located, etc. By capturing and understanding these scene features, it can identify various different music scenes, providing an important basis for subsequent music recommendations, etc., so as to match the most suitable music for users according to specific music scenes, and can also be used to dynamically adjust the music playback strategy according to the change of the scene to achieve more intelligent and personalized music recommendations.

[0219] Emotion Recognition Module 1005: This module can monitor the user's driving state and emotional changes in real time, can identify and understand the user's emotional state, and dynamically adjust the music recommendation strategy accordingly. For example, when the user is in a happy state, the system may recommend some lively pop music; when the user feels frustrated, it may recommend some soothing light music or healing songs; when it detects that the user is driving fatigued, the system will automatically switch to a more relaxing and soothing music list to remind and cheer up the user; when it detects that the user is in an excited driving state, the system may recommend some music with a fast rhythm and full of dynamics.

[0220] Music Recommendation Module 1006: Based on the output of the large model, this module recommends music content that meets the user's personalized needs. The recommended music content not only considers the user's music preferences and listening habits, but also factors such as vehicle status and driving environment. For example, during a long-distance drive, the system may recommend some relaxing and soothing music to relieve driving fatigue; when driving in a congested urban area, the system may recommend some music with a fast rhythm to improve the driving mood.

[0221] User Feedback Module 1007: This module allows users to evaluate and provide feedback on the recommended music. The system continuously adjusts and optimizes the recommendation strategy based on the users' feedback to improve the accuracy of recommendations and user satisfaction. In addition, users can also input information such as their music preferences and listening habits through this module to further improve the personalized model.

[0222] Large Model Pre-training Module 1008: This module uses the preprocessed data to train a large-scale deep learning model. The model adopts an advanced neural network architecture and, through the powerful semantic understanding and language generation capabilities of the large model, learns the patterns of the data and the characteristic distributions of user behavior patterns to construct a general large language model, thereby better understanding the music preferences and needs of users.

[0223] In actual implementation, the outputs of the Scene Recognition Module 1004, Emotion Recognition Module 1005, Music Recommendation Module 1006, and User Feedback Module 1007 can be uniformly used as the input of the large model. Based on the large model learning paradigm, through the Large Model Pre-training Module, an end-to-end recommendation system can be constructed. Additionally, it is worth mentioning that, leveraging the advantage of the large model in learning the general characteristics of data, the embodiments of this application map the text and speech data features of two different types into a unified feature space for learning and unified prediction of results, thus implementing the final end-to-end recommendation algorithm.

[0224] Next, the music recommendation system architecture of the personalized in-vehicle music recommendation design solution provided by the embodiments of this application will be described. As Figure 11 shown, the music recommendation system architecture of the solution of this application includes:

[0225] Data Storage Layer 1101: Used to store the underlying music originals and system basic logs.

[0226] Feature Mining Layer 1102: Responsible for music preprocessing and feature mining, as well as basic algorithms related to Natural Language Processing (NLP).

[0227] Data Index Layer 1103: Used to build indexes for the refined music features and user behaviors to ensure high-performance retrieval.

[0228] Recommendation Strategy Layer 1104: Used to implement various recommendation-related recall, ranking, and filtering strategies.

[0229] Application Interface Layer 1105: Used to provide encapsulated data interface services for the front end and ensure the high availability of the services.

[0230] In the above-mentioned recommendation system architecture, the data preprocessing module of the feature mining layer 1102 is mainly responsible for mining various features in music, including but not limited to keywords, label classification, semantic vectors, structured fields, simhash signatures, etc. The above features will be stored in the data indexing layer 1103, and the big data indexing framework provides fast conditional retrieval for the upper-layer recommendation strategy. The label system and tagging strategy for semantic tagging can be configured and optimized in the music semantic tagging module, which will use a part of NLP algorithm services and front-end tags to configure and manage the platform.

[0231] The music similarity calculation and clustering module uses the multi-dimensional features of music to mine the hot clusters of music globally, which can support the clustering requirements at different levels of granularity: overly similar music (i.e., the leaf nodes of hierarchical clustering) can be merged to avoid a large number of the same content received by the upper-layer recall strategy and improve the recommendation efficiency. The music clusters that are close but not the same can provide different music styles with the same theme and provide multi-angle recommendations.

[0232] The recommendation result recall, sorting, and filtering module is responsible for recalling relevant music using different types of strategies, such as keyword recall, label recall, semantic vector recall, etc., and fusing the recall results of different types. The recall ratio or priority of each strategy can be adjusted through business configuration.

[0233] The sorting module provides the default sorting logic, including similarity factors, popularity factors, time decay factors, etc., and can update the key parameters in the sorting algorithm according to the click-through feedback of the recommendation results to optimize the recommendation click-through rate.

[0234] The filtering module is responsible for filtering out the results that are not suitable for recommendation to users according to the music the user has listened to and business rules. If the number of results after filtering is small, the result set can be supplemented through secondary recall.

[0235] The data interface module supplements the fields such as title, author, source, and time required for front-end display for the recommendation results, and can mark the reasons for some users to understand the recommendations, such as "late-night song list", "summer freshness", "mood song list", etc. Finally, these data are encapsulated in the interface service to provide necessary load balancing, caching, and exception handling functions.

[0236] Next, the music recommendation strategy in the personalized in-vehicle music recommendation design solution provided by the embodiments of the present application will be described.

[0237] Music recommendations can adopt different recommendation strategies for different product forms. Common strategies include: recommendation based on popularity, recommendation based on context, recommendation based on user interests, and community-based recommendation (such as collaborative filtering). In the embodiments of this application, for music-related recommendations, emphasis is placed on designing recommendation strategies by considering the relevance in the in-vehicle scenario and the user's current mood.

[0238] Before making music recommendations, music features can be mined first. Music feature mining includes the following aspects:

[0239] (1) Keyword extraction. Core keywords of music titles and lyrics are refined through models such as Term Frequency-Inverse Document Frequency (TF-IDF) and entity recognition, and the relevant weights of each keyword are calculated. For keywords with specific meanings, directional weight adjustment can be performed through rules. The number of retained keyword features is determined by the weights, and long-tail words with too low weights can be selected for truncation.

[0240] (2) Entities and business tags. The solution of this application includes an algorithm service module for tagging music. The tags mainly include: a) Style tags: such as jazz, classical, pop, rock, folk, etc.; b) Mood tags: such as cheerful, sad, peaceful, excited, etc.; c) Instrument tags: such as piano, guitar, violin, saxophone, etc.; d) Theme tags: such as love, friendship, nature, history, etc.; e) Era tags: such as classical music period, romantic music period, modern music period, etc.; f) Region tags: such as European music, American music, Asian music, etc.; g) Singer / band tags: such as singer 1, singer 2, singer 3, etc.; h) Language tags: such as Chinese, English, Japanese, etc.; i) Tempo tags: such as slow, medium, fast, etc.; j) Emotion tags: such as love, friendship, happiness, etc. This part of the tags will be used in the music recommendation engine as the basis for underlying feature mining strategies.

[0241] (3) Topic vector. Above the lexical and tagging levels, semantic connotations behind music can be extracted through pre-training. Such semantic connotations can be represented in the form of topic vectors. Mainstream technologies or models used in the industry include the Word to Vector (Word2Vec) model, the Bidirectional Encoder Representation from Transformers (BERT) model, Probabilistic Latent Semantic Analysis (PLSA), etc. The solution of this application uses the music data accumulated by itself to train a deep learning semantic model in the music field, which can extract the full-text semantics of music as topic vectors and be used for subsequent relevance calculation, clustering mining, and music recommendation.

[0242] When performing music recommendation, music recall strategies based on keywords, tags, semantic vectors, etc. can be adopted. After all music has been preprocessed and features such as keywords, tags, and semantic vectors have been extracted, the system will construct these features into a big data index for the next step of recommendation recall. When a user listens to a piece of music a, the system will calculate the relevance between a and other music b1~bn in the global music library and obtain a recommended candidate set according to the following recall strategies:

[0243] (1) Keyword matching recall. This strategy is based on the keyword matching degree between music a and bn. Common calculation methods include TF-IDF, Best Matching 25 (BM25), the cosine angle between keyword vectors, etc., and music with a higher keyword matching degree is recalled first.

[0244] (2) Semantic tag recall. This strategy is based on the semantic tag matching degree between music a and bn. First, music is converted into a tag vector or a semantic topic vector, and then the vector distance is calculated and results with a closer distance are recalled first. Common methods include the cosine distance between vectors, Euclidean distance, Manhattan distance, etc. The matching of this recall strategy is not restricted by word segmentation, but it requires the prior establishment of a relatively rich semantic or business tag system and the ability to effectively tag and classify music. A piece of music can be associated with multiple tags or subject classifications, and each classification can also have a probability weight.

[0245] (3) Recall content balance. The final result of the recall strategy can be mixed based on the above multiple recall strategies. In this process, special attention should be paid to the richness of the mixed result. Otherwise, a large number of recommended music that are related to a but have little difference from each other will be recalled, bringing listening fatigue or even abandonment to users. To balance the richness of the recall results, such as Figure 12As shown in the figure, operations such as music clustering, duplicate removal of similar music, duplicate removal of music that users have already listened to, and recommendation of representative music can be performed during the recall stage. By performing cluster analysis on the recommendation results during the recall stage, overly similar recommended music can be de-duplicated, or music with representative views can be selected for recommendation to avoid overly redundant recommendation results.

[0246] When making music recommendations, a music sorting strategy can also be adopted. Sorting is the core part of the entire recommendation model and the last link in the entire recommendation system, playing an important role in re-sorting the recall set for each user, and displaying the content that users like more in a more prominent position. The music sorting strategy includes:

[0247] (1) Semantic relevance sorting. The most core sorting logic for context recommendation is the relevance between the recommendation result and the current music. This relevance will integrate various indicators such as keyword matching degree + tag matching degree + semantic vector matching degree, and be summarized as a comprehensive semantic relevance through weighted aggregation. The weights of various indicators can be configured through business experience in the initial stage; in the later stage, the optimal parameter configuration can be fitted through the actual click-through rate of users and the click volume prediction model.

[0248] (2) Time decay sorting. The timeliness and popularity of music will weaken over time, and users are more inclined to listen to newer music. Therefore, during the actual sorting process, time decay penalties are imposed on different types of music. One type of time penalty function is shown in the following formula (1):

[0249]

[0250] Among them, t represents the time length since the release of the music to be recommended; T and G are adjustable parameters, T is the smoothing value, and G is the time penalty factor. The curve graphs of the time penalty function for different values of T and G are as Figure 13 shown. When T and G take different values, the curve of the penalty function will have different trend changes.

[0251] For different types of music, the penalty function is different. Generally speaking, pop music has strong timeliness, and its timeliness decays rapidly within a few days, that is, T is set to be smaller and G is set to be larger; while for classical music, jazz music, etc., it may take 1 to 3 months to become out of date, and T can be set to be medium and G to be smaller. In application, different values of T and G can be selected according to the actual click feedback of different types of content or music to fit the actual curve, and then the optimal T value and G value for this column can be obtained.

[0252] (3) Music popularity sorting. Since the recommended music comes from different theme clusters, and different music also represents different popularity levels, therefore, this popularity can also be used as one of the factors for recommendation sorting, and the higher the popularity, the higher the ranking.

[0253] When making music recommendations, a music filtering strategy can also be adopted to reduce the redundancy of the final recommended music content. Specifically, it includes:

[0254] (1) Similarity deduplication filtering. For overly similar music, deduplication is performed through simhash calculation to prevent users from seeing redundant information. The filtering priority can be specified according to factors such as data sources or music release time.

[0255] (2) Filtering of music already listened to by the user. Generally, music that the user has already listened to is not recommended to be listed as a recommendation result again. To record and filter the music already listened to by the user in the shortest possible time, two solutions can be considered: 1) Directly cache the identification (ID) of the music already listened to on the client side and perform post-filtering on the results returned by the recommendation engine. To ensure that there is still enough music after filtering, the recommendation engine needs to return a candidate music set that exceeds the number of items in the list. 2) Cache the music already listened to by the user in a key-value (KV) manner in the background and directly filter it before the recommendation engine returns the results to the front end. If the number of filtered music is insufficient, it is directly supplemented before responding to the request.

[0256] (3) Business rule filtering. Specific business rules can be specified to filter the candidate set returned by the recommendation engine, and the management background will provide a configuration page to maintain these rules: a) Sensitive word filtering. Music containing specific sensitive words is filtered and not recommended. b) Tag filtering: Music containing specific tags is filtered and not recommended. c) Sentiment score filtering: Music with a specific range of positive and negative sentiment scores is not recommended. d) Combined condition filtering: Multiple above conditions are combined for filtering.

[0257] Next, the end-to-end multimodal large model algorithm result fusion and prediction process in the personalized in-vehicle music recommendation design solution provided by the embodiments of the present application will be described.

[0258] The overall process of this process is to separately extract features from different types of data such as text and speech and then perform feature fusion to fuse them into a unified semantic feature representation, and then learn through a unified encoder-decoder Transformer. The overall flowchart of the end-to-end multimodal large model algorithm of the solution of the present application is as Figure 14As shown, the text features after Byte Pair Encoding (BPE) and Embedding processing, and the audio features (equivalent to the above-mentioned at least one type of modal features) after Audio Spectrogram Transformer (AST) encoding and Perceiver processing are uniformly input into the Transformer model to learn a high-dimensional unified feature (equivalent to the above-mentioned fused modal feature). Then, the Transformer outputs discrete tokens that can represent text tokens and speech tokens. Subsequently, based on these semantic tokens, the results can be output. This high-dimensional unified feature (equivalent to the above-mentioned fused modal feature) is input into the Vector Quantized Generative Adversarial Network (VQGAN) (equivalent to the above-mentioned first audio model and second audio model), and an audio sequence including speech, ambient sound, and music (equivalent to the above-mentioned target audio sequence) is output.

[0259] The highlight of VQGAN is that it uses a codebook to discretely encode the intermediate features of the model and uses the Generative Pre-trained Transformer-2 (GPT-2) as the encoding and generation tool. The idea of the codebook has been proposed in the Vector Quantized-Variational Auto Encoder (VQ-VAE). The overall architecture of VQGAN is roughly to replace the encoding generator of VQVAE from pixelCNN with Transformer, which is more conducive to learning the serialized data features of text and speech. And during the training process, the discriminator of PatchGAN is added to the adversarial loss.

[0260] Among them, VAE is a powerful generative model. From the perspective of VAE, there is an Encoder that encodes the data into the latent space (z = encoder(x)), and then a Decoder is used to reconstruct the data from the latent space (x = decoder(z)). For VAE, a restriction is added, that is, making z satisfy the isotropic Gaussian distribution (referred to as prior, the prior distribution). The advantage of doing this is that after training, the Encoder can be discarded, and a z can be randomly sampled from this prior, and then an x can be generated through the Decoder, where x refers to the features representing the target text or audio.

[0261] The training of VAE needs to be understood from the perspective of probability. If we look at the process of z = encoder(x) from a probability perspective, it is to let the Encoder learn a conditional probability. And the decoder learns another conditional probability p θ (x|z). At the same time, to make z follow a Gaussian prior, this prior can be denoted as p(x). In this way, the loss function can be written according to this logic, as shown in the following formula (2):

[0262]

[0263] Among them: ELBO is the Evidence Lower Bound, p θ (x|z) are conditional probabilities respectively, and p(z) is the prior distribution.

[0264] Regarding the rates mentioned above Each quantity is a discrete integer, so this number can be written in one-hot form and thus regarded as a probability distribution. There are a total of K dimensions, and each dimension represents the probability of the corresponding e_ in the codebook i (i = 1, 2,..., K). From the perspective of VAE, give this K-dimensional distribution a uniform distribution as the prior, that is, p i = 1 / K, so in the Evidence Lower Bound (ELBO) This term becomes a constant, as shown in the following formula (3):

[0265]

[0266] Among them, the first term represents the contribution of the dimension with 1 in the one-hot to the KL divergence, and the second term represents the contribution of the other dimensions. ELBO, that is, evidence low bound. Evidence refers to x, and ELBO represents the minimum expectation of evidence. Making this lower bound as large as possible, the resulting probability model will be more likely to generate the x value seen here.

[0267] Each dimension of the latent variable z in VAE is a continuous value, while the biggest feature of VQ-VAE is that each dimension of z is a discrete integer, which conforms to some natural modalities. For example, Language is a sequence of symbols, or reasoning, planning and predictive learning. Therefore, VQ-VAE can successfully model important features that span multiple dimensions in the data space. For example, a phoneme in audio will last for many samples / frames, rather than learning some very detailed things. In the discrete coding method proposed in VQVAE, each dimension of the encoded feature is a discrete value, which conforms to some natural modalities.

[0268] Then, the VQGAN used in this solution combines VQVAE and the GAN network. At the same time, based on the specific requirements of this task, improvements are made in the network learning loss function. Specifically, the loss function is as shown in the following formula (4):

[0269]

[0270] Among them, L is the loss function, which uses three branches to form the total loss function. The first term is used to train the encoder and decoder. During backpropagation, the gradient of z_q(x) is directly copied to z_e(x), rather than to the embedding in the codebook. So this term only trains the encoder and decoder. The second term is used to train the codebook, making the embedding in the codebook approach its nearest z_e(x) to minimize the loss. The third term is used to train the encoder, aiming to encourage the output of the encoder to remain close to the selected codebook vector to prevent it from fluctuating too frequently from one code vector to another, that is, to prevent the output of the encoder from jumping frequently among the codebook embeddings.

[0271] So far, the task based on the overall branch can be trained to completion.

[0272] A music recommendation system with high recommendation accuracy first requires an accurate speech recognition algorithm. The accurate recognition and prediction of speech determine the quality of the recommendation. In the speech recognition link of the large model in the embodiment of this application, a fine-grained VQ-GAN is added as a speech decoder, which can fuse and predict the results of the multi-modal tokens output by the Transformer. The mathematical description of the self-attention mechanism used by the Transformer is as shown in the following formula (5):

[0273]

[0274] Among them, Q, K, and V are three Query, Key, and Value matrices for attention calculation after linear mapping.

[0275] The conditional probability is as shown in the following formula (6):

[0276] p(s|c) = ∏ i p(s i |s<i, c) Formula (6);

[0277] Among them, p(s|c) represents the probability of sequence s given the condition c. p(s i |s<i, c) represents the probability of the i-th element s given the first i-1 elements (i.e., s<i) in the known sequence and the condition c. i of.

[0278] Based on the speech feature c, the sequence feature sequence (equivalent to the above-mentioned target audio sequence) is obtained, which is called s here, and finally a likelihood function value (equivalent to the above-mentioned matching degree) is calculated. Specifically, in this task, the calculation goal is that for the expected speech prediction result, the likelihood function can estimate the speech feature sequence combination (equivalent to the above-mentioned target music content) when the maximum probability of obtaining this result can be achieved.

[0279] The beneficial effects of this solution include: starting from multi-modal learning, realizing end-to-end recommendation based on a large model, and can be applied to the in-vehicle music recommendation scenario.

[0280] In the second aspect, an embodiment of the present application provides a music recommendation device, as Figure 15 shown, the music recommendation device 150 includes: an acquisition unit 1501, a fusion unit 1502, a generation unit 1503, a determination unit 1504, and an output unit 1505. Among them:

[0281] The acquisition unit 1501 is used to acquire at least one type of modal feature for music recommendation for the target user; each type of modal feature in the at least one type of modal feature is obtained by extracting features from a type of modal data;

[0282] The fusion unit 1502 is used to perform feature fusion on the at least one type of modal feature to obtain a fused modal feature;

[0283] The generation unit 1503 is used to generate a target audio sequence for characterizing the music recommendation requirement of the target user based on the fused modal feature;

[0284] A determining unit 1504, configured to determine at least one target music content based on at least the target audio features of the target audio sequence, so as to form a music recommendation list;

[0285] An output unit 1505, configured to output the music recommendation list.

[0286] In some embodiments, the generating unit 1503 is specifically configured to: process the fusion modality features by using a first audio model to predict the music recommendation type of the target user; the music recommendation type includes at least one of the following: melody type, rhythm type, emotion type, and music style type; the first audio model belongs to an inference-based large model; process the music recommendation type by using a second audio model to generate a target audio sequence that conforms to the music recommendation type; the second audio model belongs to a generative large model.

[0287] In some embodiments, the determining unit 1504 is specifically configured to: determine at least one candidate music content in the music library based on the target audio features; or, determine the at least one candidate music content in the music library based on the target audio features and the target audio change features of the target audio sequence; determine the at least one target music content from the at least one candidate music content; sort the at least one target music content to obtain the music recommendation list.

[0288] In some embodiments, the determining unit 1504 is further configured to: for each of at least one first music content included in the music library, perform the following processing to obtain the at least one candidate music content: determine the audio feature matching degree between the target audio features and the first audio features of the first music content, and the audio change feature matching degree between the target audio change features and the first audio change features of the first music content; in a case where the audio feature matching degree is greater than a first threshold and the audio change feature matching degree is greater than a second threshold, determine the first music content as one of the candidate music contents.

[0289] In some embodiments, when the target audio feature includes a target rhythm feature, a target melody feature, and a target timbre feature, and the first audio feature includes a first rhythm feature, a first melody feature, and a first timbre feature, the determining unit 1504 is further configured to: respectively determine the rhythm matching degree between the target rhythm feature and the first rhythm feature, the melody matching degree between the target melody feature and the first melody feature, and the timbre matching degree between the target timbre feature and the first timbre feature; respectively determine a first weight corresponding to the rhythm matching degree, a second weight corresponding to the melody matching degree, and a third weight corresponding to the timbre matching degree; based on the first weight, the second weight, and the third weight, perform a weighted average on the rhythm matching degree, the melody matching degree, and the timbre matching degree to obtain the audio feature matching degree.

[0290] In some embodiments, the determining unit 1504 is further configured to: determine a representative identifier for each of the at least one candidate music content; the representative identifier is used to represent whether the candidate music content is representative in the music type to which it belongs; determine, as the at least one target music content, the candidate music content among the at least one candidate music content whose representative identifier takes a first value; the first value represents that the candidate music content is representative in the music type to which it belongs.

[0291] In some embodiments, the determining unit 1504 is further configured to: determine a timeliness score for each of the at least one target music content; the timeliness score is used to represent the closeness between the target music content and the current hot event; sort the at least one target music content in descending order of the timeliness score to obtain the music recommendation list.

[0292] In some embodiments, the determining unit 1504 is further configured to: for each of the at least one target music content, perform the following processing: obtain the release time of the target music content and the relevance between the target music content and the current hot event; based on the target music type of the target music content, determine the target timeliness function corresponding to the target music content; based on the release time, the relevance, and the target timeliness function, determine the timeliness score of the target music content.

[0293] In some embodiments, the music recommendation device 150 further includes an adjustment unit. The adjustment unit is configured to: obtain environmental data related to acoustic characteristics in the current environment of the target user; generate an environmental audio corresponding to the environmental data based on the environmental data; add the environmental audio to the target audio sequence to obtain the adjusted target audio sequence.

[0294] It should be noted that each unit included in the music recommendation device provided in the embodiments of the present application can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; during the implementation process, the processor can be a central processing unit (CPU, Central Processing Unit), a microprocessor (MPU, Micro Processor Unit), a digital signal processor (DSP, Digital Signal Processor), or a field-programmable gate array (FPGA, Field-Programmable Gate Array), etc.

[0295] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0296] It should be noted that in the embodiments of the present application, if the above music recommendation method is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read Only Memory), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0297] In a third aspect, the embodiments of the present application provide an electronic device, which at least includes a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements the steps in the music recommendation method provided in the above embodiments.

[0298] Next, in combination with Figure 16 the electronic device 160 shown, the structural diagram of the electronic device will be described.

[0299] In one example, as Figure 16As shown, the electronic device 160 includes: a processor 1601, at least one communication bus 1602, at least one external communication interface 1603, and a memory 1604. Among them, the communication bus 1602 is configured to enable connection communication between these components. Among them, the external communication interface 1603 may include a standard wired interface and a wireless interface.

[0300] The memory 1604 is configured to store instructions and applications executable by the processor 1601, and can also cache data to be processed or already processed by the processor 1601 and each module in the electronic device (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0301] In another example, the electronic device may be a terminal device (vehicle device or mobile phone) or a controller. The terminal device or the controller is used to execute the steps in the music recommendation method provided in the above embodiments.

[0302] Fourthly, an embodiment of the present application also provides a storage medium, that is, a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the music recommendation method provided in the above embodiments are implemented.

[0303] Fifthly, an embodiment of the present application also provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps in the music recommendation method provided in the above embodiments are implemented.

[0304] It should be noted here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0305] It should be understood that "an embodiment" or "one embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in an embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0306] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0307] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0308] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0309] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately a unit alone, or two or more units can be integrated in a unit; the above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0310] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks or optical discs that can store program codes.

[0311] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as removable storage devices, ROMs, magnetic disks, or optical discs that can store program codes.

[0312] As described above, the above are only the implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A music recommendation method, characterized in that: The music recommendation method comprises: Acquire at least one type of modal features for recommending music to a target user; each type of modal features in the at least one type of modal features is obtained by extracting features from a type of modal data; Performing feature fusion on the at least one type of modal features to obtain fused modal features; Based on the fused modal features, generating a target audio sequence for representing the music recommendation needs of the target user; determining at least one target music content based at least on the target audio feature of the target audio sequence to form a music recommendation list; Outputting the music recommendation list; The step of generating a target audio sequence for representing the music recommendation needs of the target user based on the fused modal features includes: The fused modal features are processed using a first audio model to predict the music recommendation type of the target user; the music recommendation type includes at least one of the following: melody type, rhythm type, emotion type and music style type; the first audio model is an inference-based large model; The music recommendation type is processed using a second audio model to generate a target audio sequence that matches the music recommendation type; the second audio model belongs to a generative large model.

2. The music recommendation method according to claim 1, characterized in that: The step of determining at least one target music content based at least on the target audio feature of the target audio sequence to form a music recommendation list comprises: Based on the target audio feature, determining at least one candidate music content in the music library; or based on the target audio feature and the target audio variation feature of the target audio sequence, determining the at least one candidate music content in the music library; determining the at least one target music content among the at least one candidate music content; The at least one target music content is sorted to obtain the music recommendation list.

3. The music recommendation method according to claim 2, characterized in that: The determining, in the music library, the at least one candidate music content based on the target audio feature and the target audio variation feature of the target audio sequence comprises: For each of the at least one first music content included in the music library, the following processing is performed to obtain the at least one candidate music content: Determining an audio feature matching degree between the target audio feature and a first audio feature of the first music content, and an audio change feature matching degree between the target audio change feature and a first audio change feature of the first music content; When the audio feature matching degree is greater than a first threshold and the audio change feature matching degree is greater than a second threshold, the first music content is determined as one of the candidate music contents.

4. The music recommendation method according to claim 3, characterized in that: In a case where the target audio feature includes a target rhythm feature, a target melody feature, and a target timbre feature, and the first audio feature includes a first rhythm feature, a first melody feature, and a first timbre feature, determining an audio feature matching degree between the target audio feature and the first audio feature of the first music content includes: respectively determining a rhythm matching degree between the target rhythm feature and the first rhythm feature, a melody matching degree between the target melody feature and the first melody feature, and a timbre matching degree between the target timbre feature and the first timbre feature; respectively determining a first weight corresponding to the rhythm matching degree, a second weight corresponding to the melody matching degree, and a third weight corresponding to the timbre matching degree; Based on the first weight, the second weight and the third weight, the rhythm matching degree, the melody matching degree and the timbre matching degree are weighted averaged to obtain the audio feature matching degree.

5. The music recommendation method according to claim 2, characterized in that: The determining the at least one target music content from the at least one candidate music content comprises: Determine a representative identifier of each of the at least one candidate music content; the representative identifier is used to indicate whether the candidate music content is representative in the music genre to which it belongs; Among the at least one candidate music content, a candidate music content that is representatively identified as a first value is determined as the at least one target music content; the first value indicates that the candidate music content is representative in the corresponding music type.

6. The music recommendation method according to claim 2, characterized in that: The step of sorting the at least one target music content to obtain the music recommendation list includes: Determine the timeliness score of each of the at least one target music content; the timeliness score is used to characterize the closeness between the target music content and the current hot event; The at least one target music content is sorted in descending order according to the timeliness score to obtain the music recommendation list.

7. The music recommendation method according to claim 6, characterized in that: The determining of the timeliness score of each of the at least one target music content comprises: For each of the at least one target music content, the following processing is performed: Acquire the release time of the target music content and the relevance of the target music content to the current hot event; Determining a target time-effect function corresponding to the target music content based on a target music type of the target music content; Based on the release time, the relevance, and the target timeliness function, a timeliness score of the target music content is determined.

8. The music recommendation method according to claim 1, characterized in that: After generating a target audio sequence for representing the music recommendation demand of the target user based on the fused modal features, the method further includes: Acquiring environmental data related to acoustic characteristics of the target user's current environment; Based on the environmental data, generating environmental audio corresponding to the environmental data; The environmental audio is added to the target audio sequence to obtain the adjusted target audio sequence.

9. An electronic device, characterized in that: The electronic device comprises at least a memory and a processor, the memory stores a computer program executable on the processor, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Personalized music recommendation method and device, equipment, medium and wearable equipment

    CN118013071A