Intelligent audio player based on AI and control method thereof

By collecting multi-dimensional data and sensing the environment in real time, combined with machine learning and deep learning, the intelligent audio player solves the problem of traditional players' inability to adapt, realizes personalized recommendations and intelligent interaction, and improves the accuracy and stability of the user experience.

CN121597159APending Publication Date: 2026-03-03SHENZHEN JINRUI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511974237.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Traditional audio players cannot adapt to users' habits, emotional state, and context, resulting in cumbersome operation, difficulty in quickly and accurately locating desired content, and a lack of intelligent interaction capabilities.

Method used

By collecting multi-dimensional data, analyzing machine learning algorithms and deep learning, and adjusting parameters in real time using environmental sensors, a personalized song recommendation list is generated. User feedback is collected in real time to optimize the model, and early warning indicators are set to provide intelligent interactive suggestions.

Benefits of technology

It improves the accuracy and personalization of recommendations, enhances the flexibility of playback strategies and the stability of user experience, and reduces the degradation of experience caused by environmental changes or recommendation failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597159A_ABST
    Figure CN121597159A_ABST
Patent Text Reader

Abstract

The invention provides an AI-based intelligent audio player and a control method thereof. Belongs to the technical field of artificial intelligence and music players. The method comprises the steps of performing multi-dimensional data collection on a user of the intelligent audio player, and generating a user music preference comprehensive data set; based on the comprehensive data set, constructing an initial user music preference analysis model through a machine learning algorithm; a plurality of environment sensors are integrated on the intelligent audio player, data of a playing environment are collected in real time through the environment sensors, and environment real-time data are generated. Through multi-dimensional data collection and machine learning algorithm analysis, in combination with the song listening history, collection preference and active feedback of the user, the intelligent audio player can accurately generate a music preference comprehensive data set of the user and construct an initial personalized music recommendation model, so that the deviation between recommended songs and user preferences is reduced, and the user experience is improved. And the recommendation accuracy and personalized experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes an AI-based intelligent audio player and its control method, belonging to the fields of artificial intelligence and music player technology. Background Technology

[0002] With the rapid development of digital audio technology, traditional audio players can no longer meet the increasingly diverse and personalized needs of users. Traditional players mainly rely on manual operation, requiring users to search and select audio content one by one, which is cumbersome and inefficient. Especially when faced with a massive amount of audio resources, users find it difficult to quickly and accurately locate their desired content.

[0003] Meanwhile, traditional media players lack intelligent interaction capabilities and cannot adaptively adjust based on user habits, emotional state, and context. For example, users have different needs for audio type and volume at different times (such as working, resting, and exercising), but traditional media players cannot automatically adapt to these changes.

[0004] In recent years, artificial intelligence (AI) technology has achieved significant breakthroughs, with technologies such as speech recognition, natural language processing, and machine learning being widely applied in numerous fields. Integrating AI technology into audio player control enables functions such as voice command operation, intelligent content recommendation, and scene-adaptive playback. By analyzing users' historical playback data and favorite preferences, AI can accurately predict user needs and provide personalized audio services. Summary of the Invention

[0005] This invention provides an AI-based intelligent audio player and its control method to solve the problems mentioned in the background section above:

[0006] The present invention proposes a control method for an AI-based intelligent audio player, the method comprising:

[0007] S1. Collect multi-dimensional data from users of smart audio players to generate a comprehensive dataset of user music preferences; based on the comprehensive dataset, construct an initial user music preference analysis model using machine learning algorithms;

[0008] S2. Integrate multiple environmental sensors on the smart audio player, collect data of the playback environment in real time through the environmental sensors, and generate real-time environmental data; based on the real-time environmental data, combined with pre-set rules and algorithms, adaptively adjust the audio player parameters to generate adjusted playback parameter data;

[0009] S3. Input the comprehensive dataset of user music preferences and the adjusted playback parameter data into the AI ​​analysis module based on deep learning. Through matching and analysis of massive music data, generate a personalized song recommendation list for each user and obtain personalized recommendation result data.

[0010] S4. During the process of a user playing recommended songs using a smart audio player, real-time feedback data on the playback effect is collected. The feedback data is correlated with the personalized recommendation result data to evaluate the accuracy of the recommended songs and user satisfaction. Based on the evaluation results, the initially constructed user music preference analysis model is dynamically optimized and adjusted to generate an optimized music preference analysis model.

[0011] S5. Based on the optimized music preference analysis model and real-time collected environmental and user feedback data, a series of early warning indicators are set; relevant suggestions are provided to users through an intelligent interactive interface, while generating intelligent interactive prompt data; based on the user's response, the playback strategy is further adjusted, and finally, complete intelligent audio playback control scheme data is generated.

[0012] The AI-based intelligent audio player proposed in this invention includes:

[0013] One or more processors;

[0014] Memory, used to store one or more programs.

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the control method described in any one of the above descriptions.

[0016] The beneficial effects of this invention are as follows: By collecting multi-dimensional data and analyzing machine learning algorithms, combined with users' listening history, collection preferences, and active feedback, the intelligent audio player can accurately generate a comprehensive dataset of users' music preferences, constructing an initial personalized music recommendation model. This reduces the deviation between recommended songs and user preferences, improving the accuracy and personalized experience of recommendations. Furthermore, leveraging deep learning algorithms, the intelligent audio player can deeply mine users' potential music preferences and the changing patterns of environmental parameters. Through matching and analyzing massive amounts of music data, it automatically generates personalized recommendation lists, not only improving the accuracy of recommendations but also predicting the song styles, rhythms, and other characteristics that users might like based on user behavior and environmental data, ensuring accurate recommendations. The intelligent audio player ensures the timeliness and personalization of recommendation results. By collecting user feedback on recommended songs through a real-time feedback mechanism, it can dynamically optimize the user music preference analysis model based on feedback data and adjust the recommendation algorithm in real time. This avoids reduced recommendation accuracy due to model lag and enhances the adaptability and self-adjustment capabilities of the personalized recommendation system. By setting early warning indicators and combining user feedback and environmental changes, the intelligent audio player can promptly identify and respond to potential problems (such as frequent song skipping or sudden environmental changes), provide relevant suggestions to users through an intelligent interactive interface, and further adjust the playback strategy based on user responses. This reduces user churn caused by unsuitable playback experiences and improves the quality and satisfaction of the user experience. Attached Figure Description

[0017] Figure 1 This is a diagram illustrating the steps of the method described in this invention. Detailed Implementation

[0018] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0019] One embodiment of the present invention, such as Figure 1 As shown, an AI-based intelligent audio player control method includes:

[0020] S1. Collect multi-dimensional data from users of the smart audio player. The multi-dimensional data includes users' listening history, including the type, duration, and frequency of songs played in different time periods and scenarios; users' favorite preference data, such as favorite songs, playlists, and artists; and user-initiated feedback on preferences for song style, rhythm, and artist timbre, etc., to generate a comprehensive user music preference dataset. Based on the comprehensive dataset, construct an initial user music preference analysis model using machine learning algorithms. The music preference analysis model serves as the basic framework for subsequent intelligent control.

[0021] S2. Integrate multiple environmental sensors into the smart audio player, including a light sensor, a sound sensor, and a temperature sensor; collect real-time data of the playback environment through these sensors, including ambient light intensity, ambient noise level, ambient temperature, etc., to generate real-time environmental data; based on the real-time environmental data, and combined with pre-set rules and algorithms, adaptively adjust the audio player parameters (volume, sound effect mode, etc.); for example, automatically increase the volume when the ambient noise is high; switch to a softer sound effect mode in low-light scenes, and generate adjusted playback parameter data;

[0022] S3. Input the comprehensive dataset of user music preferences and the adjusted playback parameter data into the AI ​​analysis module based on deep learning. The AI ​​analysis module uses complex neural network algorithms to deeply explore the user's potential music preference patterns and the changes in the user's music preferences under different environmental parameters. Through matching and analyzing massive amounts of music data, a personalized song recommendation list is generated for each user. At the same time, it predicts the song styles, rhythms and other features that the user may like under the current environment and preference state, and obtains personalized recommendation result data.

[0023] S4. During the process of a user playing recommended songs using a smart audio player, real-time feedback data on the playback effect is collected. The feedback data includes whether the user skipped the song, the number of times the song was repeated, and the user's rating of the song. The feedback data is correlated with the personalized recommendation result data to evaluate the accuracy of the recommended songs and user satisfaction. Based on the evaluation results, the initially constructed user music preference analysis model is dynamically optimized and adjusted, the model parameters are continuously corrected, the model's prediction accuracy of user music preferences is improved, and an optimized music preference analysis model is generated.

[0024] S5. Based on the optimized music preference analysis model and real-time collected environmental and user feedback data, a series of early warning indicators are set. For example, when it is detected that a user skips recommended songs multiple times in a row, or when environmental parameters change drastically and may affect the playback experience, an early warning mechanism is triggered. Through an intelligent interactive interface, relevant suggestions are provided to the user in the form of voice prompts, pop-up notifications, etc. The relevant suggestions include changing the song style and adjusting the playback environment, while generating intelligent interactive prompt data. Based on the user's response, the playback strategy is further adjusted to ensure that the user is provided with a playback experience that always meets their needs, and finally, complete intelligent audio playback control scheme data is generated.

[0025] The working principle and effects of the above technical solution are as follows:

[0026] By collecting user listening history, collection preferences, and active feedback from multiple dimensions, and combining this with machine learning to build an initial preference model, we reduce the bias of recommendations relying solely on single data points, avoid recommendation biases caused by insufficient information, improve the relevance of initial recommendations, and enhance the model's ability to capture users' basic preferences. Real-time collection of ambient light, noise, and temperature data automatically adjusts parameters such as volume and sound effects, reducing the manual adjustment costs for users, preventing experience degradation caused by untimely environmental changes, improving the real-time performance of parameter adjustments, and enhancing the adaptability of playback effects to the environment. Deep learning is used to uncover latent user preferences and patterns of environmental influence, matching massive amounts of music for generation. Personalized recommendations reduce the problem of recommendations being limited to explicit preferences, avoid ignoring the dynamic influence of the environment on preferences, improve recommendation accuracy, and enhance user acceptance of recommended content. Real-time collection and correlation analysis of user feedback dynamically optimizes the preference model, reducing the likelihood of the model becoming ineffective due to changes in user preferences, preventing long-term decline in recommendation accuracy, improving model adaptability, and enhancing the sustainability of recommendation effects. Setting early warning indicators and triggering intelligent interaction suggestions, adjusting playback strategies based on user responses, reduces the problem of untimely handling of sudden environmental changes or recommendation failures, prevents sudden deterioration of the experience, improves the flexibility of playback strategies, and enhances the stability of the user experience.

[0027] In one embodiment of the present invention, S1 includes:

[0028] S11: Construct a multi-dimensional user music data collection framework and clarify the dimensional boundaries and indicator system of data collection. The dimensions include time dimension (e.g., weekdays / weekends, morning, noon and evening), scenario dimension (e.g., commuting, exercise, rest), behavior dimension (e.g., play / pause, skip songs, favorites) and feedback dimension (e.g., star rating, text review).

[0029] S12: Based on the acquisition framework, user raw data is collected through the local storage module of the smart audio player and the cloud synchronization interface. The user raw data includes playback records for different time periods, song type selections in various scenarios, update records of favorite playlists, and active feedback text evaluations, generating a user music behavior raw dataset.

[0030] S13: Clean and standardize the original dataset of user music behavior, remove abnormal data (such as second-level playback records caused by accidental operation), unify the data format (such as standardizing different scene labels into preset scene categories), and integrate multi-source data through data fusion technology to generate a comprehensive dataset of user music preferences.

[0031] S14: Based on a comprehensive dataset of user music preferences, an initial user music preference analysis model is constructed using machine learning algorithms. The model includes user basic preference feature vectors (such as the weight of preferred music genres and the distribution of commonly played times), which serve as the basic framework for subsequent intelligent control.

[0032] The working principle and effects of the above technical solution are as follows:

[0033] By constructing a multi-dimensional data collection framework and clearly defining dimensional boundaries, the aim was to avoid the blindness of data collection, reduce the omission of key dimensions, and improve the comprehensiveness of data collection, enabling subsequent analysis to cover multiple aspects of user music behavior. Based on this framework, raw data was collected, and combined with local storage and cloud synchronization, reducing the problems of missing or fragmented data, avoiding analytical biases caused by incomplete information, ensuring the integrity of the raw user music behavior data, and providing sufficient material for subsequent processing. Cleaning abnormal data, standardizing formats, and integrating multi-source information reduced the interference of erroneous records and formatting issues on the analysis, reduced model errors caused by data noise, improved the quality of the comprehensive user music preference dataset, and made the data more aligned with actual analytical needs. Using machine learning to construct an initial model containing basic preference vectors avoided the subjectivity of relying solely on manual judgment of user preferences, reduced the blindness of initial model design, enhanced the model's ability to capture basic user music preferences, and laid a reliable foundation for subsequent intelligent control.

[0034] In one embodiment of the present invention, S2 includes:

[0035] S21: Based on the usage scenario requirements of the smart audio player, select an appropriate combination of environmental sensors, including a high-precision light sensor (detecting light intensity from 0-10000 lux), a wideband sound sensor (collecting ambient noise from 30-18000 Hz), and a temperature and humidity sensor (monitoring ambient temperature from 10-40℃), and integrate the sensors into the reserved interface on the player's casing.

[0036] S22: Real-time environmental data is collected at a frequency of 100ms / time through the sensor data acquisition interface. After removing data noise through Kalman filtering, a real-time environmental data stream is generated, which includes light intensity value, noise decibels, and ambient temperature value.

[0037] S23: Construct an environment parameter and playback parameter mapping rule library, the rules include basic rules (e.g., volume increase by 15% when noise ≥ 60dB) and scenario-based rules (e.g., cooling sound effect mode is enabled by default when temperature ≥ 30℃), and store the rule library in the player's local algorithm module;

[0038] S24: Input the real-time environmental data stream into the rule base for matching calculation, adaptively adjust the core parameters of the audio player (volume, equalizer curve, surround sound mode, etc.), and generate adjusted playback parameter data.

[0039] The working principle and effects of the above technical solution are as follows:

[0040] By selecting and integrating suitable high-precision sensors, environmental data acquisition deviations caused by insufficient sensor performance or incompatible types are avoided. This reduces misjudgments of information such as light, noise, temperature, and humidity, improving the accuracy of environmental perception and enabling the player to accurately capture the surrounding environment. High-frequency data acquisition and filtering reduce interference signals in the environmental data (such as momentary noise or light flicker), minimizing the impact of data fluctuations on subsequent analysis and improving the stability of real-time environmental data streams, ensuring that the input environmental information is more realistic. A mapping rule library combining basic and scenario-based approaches is built and stored locally, avoiding blind parameter adjustments, reducing confusion in playback parameter settings under different environments, improving the efficiency of rule calls, and providing a clear basis for parameter adjustments. Adaptive adjustment of core parameters through rule matching reduces the hassle of manual volume and sound effect adjustments for users, preventing poor playback experience due to untimely environmental changes (such as unclear hearing in noisy environments or muffled sound in high temperatures), improving the adaptability of playback parameters to the environment, and enhancing ease of use and user comfort.

[0041] In one embodiment of the present invention, S24 includes:

[0042] Key feature values ​​are extracted from the real-time environmental data stream. These key feature values ​​include the range of light intensity (e.g., 0-500 lux for low light, 500-5000 lux for medium light, and 5000-10000 lux for high light), the level of noise in decibels (e.g., 30-50 dB for low noise, 50-70 dB for medium noise, and above 70 dB for high noise), and the temperature range (e.g., 10-20℃ for low temperature, 20-30℃ for normal temperature, and 30-40℃ for high temperature). Environmental feature identification data is then generated.

[0043] The environmental feature identification data is matched with the environmental parameter and playback parameter mapping rule library to filter out the suitable basic rules and scenario-based rules (such as volume adjustment rules for medium noise and standard sound effect rules for normal temperature) and generate a rule matching result set.

[0044] Based on the rule matching result set, the specific adjustment amount of each core parameter is calculated (e.g., volume is increased by 10% when there is medium noise, and surround sound intensity is increased by 20% when there is high brightness). The rule conflict is eliminated through the parameter coordination algorithm (e.g., when both high noise and high temperature are met, volume adjustment is performed first and then sound effects are adapted), and a parameter adjustment instruction set is generated.

[0045] The parameter adjustment instruction set is loaded into the control module of the audio player to complete the adaptive adjustment of core parameters (volume, equalizer curve, surround sound mode, etc.) and generate playback parameter data containing the specific parameter values ​​after adjustment.

[0046] The working principle and effects of the above technical solution are as follows:

[0047] Extracting key environmental features and classifying them reduces the clutter of raw environmental data, avoids difficulties in rule matching caused by fragmented data, improves the efficiency of matching environmental information with the rule base, and makes subsequent matching more accurate. Matching feature identifiers with the rule base to filter suitable rules avoids interference from irrelevant rules, reduces the blindness of rule matching, improves the targeting of rule selection, and provides a clear basis for parameter adjustment. Calculating specific adjustment amounts and eliminating rule conflicts through algorithms reduces parameter contradictions when multiple rules are triggered simultaneously, reduces adjustment chaos, improves the rationality of parameter coordination, and makes the adjustment of different parameters more coordinated. Loading instruction sets to complete adaptive parameter adjustment reduces the trouble of manual adjustment by users, avoids playback discomfort caused by untimely adaptation to environmental changes (such as insufficient volume in noisy environments or abrupt sound effects in bright light), improves the real-time performance and accuracy of parameter adjustment, and enhances the fit between playback effect and environment.

[0048] In one embodiment of the present invention, S3 includes:

[0049] S31: Build an AI analysis module based on the Transformer architecture. The analysis module includes a feature extraction layer (extracting high-dimensional features of user preferences and environmental parameters), an interaction layer (modeling the dynamic relationship between user preferences and the environment), and an output layer (generating a recommendation probability distribution), and is deployed on the edge computing unit of the player.

[0050] S32: The user music preference dataset and the adjusted playback parameter data are fused to generate an input feature matrix. The feature matrix includes user historical preference features (e.g., the features of the top 10 songs played in the last 30 days) and real-time environmental features (e.g., the feature vector corresponding to the current noise level).

[0051] S33: The AI ​​analysis module performs in-depth mining of the input feature matrix to identify potential user preference patterns (e.g., a preference for playing slow-paced songs on rainy days) and environmental influence patterns (e.g., a high requirement for clear lyrics during commuting), generating a user and environment-related feature map.

[0052] S34: Based on the user and environment association feature map, perform similarity matching and priority sorting in the music resource library to generate a personalized recommendation list containing 20 candidate songs, and predict the user acceptance probability of each song to obtain personalized recommendation result data.

[0053] The working principle and effects of the above technical solution are as follows:

[0054] A Transformer-based AI analysis module was built and deployed on the edge computing unit, avoiding latency caused by cloud reliance on complex analyses, reducing data transmission losses, and improving feature processing efficiency, enabling rapid local deep analysis. By fusing user preferences with playback parameter features to generate a matrix, the limitations of relying on a single feature are avoided, reducing the disconnect between user history and real-time environment, improving the completeness of input features, and allowing analysis to consider both user habits and current state. The AI ​​module mines potential preference patterns and environmental influences, avoiding the limitations of only capturing surface preferences, reducing the neglect of implicit user needs, and enhancing the depth of understanding of the relationship between users and the environment, making recommendations more aligned with real needs. Based on association graph matching and ranking, and predicting acceptance probability, inappropriate content caused by blind recommendations is avoided, reducing the inconvenience of frequent song switching for users, improving the accuracy of the recommendation list, enhancing user acceptance of recommended content, and making each recommendation more relevant.

[0055] In one embodiment of the present invention, S32 includes:

[0056] Core historical preference features are extracted from the comprehensive dataset of user music preferences, including style tag weights of frequently played songs in the past 30 days, the theme distribution of favorite playlists, and the proportion of playback time in different time periods. Through feature normalization processing (e.g., converting the proportion to a 0-1 range value), a user historical preference feature vector is generated.

[0057] Based on the adjusted playback parameter data, the real-time environment features are inferred, including the noise features corresponding to the current volume, the environment type matched by the sound effect mode (e.g., soothing mode corresponds to a quiet environment), and the scene requirements reflected by the equalizer curve (e.g., bass enhancement corresponds to a sports scene). After feature encoding, a real-time environment feature vector is generated.

[0058] The user's historical preference feature vector and the real-time environment feature vector are dimensionally aligned. The fusion weight (e.g., historical preference weight 0.6, real-time environment weight 0.4) is determined by a feature association algorithm (e.g., calculating the mutual information value of the two types of features) to eliminate feature dimension conflicts.

[0059] The aligned feature vectors are concatenated into a matrix according to the fusion weights to form an input feature matrix containing user and environmental interaction information. The row dimension of the matrix corresponds to the feature category, and the column dimension corresponds to the feature quantization value.

[0060] The working principle and effects of the above technical solution are as follows:

[0061] By extracting and normalizing core historical features from user preference data, the system reduces format differences between different data types (such as playback duration and style tags), avoids feature distortion caused by inconsistent numerical ranges, improves the consistency of historical preference features, and allows users' long-term habits to be captured more clearly. By inferring and encoding environmental correlation features based on playback parameters, the system avoids the one-sidedness of directly collecting environmental information, reduces over-reliance on sensor data, enhances the correlation between environmental features and playback status, and makes the impact of the real-time environment on playback easier to analyze. Aligning feature dimensions and determining fusion weights through algorithms reduces dimensional conflicts between historical user features and real-time environmental features, reduces information loss when splicing the two types of features, improves the rationality of feature fusion, and allows features from different sources to work synergistically. By splicing feature vectors according to weights to form an interaction matrix, the system avoids the isolated existence of user preferences and environmental features, reduces the problem of ignoring their correlation during analysis, enhances the completeness of input features, and allows the AI ​​module to more comprehensively uncover users' real needs in specific environments.

[0062] In one embodiment of the present invention, S33 includes:

[0063] The input feature matrix is ​​fed into the feature extraction layer of the AI ​​analysis module. The multilayer perceptron performs high-dimensional mapping between the user's historical preference features and real-time environmental features in the matrix, extracting fine-grained features (such as the implicit correlation between song rhythm and environmental noise, and the coupling features between playback time and temperature), and generating a deeper feature set.

[0064] Based on the deepened feature set, the attention mechanism of the interaction layer is used to calculate the association weight between user preference features and environmental features (for example, the association weight between rhythm intensity features and noise level in a sports scene is 0.7), identify potential user preference patterns (for example, users prefer high-tempo songs during weekday morning rush hour), and generate a preference pattern set.

[0065] By using dynamic modeling algorithms in the interaction layer, we analyze the impact of changes in different environmental parameters on preference patterns (for example, for every 10dB increase in environmental noise, the user's demand weight for "lyric clarity" increases by 0.2), and generate a set of environmental impact patterns.

[0066] By associating and integrating the preference pattern set with the environmental impact pattern set, and using nodes to represent features (such as rainy weather environment, soothing rhythm) and edges to represent the association strength, a structured user and environment association feature map is constructed.

[0067] The working principle and effects of the above technical solution are as follows:

[0068] By using a multilayer perceptron to perform high-dimensional mapping of features and extract fine-grained features, the limitations of only capturing surface features are avoided. This reduces the omission of implicit correlations such as song rhythm and environmental noise, playback time and temperature, and improves the depth of feature mining, allowing subsequent analysis to reach users' unmanifested potential needs. Attention mechanisms are used to calculate association weights and identify potential preference patterns, avoiding the one-sidedness of judging preferences solely based on explicit listening behavior. This reduces misjudgments of users' actual habits (e.g., listening to high-tempo songs during weekday morning rush hour), enhances the accuracy of preference pattern recognition, and makes the mined patterns more aligned with users' actual usage scenarios. Dynamic modeling algorithms are used to analyze the influence of the environment on preferences, avoiding the problem of ignoring changes in environmental parameters. This reduces the bias of recommendations that do not consider dynamic influences such as noise and temperature, improves the practicality of the patterns, and allows subsequent recommendations to flexibly adjust to environmental changes. Integrating preference patterns and environmental patterns to generate a structured graph avoids scattered and disorganized information, reduces the cost of information search and sorting during subsequent matching, enhances the intuitiveness of feature associations, and makes the AI ​​module more efficient and accurate when using this association information.

[0069] In one embodiment of the present invention, S34 includes:

[0070] Features are extracted from songs in the music resource library, including genre tags (e.g., pop, rock), rhythm BPM value, vocal clarity parameters, instrument composition, etc., and a standardized song feature library is generated through feature standardization (e.g., normalizing BPM values ​​to the 0-1 range).

[0071] The cosine similarity algorithm is used to calculate the matching degree between the core features in the user and environment association feature map (such as commuting environment + high rhythm preference) and the features of each song in the standardized song feature library. Songs with a matching degree higher than a preset threshold (such as 0.6) are selected to generate a preliminary candidate song set.

[0072] By combining the weight of users’ historical interactions (such as the completion rate of similar songs in the past 30 days) and the environmental adaptability coefficient (such as the requirement coefficient of vocal clarity in noisy environments), the initial candidate song set is prioritized and the top 20 songs are selected to generate an ordered candidate recommendation list.

[0073] The pre-trained user acceptance probability prediction model is invoked. The user and environment-related features of each song in the ordered candidate recommendation list are input, and the user acceptance probability of each song (e.g., 85%, 62%) is output. The ordered candidate recommendation list and the acceptance probability are integrated to generate personalized recommendation result data.

[0074] The working principle and effects of the above technical solution are as follows:

[0075] Multi-dimensional features are extracted from songs and standardized to generate a feature library. This reduces format differences in features (such as BPM values ​​and clarity parameters) across different songs, avoids matching bias caused by inconsistent numerical ranges, and improves feature consistency, laying the foundation for accurate matching in the future. Cosine similarity is used to calculate the matching degree between core features and songs and to filter them, avoiding the problem of blindly selecting songs, reducing the probability of low-fit songs entering the candidate set, improving the accuracy of the initial screening, and making the candidate songs more suitable for the user's current environment and preferences. The top 20 songs are selected by combining historical interaction weights and environmental fit coefficients, avoiding the limitation of only looking at the matching degree and ignoring the user's past reactions and environmental needs. This reduces the situation of unreasonable ranking (such as highly matched songs that the user has skipped being ranked first), and enhances the practicality of the recommendation list. The probability of acceptance is predicted and integrated into the generated results, avoiding the problem of providing a list without reference, reducing the risk of low user acceptance after recommendation, improving the reliability of the recommendation results, and providing a more accurate basis for subsequent model optimization.

[0076] In one embodiment of the present invention, step S4 includes:

[0077] S41: Design a multimodal feedback acquisition mechanism to capture user feedback signals on recommended songs in real time through the player's touch interface (e.g., swiping to rate), voice interaction (e.g., disliking a song), and behavior tracking (e.g., the interval between consecutively skipped songs);

[0078] S42: Quantify the feedback signals, convert subjective feedback (e.g., five-star rating) into a quantitative score of 0-10, and convert behavioral feedback (e.g., number of repeated playbacks) into weight coefficients (e.g., weight +0.2 for playbacks more than 3 times), and generate a quantitative dataset of user feedback.

[0079] S43: Perform correlation analysis between the quantitative dataset of user feedback and the personalized recommendation results data, calculate the recommendation accuracy (e.g., the percentage of recommended songs accepted by users) and satisfaction score (e.g., average rating), and generate a model evaluation report;

[0080] S44: Based on the model evaluation report, use online learning algorithms (such as incremental SVM) to correct the parameters of the initial user music preference analysis model, dynamically update the user preference feature vector (e.g., reduce the weight of song genres that are skipped multiple times), and generate an optimized music preference analysis model.

[0081] The working principle and effects of the above technical solution are as follows:

[0082] A multimodal feedback collection mechanism was designed, covering touch ratings, voice feedback, and behavior tracking. This reduces blind spots caused by single feedback methods, avoids situations where users want to express dissatisfaction but cannot easily operate the system, and improves the comprehensiveness of feedback collection, allowing users' true attitudes to be captured more completely. Feedback signals are quantified, converting subjective ratings and behavioral feedback into numerical values ​​or weights, reducing the problem of disorganized feedback data, avoiding the difficulty in quantifying subjective evaluations, improving the usability of feedback data, and providing well-organized material for subsequent correlation analysis. Correlation analysis of feedback data with recommendation results generates an evaluation report, clearly presenting recommendation accuracy and satisfaction, avoiding the blindness of judging recommendation effectiveness based solely on feelings, reducing the lack of basis for model optimization, improving the targeting of evaluation, and making the optimization direction clearer. Online learning algorithms are used to correct model parameters and update preference vectors, avoiding preference disconnect caused by a static model, reducing the frequent recommendation of skipped categories, improving model adaptability, enhancing the continuous accuracy of recommendation effects, and allowing the model to continuously adjust according to changes in user preferences.

[0083] In one embodiment of the present invention, S44 includes:

[0084] The key indicators (recommendation accuracy, satisfaction score) and problem attributions (e.g., high skip rate of songs in a certain genre, recommendation bias in a specific environment) in the model evaluation report are broken down to extract the core directions that need to be optimized (e.g., adjusting genre weights, correcting environment and preference correlation parameters) and generate a list of model optimization requirements.

[0085] Based on the optimization requirements list, incremental SVM was selected as the core online learning algorithm. The parameter correction range was defined (e.g., the weight range of school of thought in the user preference feature vector and the threshold of the environmental correlation coefficient), and a parameter adjustment execution plan was formulated.

[0086] The parameter adjustment execution plan is input into the initial user music preference analysis model. The model parameters are iteratively corrected through incremental learning algorithm (for example, the weight of genres that will be skipped multiple times is reduced from 0.3 to 0.1, and the correlation coefficient between commuting environment and rhythm preference is updated). The user preference feature vector is updated simultaneously to generate a temporary optimized model.

[0087] The effectiveness of the temporary optimization model is verified using quantitative user feedback data from the past 7 days. If the recommendation accuracy improves by ≥10% and the satisfaction score increases by ≥0.8 points, the model optimization is confirmed to be effective, and the optimized music preference analysis model is finally generated.

[0088] The working principle and effects of the above technical solution are as follows:

[0089] By breaking down the key indicators and identifying the root causes of problems in the evaluation report, we clarified the optimization direction and generated a list, avoiding the pitfalls of blindly adjusting model parameters, reducing ineffective optimizations caused by failing to grasp the core issues, and improving the targeting of model optimization so that adjustments are more aligned with actual needs. Selecting an incremental SVM algorithm and clarifying the parameter correction range and developing an execution plan avoided problems of poor algorithm adaptability or boundless parameter adjustments, reduced the risk of model crashes during optimization, improved the controllability of parameter adjustments, and made optimization operations more orderly. Iteratively correcting parameters and updating preference vectors to generate a temporary model avoided abrupt, one-size-fits-all parameter adjustments, reduced the situation where skipped categories were still recommended with high weights, improved the alignment of model parameters with current user preferences, and made the temporary model closer to users' true preferences. Using recent feedback data to verify the effectiveness of the temporary model and setting target thresholds avoided the problem of using the optimized model despite its unsatisfactory results, reduced the waste of optimization resources, improved the reliability of model optimization, and ensured that the final optimized model could effectively improve recommendation accuracy and user satisfaction.

[0090] In one embodiment of the present invention, step S5 includes:

[0091] S51: Construct a multi-dimensional early warning indicator system, the multi-dimensional indicators include recommendation effectiveness indicators (e.g., five consecutive songs are skipped), environmental adaptability indicators (e.g., environmental noise suddenly increases by more than 30dB), and user status indicators (e.g., detecting user emotional fluctuations through heart rate linkage data).

[0092] S52: The player's real-time monitoring module continuously monitors the warning indicators. When the indicators reach the threshold, the warning mechanism is triggered to generate warning signals (such as warnings of recommendation strategy failure or warnings of sudden environmental changes).

[0093] S53: Based on the warning signal, push targeted suggestions to the user through the intelligent interactive interface (such as voice assistant, pop-up window). The targeted suggestions include emergency adjustment suggestions (such as whether to turn on noise reduction effect in the current noisy environment) and long-term optimization suggestions (such as recommending new playlists of similar styles for you, whether to view them), and generate intelligent interactive prompt data;

[0094] S54: Collect user response data to intelligent interactive prompts (e.g., accepting / rejecting suggestions), combine it with the optimized music preference analysis model, dynamically adjust playback strategies (e.g., switching recommendation algorithms, updating environmental parameter mapping rules), and generate complete intelligent audio playback control scheme data covering data collection, environmental adaptation, recommendation generation, feedback optimization, and early warning adjustment.

[0095] The working principle and effects of the above technical solution are as follows:

[0096] A multi-dimensional early warning indicator system is constructed, covering dimensions such as recommendation effectiveness, environmental adaptability, and user status. This avoids blind spots in single-indicator monitoring, reduces experience degradation caused by missed critical issues (such as continuous song skipping or sudden noise increases), and improves the comprehensiveness of anomaly monitoring, allowing potential problems to be discovered in a timely manner. A real-time monitoring module continuously tracks indicators and triggers early warnings, preventing situations where no intervention occurs after anomalies, reducing the risk of problem escalation (such as not adjusting volume in time after a sudden noise increase), and improving the timeliness of anomaly response, ensuring problems are detected as soon as they appear. Based on early warnings, targeted suggestions for both urgent and long-term issues are pushed, avoiding the problem of vague and ineffective suggestions, reducing user confusion when facing anomalies (such as directly suggesting noise reduction in noisy environments), improving the practicality of suggestions, and enhancing user ease of operation. Collecting user responses to suggestions and adjusting playback strategies to form a complete control scheme avoids the problem of rigid and unchanging playback strategies, reduces the recurrence of similar anomalies, improves the closed-loop nature of the control scheme, and allows the player to continuously adapt to dynamic user needs, enhancing the stability of the user experience.

[0097] One embodiment of the present invention provides an AI-based intelligent audio player, comprising:

[0098] One or more processors;

[0099] Memory, used to store one or more programs.

[0100] When the one or more programs are executed by the one or more processors, the one or more processors implement the control method described in any one of the above descriptions.

[0101] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A control method for an AI-based intelligent audio player, characterized in that, The method includes: S1. Collect multi-dimensional data from users of smart audio players to generate a comprehensive dataset of user music preferences; based on the comprehensive dataset, construct an initial user music preference analysis model; S2. Real-time data of the playback environment is collected by environmental sensors to generate real-time environmental data; based on the real-time environmental data, combined with pre-set rules and algorithms, the audio player parameters are adaptively adjusted to generate adjusted playback parameter data. S3. Input the comprehensive dataset of user music preferences and the adjusted playback parameter data into the AI ​​analysis module based on deep learning. Through matching and analysis of massive music data, generate a personalized song recommendation list for each user and obtain personalized recommendation result data. S4. Collect user feedback data on playback effect in real time, perform correlation analysis between the feedback data and personalized recommendation result data, evaluate the accuracy of recommended songs and user satisfaction; based on the evaluation results, dynamically optimize and adjust the initially constructed user music preference analysis model to generate an optimized music preference analysis model. S5. Based on the optimized music preference analysis model and real-time collected environmental and user feedback data, set early warning indicators; provide relevant suggestions to users through an intelligent interactive interface, and generate intelligent interactive prompt data; further adjust the playback strategy according to the user's response, and generate complete intelligent audio playback control scheme data.

2. The control method for an AI-based intelligent audio player according to claim 1, characterized in that, S1 includes: S11: Construct a multi-dimensional user music data collection framework and determine the dimensional boundaries and indicator system for data collection; S12: Based on the user data collection framework, the user's raw data is collected through the local storage module of the smart audio player and the cloud synchronization interface to generate the user's raw music behavior dataset; S13: Clean and standardize the original dataset of user music behavior, and integrate multi-source data through data fusion technology to generate a comprehensive dataset of user music preferences; S14: Based on a comprehensive dataset of user music preferences, an initial user music preference analysis model is constructed using machine learning algorithms.

3. The control method for an AI-based intelligent audio player according to claim 1, characterized in that, S2 includes: S21: Based on the usage scenarios of the smart audio player, select a suitable combination of environmental sensors and integrate the sensors into the reserved interface on the player's shell. S22: Real-time environmental data is collected at a frequency of 100ms / time through the sensor data acquisition interface. After removing data noise through Kalman filtering, a real-time environmental data stream is generated. S23: Construct a rule base for mapping environmental parameters and playback parameters. The rules include basic rules and scenario-based rules. The rule base is stored in the local algorithm module of the player. S24: Input the real-time environmental data stream into the rule base for matching calculation, adaptively adjust the core parameters of the audio player, and generate adjusted playback parameter data.

4. The control method for an AI-based intelligent audio player according to claim 3, characterized in that, S24 includes: Extract key feature values ​​from real-time environmental data streams to generate environmental feature identification data; The environmental feature identification data is matched with the environmental parameter and playback parameter mapping rule base to filter out the suitable basic rules and scenario-based rules and generate a rule matching result set. Based on the rule matching result set, the specific adjustment amount of each core parameter is calculated, and rule conflicts are eliminated through the parameter coordination algorithm to generate a parameter adjustment instruction set. The parameter adjustment instruction set is loaded into the control module of the audio player to complete the adaptive adjustment of the core parameters and generate playback parameter data.

5. The control method for an AI-based intelligent audio player according to claim 1, characterized in that, The S3 includes: S31: Build an AI analysis module based on the Transformer architecture. The analysis module includes a feature extraction layer, an interaction layer, and an output layer, and is deployed on the edge computing unit of the player. S32: The user music preference dataset is fused with the adjusted playback parameter data to generate an input feature matrix, which includes user historical preference features and real-time environment features. S33: Through the AI ​​analysis module, the input feature matrix is ​​deeply mined to identify potential user preference patterns and environmental influence patterns, and to generate a feature map of user and environment association. S34: Based on the user and environment association feature map, perform similarity matching and priority sorting in the music resource library to generate a personalized recommendation list containing 20 candidate songs, and predict the user acceptance probability of each song to obtain personalized recommendation result data.

6. The control method for an AI-based intelligent audio player according to claim 5, characterized in that, S34 includes: Feature extraction is performed on songs in the music resource library, and a standardized song feature library is generated through feature standardization. The cosine similarity algorithm is used to calculate the matching degree between the core features in the user and environment association feature map and the features of each song in the standardized song feature library. Songs with matching degree higher than the preset threshold are selected to generate a preliminary candidate song set. By combining the weight of users' historical interactions with the environmental adaptability coefficient, the initial candidate song set is prioritized and the top 20 songs are selected to generate an ordered candidate recommendation list. The pre-trained user acceptance probability prediction model is invoked. The user and environment-related features of each song in the ordered candidate recommendation list are input, and the user acceptance probability of each song is output. The ordered candidate recommendation list and the acceptance probability are integrated to generate personalized recommendation result data.

7. The control method for an AI-based intelligent audio player according to claim 1, characterized in that, The S4 includes: S41: Design a multimodal feedback acquisition mechanism to capture user feedback signals for recommended songs in real time through the player's touch interface, voice interaction, and behavior tracking; S42: Quantize the feedback signal, convert subjective feedback into a quantitative score of 0-10, convert behavioral feedback into weight coefficients, and generate a quantitative dataset of user feedback. S43: Perform correlation analysis between the quantitative dataset of user feedback and the personalized recommendation results data, calculate the recommendation accuracy and satisfaction score, and generate a model evaluation report; S44: Based on the model evaluation report, an online learning algorithm is used to correct the parameters of the initial user music preference analysis model, dynamically update the user preference feature vector, and generate an optimized music preference analysis model.

8. The control method for an AI-based intelligent audio player according to claim 7, characterized in that, S44 includes: The key indicators and problem attributions in the model evaluation report are broken down to extract the core directions that need to be optimized and generate a list of model optimization requirements. Based on the optimization requirements list, incremental SVM was selected as the core online learning algorithm, the parameter correction range was determined, and a parameter adjustment execution plan was formulated. Input the parameter adjustment execution plan into the initial user music preference analysis model, iteratively correct the model parameters through incremental learning algorithm, synchronously update the user preference feature vector, and generate a temporary optimized model. The effectiveness of the temporary optimization model is verified using quantitative user feedback data from the past 7 days. If the recommendation accuracy improves by ≥10% and the satisfaction score increases by ≥0.8 points, the model optimization is confirmed to be effective, and the optimized music preference analysis model is finally generated.

9. The control method for an AI-based intelligent audio player according to claim 1, characterized in that, The S5 includes: S51: Construct a multi-dimensional early warning indicator system, wherein the multi-dimensional indicators include recommendation effectiveness indicators, environmental adaptation indicators, and user status indicators; S52: The player's real-time monitoring module continuously monitors the warning indicators. When the indicators reach the threshold, the warning mechanism is triggered to generate a warning signal. S53: Based on the early warning signal, push targeted suggestions to the user through the intelligent interactive interface and generate intelligent interactive prompt data; S54: Collect user response data to intelligent interactive prompts, combine it with the optimized music preference analysis model, dynamically adjust the playback strategy, and generate complete intelligent audio playback control scheme data.

10. AI-based smart audio players, including: One or more processors; Memory, used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the control method according to any one of claims 1 to 9.