A music scene recognition method and system based on artificial intelligence

CN122598680APending Publication Date: 2026-08-18GANSU UNIV OF CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610714081.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本发明的主要目的在于提供一种基于人工智能的音乐场景识别方法及系统,通过多维音频特征提取、特征加权分析、深度学习模型融合、反馈优化机制与标签稳定性校验,解决了现有音乐场景识别中准确性不足、风格模糊判断困难及缺乏自适应优化能力的问题

Benefits of technology

本发明通过对音频信号的时频分析提取频率、音量、节奏、音高与时长等多维度特征,并结合滤波去噪和分帧处理,将音频信号以帧为单位在频域中进行精细表达,提升音频特征在复杂音乐结构下的可分辨度。多维特征的重要性在不同场景中存在显著差异,通过对这些差异进行量化分析,并依照场景类型赋予特征动态权重,使得音频表达更具场景辨识性,避免统一特征处理带来的误判倾向。在此基础上,通过融合多种音乐风格模型,将各子模型权重依据输入特征进行动态调整,并以加权平均方式融合输出结果,实现在多风格共存或过渡片段中对主导场景的精准判断,显著提升对边缘风格的识别稳定性。通过对识别标签与原始特征的匹配度分析,可对不一致结果进行反馈纠偏,系统自我优化能力得到显著增强。对优化后的标签进行相似度排序与二次确认,确保输出结果在模糊风格交汇场景下具备更强的合理性与鲁棒性。通过上述一系列技术动作构建起具备精度可调、模型自适应、标签可信度动态校验等能力的识别路径,使输出标签在复杂音乐信号中具备更高的准确率与稳定性,增强场景识别在多源音频数据中的泛化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598680A_ABST
    Figure CN122598680A_ABST
Patent Text Reader

Abstract

The application discloses a music scene recognition method and system based on artificial intelligence, and particularly relates to the field of audio signal processing, wherein frequency, volume, rhythm, pitch and duration and other features are extracted by performing time-frequency analysis and filtering processing on input audio signals to construct an audio feature matrix; then the features are weighted in combination with scene differences to generate a weighted feature vector which is input into a deep learning model for recognition. Through model fusion and feedback optimization mechanism, the output result is dynamically adjusted, and the fuzzy label is secondarily confirmed, so that a stable and accurate scene label is finally output. The application realizes fine description of complex music signals and accurate scene recognition, effectively avoids the deficiency that traditional methods are easy to confuse in multi-style transition and fuzzy scenes, improves the ability of the system in label self-correction and result consistency, makes the recognition result have higher accuracy, robustness and practicality, and can be widely applied to intelligent music recommendation and multi-scene audio processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing, and in particular to a music scene recognition method and system based on artificial intelligence. Background Technology

[0002] The field of audio signal processing technology encompasses the acquisition, analysis, processing, and application of audio data. Its core content mainly involves technologies such as sound signal processing, feature extraction, classification, and recognition, and is widely applied in areas such as speech recognition, audio compression, noise suppression, music recommendation, and sound effect optimization. The goal of audio signal processing is to adapt sound signals to different application needs through analysis and processing, such as improving sound quality, achieving speech recognition, or automatic classification. Key technologies involved in this field include time-frequency analysis, feature extraction methods, and pattern recognition algorithms. With the development of artificial intelligence and machine learning technologies, audio signal processing technology has been widely applied in various industries, especially in smart homes, entertainment, and communications.

[0003] Among them, the AI-based music scene recognition method refers to the method of automatically identifying and classifying music scenes using artificial intelligence technology. The core technology involves using deep learning or other machine learning methods to analyze and process features in audio signals to identify different music scenes. Specifically, it involves extracting multi-dimensional features from audio signals and combining them with a trained model to identify music scenes. This not only covers feature analysis methods for audio signals but also how to effectively classify and identify these features using AI models. Through the training and optimization of machine learning models, this technology can automatically identify and classify different music scenes and is widely used in fields such as intelligent music recommendation and intelligent speaker control.

[0004] In existing technologies, music scene recognition often relies on static feature extraction from audio signals, directly inputting them into a single model for classification. The weights between feature dimensions fail to dynamically adjust according to scene differences, resulting in weak feature generalization ability and difficulty adapting to audio content with mixed or ambiguous styles. Without a weight adaptation mechanism, rhythm-driven rock and pitch-driven classical music often overlap in the feature space, causing the model to confuse edge segments. The single-model output structure lacks a fusion strategy for different style models, making it difficult to determine the dominant features of the current scene from multiple music style dimensions, leading to increased error rates when encountering style transitions or blended genres. Systems generally lack feedback adjustment mechanisms; once the initial recognition deviates from the real scene, subsequent outputs cannot self-correct, and incorrect labels will be misused in recommendation systems, impacting user experience. Some systems fail to perform stability testing on ambiguous labels, resulting in frequent label jumps, especially in meditative or experimental music with weak rhythms and many transitions, further reducing the reliability and practicality of the recognition results. The aforementioned shortcomings directly affect the accuracy of automatic tag generation, personalized recommendations, and content classification in practical applications, limiting their promotional value in complex audio environments. Summary of the Invention

[0005] The main objective of this invention is to provide an artificial intelligence-based music scene recognition method and system. Through multi-dimensional audio feature extraction, feature weighted analysis, deep learning model fusion, feedback optimization mechanism and label stability verification, it solves the problems of insufficient accuracy, difficulty in style ambiguity judgment and lack of adaptive optimization capability in existing music scene recognition.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A music scene recognition method and system based on artificial intelligence, the method comprising: The system acquires the input audio signal, extracts audio features using time-frequency analysis, removes noise using filtering techniques, and performs frame-by-frame processing on the signal. Based on frequency domain transformation, it generates an audio feature matrix for each frame of the audio signal. The audio features include frequency, volume, rhythm, pitch, and duration. The importance of audio features in different scenarios is analyzed, and weighted adjustments are made according to scenario type. Different weights are assigned to features in different scenarios to generate weighted feature vectors. Based on the audio feature matrix and weighted feature vector, the rhythm, spectrum and volume features of the audio signal are analyzed. Combined with models of different music styles, the weights of each sub-model are dynamically adjusted. Through weighted average fusion results, scene recognition labels are generated. Based on scene recognition labels, the matching degree between the labels and audio features is analyzed and optimized. If the labels do not match, the model output is adjusted through a feedback mechanism to generate optimized scene labels. The system checks the stability and accuracy of the optimized scene labels, performs secondary verification on fuzzy labels to confirm their rationality, outputs the final confirmed labels, and completes the output of the recognition results.

[0007] Preferably, the audio features include any one or more of the following: frequency, volume, rhythm, pitch, and duration. The pitch feature is a spectral feature extracted by performing a Fourier transform on the time-domain waveform of the audio signal, and the rhythm feature is rhythm information obtained through periodic analysis of the audio signal.

[0008] Preferably, the weighted feature vector is generated by performing a differential analysis of the importance of each audio feature in multiple music scenarios and then dynamically adjusting its weight in combination with the scenario type. The dynamic adjustment process includes assigning higher weights to specific audio features based on the frequency and rhythm variation patterns in the audio signal.

[0009] Preferably, the scene types include electronic music, classical music, rock music, pop music, jazz music, etc., wherein the classification of the scene types is based on the learning results of the training model on the features of different music styles, and the model is used to infer unknown scenes.

[0010] Preferably, the model is a deep learning network model based on convolutional neural networks or long short-term memory networks. It is trained on a large amount of audio data from different music scenes to extract feature information for each scene. The model classifies the scene based on the features of the input audio signal.

[0011] Preferably, the feedback mechanism includes adjusting model weight values ​​through error feedback to optimize the label matching degree. The feedback mechanism adjusts model parameters by comparing the difference between the current label and the ideal label, and uses gradient descent algorithm or other optimization algorithms for model training and weight update.

[0012] Preferably, the generation of the optimized scene label includes sorting multiple label candidates by matching degree and selecting the label that best matches the audio signal. The matching degree sorting process calculates the similarity between each candidate label and the current audio feature and uses a weighted average method to select the final label.

[0013] This invention also discloses an artificial intelligence-based music scene recognition system for performing the above-described method, comprising: Audio Acquisition and Feature Extraction Module: Used to acquire audio signals and perform denoising, normalization, and time window division, extracting features such as spectrum, volume, rhythm, and duration, and generating an audio feature matrix; Feature Differentiation Analysis Module: Used to call the audio feature matrix, compare the differences in spectrum, rhythm, volume and duration features in different scenarios, assign feature weights, and generate scenario feature weight values; Scene recognition and label generation module: This module performs weighted matching of scene labels based on the audio feature matrix and scene feature weight values, calculates similarity, and outputs scene recognition labels. Cross-scene adaptation and optimization module: used to call scene recognition labels, analyze the changes in audio signals under different scenes, adjust the weights of sub-models, and obtain cross-scene adaptation values; The results output and verification module is used to output the recognition results based on the cross-scene adaptation value and scene recognition label, and compare them with the preset label to obtain the recognition result verification value.

[0014] Compared with the prior art, the present invention has the following beneficial effects: This invention extracts multi-dimensional features such as frequency, volume, rhythm, pitch, and duration from audio signals through time-frequency analysis. Combined with filtering, noise reduction, and frame segmentation, it finely expresses the audio signal in the frequency domain frame by frame, improving the discriminability of audio features in complex musical structures. The importance of multi-dimensional features varies significantly across different scenarios. By quantifying these differences and assigning dynamic weights to features according to scenario type, the audio expression becomes more scenario-specific, avoiding misjudgments caused by uniform feature processing. Furthermore, by fusing multiple music style models and dynamically adjusting the weights of each sub-model based on input features, the output results are fused using a weighted average. This achieves accurate identification of the dominant scenario in multi-style coexistence or transitional segments, significantly improving the stability of edge style recognition. Matching analysis between the identified labels and original features allows for feedback correction of inconsistent results, significantly enhancing the system's self-optimization capability. The optimized labels undergo similarity ranking and secondary verification to ensure the output results have stronger rationality and robustness in scenarios with overlapping styles. By employing the aforementioned series of technical actions, a recognition path is constructed that possesses capabilities such as adjustable precision, adaptive model, and dynamic verification of label credibility. This enables the output labels to achieve higher accuracy and stability in complex music signals and enhances the generalization ability of scene recognition in multi-source audio data. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the working steps of the present invention; Detailed Implementation

[0016] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0017] It should be understood that the terms "system," "device," "unit," and / or "module" as used in this specification are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0018] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0019] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0020] The following is a detailed description of the AI-based music scene recognition method provided in the embodiments of this specification.

[0021] like Figure 1 As shown, a music scene recognition method based on artificial intelligence has the following specific steps: Step 1: Acquire the input audio signal, extract audio features using time-frequency analysis, remove noise using filtering techniques, and perform frame segmentation on the signal. Generate the audio feature matrix for each frame of the audio signal based on frequency domain transformation. The audio features include frequency, volume, rhythm, pitch, and duration. In the specific implementation process, the system first receives externally input audio signals, which can originate from user-uploaded audio files, streaming audio data, or real-time microphone signals. The audio data is recorded into the system at a standard sampling rate (e.g., 44.1kHz) and undergoes preliminary preprocessing steps to ensure the stability and accuracy of subsequent analysis. Preprocessing includes background noise removal, signal amplitude normalization, and frame segmentation in the time dimension. In this embodiment, the system uses Short-Time Fourier Transform (STFT) to perform time-frequency analysis on the signal, converting the audio signal into a spectrogram, which serves as the basis for subsequent feature extraction. Audio features include any one or more of the following: frequency, volume, rhythm, pitch, and duration. The pitch feature is the spectral feature extracted by performing a Fourier transform on the time-domain waveform of the audio signal, and the rhythm feature is the rhythm information obtained through periodic analysis of the audio signal.

[0022] In the audio feature extraction stage, the system extracts multi-dimensional feature information for each frame of signal. These features include, but are not limited to, frequency features, volume features, rhythm features, pitch features, and duration features. Frequency features reflect the energy distribution of audio across different frequency bands; volume features characterize the overall intensity and dynamic range of the audio; rhythm features are obtained by detecting periodic fluctuations in the signal, such as the density and periodic changes of rhythm points; pitch features are obtained by performing a Fourier transform on the original time-domain waveform signal and extracting the spectral features composed of the fundamental frequency and its harmonic structure; duration features measure the duration of musical events. These features are uniformly represented in matrix form, forming an audio feature matrix, which provides the input basis for subsequent scene recognition.

[0023] Step 2: Analyze the differences in the importance of audio features in different scenarios, and adjust the weights according to the scenario type. Assign different weights to features in different scenarios to generate a weighted feature vector. The weighted feature vector is generated by dynamically adjusting the weights of each audio feature in combination with the scenario type after conducting a differential analysis of the importance of each audio feature in multiple music scenarios. The dynamic adjustment process includes assigning higher weights to specific audio features based on the frequency and rhythm changes in the audio signal.

[0024] The scene types include electronic music, classical music, rock music, pop music, jazz music, etc. The scene type classification is based on the learning results of the training model on the characteristics of different music styles, and the model is used to infer unknown scenes.

[0025] Simple feature extraction is insufficient for accurate scene classification. Since different music scenes exhibit significantly different sensitivities to audio features—for example, electronic music relies more on low-frequency energy fluctuations, classical music emphasizes melodic structure and pitch stability, and rock music typically contains high-intensity rhythmic percussion elements—this invention introduces a feature weighting mechanism based on scene differences. Specifically, the system incorporates multiple music scene models. After analyzing the importance of each scene to each feature dimension, it assigns differentiated weights to each feature based on the preliminary feature distribution of the input audio, generating a weighted feature vector. The weight assignment is not statically set but rather a dynamic adjustment strategy built upon learning from a large amount of labeled sample data. In practice, if the input audio exhibits highly regular periodicity in rhythmic features and a clear mid-frequency focusing trend in frequency distribution, the system will increase the weights of rhythm and frequency dimensions while weakening the influence of secondary dimensions such as duration and volume, thereby enhancing the model's responsiveness to key information.

[0026] Step 3: Based on the audio feature matrix and weighted feature vector, analyze the rhythm, spectrum, and volume features of the audio signal. Combine models of different music styles, dynamically adjust the weights of each sub-model, and generate scene recognition labels by weighted averaging the results. The system inputs this vector and the audio feature matrix into the scene recognition module. This module performs inference and judgment based on multiple pre-trained neural network sub-models, including typical deep learning architectures such as Convolutional Neural Networks (CNN) and Long Short-Term Memory Networks (LSTM). The CNN model is mainly responsible for extracting local spatial features from the spectrogram, such as harmonic structure and frequency band jumps; while the LSTM network is good at modeling time series information and is suitable for capturing rhythmic patterns and structural change trends. When performing the recognition task, the system dynamically adjusts the weights of each sub-model to adapt to the feature bias of the current audio. For example, when the rhythmic features of the audio are dominant, the system automatically enhances the response of the LSTM sub-model, giving it a larger weight in the recognition result. Finally, the system fuses the output results of multiple sub-models by weighted averaging to generate preliminary scene recognition labels.

[0027] Step 4: Based on scene recognition labels, analyze the matching degree between the labels and audio features, and perform optimization processing. If the labels do not match, adjust the model output through a feedback mechanism to generate optimized scene labels. The feedback mechanism includes adjusting the model weight values ​​through error feedback to optimize the label matching degree. The feedback mechanism adjusts the model parameters by comparing the difference between the current label and the ideal label, and uses gradient descent algorithm or other optimization algorithms for model training and weight updates.

[0028] Scene recognition results may not be stable, especially in audio segments with ambiguous musical styles or strong transitions, where the results may fluctuate or become ambiguous. To address this, this invention proposes a feedback optimization mechanism to detect and correct the accuracy of initial labels. After initial label generation, the system calls a matching degree calculation module to calculate the similarity between the scene model represented by the label and the feature distribution of the current audio. If the matching degree does not reach a set threshold, the feedback mechanism is triggered. This mechanism adjusts the weights of the intermediate layers of the neural network through error backpropagation, optimizing the model towards higher matching degrees. It can also be combined with gradient descent algorithms or other mainstream optimization strategies such as Adam to update weights. The optimized model outputs recognition results again; if the new label has high consistency with the features, it proceeds to the next stage of the confirmation process.

[0029] Optimizing scene label generation involves ranking multiple candidate labels by matching degree and selecting the label that best matches the audio signal. The matching degree ranking process calculates the similarity between each candidate label and the current audio features and uses a weighted average method to select the final label.

[0030] Step 5: Detect the stability and accuracy of the optimized scene labels, perform secondary verification on fuzzy labels, confirm the rationality of the labels, output the final confirmed labels, and complete the output of the recognition results.

[0031] During the label confirmation phase, the system further tests the stability of the labels. By performing sliding window processing along the audio time dimension, it extracts recognition results for several short time periods and analyzes label change trends. If labels frequently change or multiple labels coexist, the system determines that the current label is unstable. In this case, the system initiates a secondary verification mechanism, re-sorting multiple candidate labels based on their matching degree and selecting the most suitable final label for the current audio using a weighted average strategy. The label selection considers not only static feature matching degree but also factors such as temporal consistency and stylistic coherence to ensure that the recognition results have higher semantic integrity and practical application value.

[0032] This invention also discloses an artificial intelligence-based music scene recognition system, the system comprising: Audio Acquisition and Feature Extraction Module: Used to acquire audio signals and perform denoising, normalization, and time window division, extracting features such as spectrum, volume, rhythm, and duration, and generating an audio feature matrix; Feature Differentiation Analysis Module: Used to call the audio feature matrix, compare the differences in spectrum, rhythm, volume and duration features in different scenarios, assign feature weights, and generate scenario feature weight values; Scene recognition and label generation module: This module performs weighted matching of scene labels based on the audio feature matrix and scene feature weight values, calculates similarity, and outputs scene recognition labels. Cross-scene adaptation and optimization module: used to call scene recognition labels, analyze the changes in audio signals under different scenes, adjust the weights of sub-models, and obtain cross-scene adaptation values; The results output and verification module is used to output the recognition results based on the cross-scene adaptation value and scene recognition label, and compare them with the preset label to obtain the recognition result verification value.

[0033] The present invention is further disclosed below using the intelligent recognition and optimization process of music scenes in an in-vehicle entertainment system as an example: Imagine a user driving on a highway, and the in-car entertainment system is playing a 3-minute track. The track is a fusion piece, with a clean, electronic melody in the first half, gradually transitioning to a rock style with heavy drum beats and powerful vocals in the second half. The user hasn't manually categorized or tagged the track; the system needs to automatically identify the context based on the music's content and output the most appropriate tag for use in the recommendation engine or the in-car human-machine interface system (e.g., adjusting ambient lighting or driving mode based on the current style).

[0034] After the user initiates playback, the system automatically calls the audio acquisition module to sample the audio signal in real time. Once the audio data is input in a standard format, it immediately enters the preprocessing stage. Due to the presence of engine noise and wind noise in the vehicle environment, the system first uses spectral subtraction filtering to reduce noise in the audio signal, and then processes the signal in frames with a time window of 25 milliseconds per frame. Next, the audio signal is converted to the frequency domain, and a frame-level audio feature matrix is ​​generated.

[0035] The features extracted from the audio by the system include the fundamental frequency peak, frequency band energy distribution (for frequency features), root mean square volume value (for volume features), beat cycle detection (rhythm features), Fourier fundamental frequency and harmonic coefficients (pitch features), and cumulative frame length for duration features. At this stage, the system also performs window function-based pitch change rate calculation and rhythm density measurement to obtain higher-level structural information.

[0036] With the feature matrix constructed, the system calls the trained feature difference analysis model to compare the performance of the aforementioned multi-dimensional features in different music scenarios. By comparing the previously constructed training sample data, the system found that the frequency of the music in the first 45 seconds is concentrated in the mid-to-high frequency range, the rhythm is relatively weak, the low-frequency drum beats are sparse, and the pitch changes are gradual—these features are highly consistent with the "electronic music" scenario; while in the latter half from 45 seconds to the end, the low-frequency energy suddenly increases, the rhythm is dense, the lead singer enters a strong screaming section, the pitch vibrates violently, and the spectrum distribution is closer to the rock style.

[0037] Based on these observations, the system implements a feature weighting mechanism, giving higher weights (e.g., 0.35 and 0.30) to timbre (frequency) and pitch in the first part, while increasing the weights of rhythm and volume dimensions in the second part (e.g., rhythm 0.40, volume 0.30), thus constructing a time-related weighted feature vector.

[0038] Subsequently, the system inputs these weighted features into the scene recognition network. In this embodiment, the system integrates a deep recognition model based on a parallel architecture of convolutional neural network (CNN) and long short-term memory network (LSTM). The CNN model analyzes the structural features of the spectrogram to identify harmonic texture and timbre fingerprint; the LSTM model models the time-series dynamic information such as rhythm changes and structural transitions.

[0039] The recognition results show that the CNN model outputs a "electronic music" label with a confidence level of 0.84 for the first part of the music, while the LSTM model identifies the second part as "rock music" with a confidence level of 0.88. Due to the stylistic transitions throughout the music, the system integrates the labels using a sub-model weighted fusion strategy and initially generates an "electronic-rock" fused label.

[0040] Next, the system evaluates the preliminary results using a label matching calculation mechanism. It finds that the "electronic-rock" label lacks a corresponding subclass in the system's label library and cannot be used for recommendation system calls. Therefore, a feedback mechanism is triggered, and the system uses the preset ideal label "rock music" as the target to perform lightweight retraining of the model, locally adjusting the weights of the rhythm gating unit in the LSTM sub-model. The optimized model then identifies the entire piece of music again, outputting "rock music" with a confidence level increased to 0.91.

[0041] To verify the effectiveness of the tag, the system divided the audio into eight consecutive segments over time, performed sub-tag recognition for each segment, and evaluated the tag stability using a moving average method. Analysis showed that six segments were highly consistent with the "rock music" tag, with a similarity greater than 0.85. The remaining two segments were style transition areas, leading the system to determine that the tag had high stability and reasonableness. Finally, the system output the "rock music" tag as the scene recognition result for the song and synchronized it to the in-vehicle main interface and driving behavior prediction module.

[0042] The entire recognition process is completed by the music scene recognition system in the in-vehicle entertainment system. This system integrates functional units such as audio acquisition, feature extraction, feature analysis and weighting, model fusion and label generation, feedback optimization, and output verification. All modules work together to achieve accurate understanding and label output of complex music styles, providing a reliable basis for subsequent personalized recommendations and interactive control.

[0043] The advantages of this invention are as follows: it achieves dynamic recognition in multiple scenarios by extracting and weighting features of audio signals across all dimensions, combined with deep learning technology. It also possesses self-feedback optimization capabilities and a fuzzy labeling mechanism, significantly improving the accuracy and stability of traditional music recognition systems in complex music scenarios. Compared to existing methods that only perform coarse classification based on a single dimension such as spectrum or rhythm, this invention can more precisely characterize the intrinsic structure of music signals and achieve automatic recognition of multiple music styles within a unified system, demonstrating broad application prospects.

[0044] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A music scene recognition method based on artificial intelligence, characterized in that, The method includes: The system acquires the input audio signal, extracts audio features using time-frequency analysis, removes noise using filtering techniques, and performs frame-by-frame processing on the signal. Based on frequency domain transformation, it generates an audio feature matrix for each frame of the audio signal. The audio features include frequency, volume, rhythm, pitch, and duration. The importance of audio features in different scenarios is analyzed, and weighted adjustments are made according to scenario type. Different weights are assigned to features in different scenarios to generate weighted feature vectors. Based on the audio feature matrix and weighted feature vector, the rhythm, spectrum and volume features of the audio signal are analyzed. Combined with models of different music styles, the weights of each sub-model are dynamically adjusted. Through weighted average fusion results, scene recognition labels are generated. Based on scene recognition labels, the matching degree between the labels and audio features is analyzed and optimized. If the labels do not match, the model output is adjusted through a feedback mechanism to generate optimized scene labels. The system checks the stability and accuracy of the optimized scene labels, performs secondary verification on fuzzy labels to confirm their rationality, outputs the final confirmed labels, and completes the output of the recognition results.

2. The music scene recognition method based on artificial intelligence according to claim 1, characterized in that, The audio features include any one or more of the following: frequency, volume, rhythm, pitch, and duration. The pitch feature is a spectral feature extracted by performing a Fourier transform on the time-domain waveform of the audio signal, and the rhythm feature is rhythm information obtained by periodic analysis of the audio signal.

3. The music scene recognition method based on artificial intelligence according to claim 1, characterized in that: The weighted feature vector is generated by performing a differential analysis of the importance of each audio feature in multiple music scenarios, and then dynamically adjusting its weight in combination with the scenario type. The dynamic adjustment process includes assigning higher weights to specific audio features based on the frequency and rhythm variation patterns in the audio signal.

4. The music scene recognition method based on artificial intelligence according to claim 1, characterized in that: The scene types include electronic music, classical music, rock music, pop music, jazz music, etc. The classification of scene types is based on the learning results of the training model on the features of different music styles, and the model is used to infer unknown scenes.

5. The music scene recognition method based on artificial intelligence according to claim 1, characterized in that, The model is based on deep learning networks such as convolutional neural networks or long short-term memory networks. It is trained on a large amount of audio data from different music scenes to extract feature information for each scene. The model classifies scenes based on the features of the input audio signal.

6. The music scene recognition method based on artificial intelligence according to claim 1, characterized in that: The feedback mechanism includes adjusting model weight values ​​through error feedback to optimize label matching. The feedback mechanism adjusts model parameters by comparing the difference between the current label and the ideal label, and uses gradient descent algorithm or other optimization algorithms for model training and weight updates.

7. The music scene recognition method based on artificial intelligence according to claim 1, characterized in that: The generation of optimized scene labels includes sorting multiple candidate labels by matching degree and selecting the label that best matches the audio signal. The matching degree sorting process calculates the similarity between each candidate label and the current audio features and uses a weighted average method to select the final label.

8. A music scene recognition system based on artificial intelligence, used to perform the method according to any one of claims 1-7, characterized in that, The system includes: Audio Acquisition and Feature Extraction Module: Used to acquire audio signals and perform denoising, normalization, and time window division, extracting features such as spectrum, volume, rhythm, and duration, and generating an audio feature matrix; Feature Differentiation Analysis Module: Used to call the audio feature matrix, compare the differences in spectrum, rhythm, volume and duration features in different scenarios, assign feature weights, and generate scenario feature weight values; Scene recognition and label generation module: This module performs weighted matching of scene labels based on the audio feature matrix and scene feature weight values, calculates similarity, and outputs scene recognition labels. Cross-scene adaptation and optimization module: used to call scene recognition labels, analyze the changes in audio signals under different scenes, adjust the weights of sub-models, and obtain cross-scene adaptation values; The results output and verification module is used to output the recognition results based on the cross-scene adaptation value and scene recognition label, and compare them with the preset label to obtain the recognition result verification value.