A Music Emotion Recognition Method and System Based on Multimodal Sliding Window Temporal Modeling

CN122575418APending Publication Date: 2026-08-14HANGZHOU ZHAOZHEN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本发明提供基于多模态滑窗时序建模的音乐情感识别方法及系统,用于至少解决如何在音乐情感识别过程中同时兼顾情感变化定位、跨模态时间对齐、冲突信息保留和全曲情感轨迹建模的问题

Benefits of technology

通过构建结构化输入数据,实现了不同来源音乐信息在统一时间基准下的组织与承接,使后续处理具有明确的数据基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575418A_ABST
    Figure CN122575418A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio signal processing, specifically to a music emotion recognition method and system based on multimodal sliding window temporal modeling. The method includes: acquiring structured input data of the target music; determining the emotion change sequence, the posterior distribution of emotion segment durations, and the grouping results of repeated segments; constructing a multi-scale sliding window and extracting window-level multimodal emotion evidence; performing cross-modal temporal alignment and fusion processing on the window-level multimodal emotion evidence, retaining conflict information, and obtaining an emotion state map; performing duration-constrained trajectory decoding to obtain the full-track emotion evolution trajectory, and determining the music emotion recognition result. This invention can improve the accuracy of emotion transition recognition, complex emotion expression, and full-track emotion analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing, specifically to a music emotion recognition method and system based on multimodal sliding window temporal modeling. Background Technology

[0002] Music content understanding technology is a crucial foundation for digital music processing, and its results directly impact the quality of music retrieval, recommendation and distribution, content annotation, creative assistance, and interactive analysis. In practical applications, systems often need to do more than just determine the general mood of a piece of music; they also need to identify the changes in mood over time, the differences between sections, and the complex emotional states formed by the combined effects of various expressive factors.

[0003] Existing technologies typically employ static representation at the whole-song level, fixed time window segmentation, or simple multimodal splicing. While these methods can provide a single emotional category, they struggle to accurately depict the continuous changes within the music. Especially when there are repetitions in verses and choruses, changes in instrumentation, shifts in singing style, and semantic transitions in lyrics, time shifts and expression conflicts often exist between different information sources. This leads to inaccurate positioning of emotional transitions, averaged weakening of local details, and difficulty in identifying complex emotions, ultimately resulting in a significant deviation between the output and the actual listening experience. Summary of the Invention

[0004] This invention provides a music emotion recognition method and system based on multimodal sliding window temporal modeling, which is used to at least solve the problems of how to simultaneously take into account emotion change localization, cross-modal time alignment, conflict information preservation and whole-song emotion trajectory modeling in the process of music emotion recognition.

[0005] In a first aspect, the present invention provides a music emotion recognition method based on multimodal sliding window temporal modeling, the method comprising: Obtain the structured input data for the target music; Based on the structured input data, determine the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments. Based on the change point detection results of the emotional change sequence, determine the anchor point position, construct a multi-scale sliding window around the anchor point position, and extract window-level multimodal emotional evidence. Based on the posterior distribution of sentiment segment duration, cross-modal temporal alignment of window-level multimodal sentiment evidence is performed, and the aligned window-level multimodal sentiment evidence is fused to retain conflict information and obtain a sentiment state map. Based on the emotional state map, the posterior distribution of emotional segment duration, and the grouping results of repeated segments, the duration-constrained trajectory is decoded to obtain the emotional evolution trajectory of the entire piece, and the music emotion recognition result is determined based on the emotional evolution trajectory of the entire piece.

[0006] Secondly, the present invention provides a music emotion recognition system based on multimodal sliding window temporal modeling, for implementing a music emotion recognition method based on multimodal sliding window temporal modeling, the system comprising: The data construction module is used to acquire structured input data for the target music. The window evidence module is used to determine the sentiment change sequence, the posterior distribution of sentiment segment duration, and the grouping results of repeated segments based on structured input data. It also determines the anchor point position based on the change point detection results of the sentiment change sequence, constructs a multi-scale sliding window around the anchor point position, and extracts window-level multimodal sentiment evidence. The alignment and fusion module is used to perform cross-modal temporal alignment of window-level multimodal sentiment evidence based on the posterior distribution of sentiment segment duration, and to fuse the aligned window-level multimodal sentiment evidence, retaining conflict information to obtain a sentiment state map. The trajectory decoding module is used to perform duration-constrained trajectory decoding based on the emotion state map, the posterior distribution of emotion segment duration, and the grouping results of repeated segments to obtain the full-song emotion evolution trajectory, and to determine the music emotion recognition result based on the full-song emotion evolution trajectory.

[0007] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By constructing structured input data, music information from different sources is organized and integrated under a unified time reference, providing a clear data foundation for subsequent processing.

[0008] By determining the sequence of emotional changes, the posterior distribution of the duration of emotional segments, and the grouping results of repeated segments, the synchronous characterization of the location, duration, and correspondence of emotional changes and repeated segments was achieved.

[0009] By constructing a multi-scale sliding window around the anchor point and extracting window-level multimodal emotional evidence, the simultaneous expression of local instantaneous changes, phrase-level changes, and paragraph-level changes was achieved.

[0010] By performing cross-modal temporal alignment on window-level multimodal sentiment evidence, the temporal correspondence between different modalities is unified, reducing misjudgments caused by temporal misalignment.

[0011] By fusing aligned window-level multimodal sentiment evidence and preserving conflict information, simultaneous representation of dominant and complex sentiments was achieved.

[0012] By combining the emotional state map, the posterior distribution of emotional segment duration, and the grouping results of repeated segments to perform duration-constrained trajectory decoding, stable recovery of the emotional evolution trajectory of the entire piece and accurate output of music emotion recognition results were achieved. Attached Figure Description

[0013] Figure 1This is a schematic flowchart of the method of the present invention; Figure 2 This is a diagram showing the emotional change sequence and anchor point distribution in a specific embodiment of the present invention; Figure 3 This is a heatmap of multimodal evidence time-series alignment in a specific embodiment of the present invention; Figure 4 This is a diagram illustrating the emotional evolution trajectory of the entire piece in a specific embodiment of the present invention; Figure 5 This is a statistical chart comparing the recognition results in a specific embodiment of the present invention; Figure 6 This is a block diagram of the module composition of the system of the present invention. Detailed Implementation

[0014] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation.

[0015] Multimodal sliding window temporal modeling is a hierarchical analysis technique for continuous audio content. Its basic idea is to use a sliding window as the unit of temporal analysis, synchronously organizing multi-source information such as acoustic texture, musical structure, performance expression, and lyrical semantics under a unified time reference. It also jointly characterizes local changes, paragraph continuations, and relationships between repeated segments through temporal correlations between windows. This technique not only focuses on the characteristic states of a single moment or segment but also emphasizes the correspondence, evolution, and constraint relationships between different modalities on a continuous time axis, thus allowing the expression process of musical content to be described in the form of a dynamic trajectory. Based on this, this invention focuses on the structured input of target music, the construction of window-level multimodal emotional evidence, cross-modal temporal alignment, conflict information preservation, and decoding of the entire song's emotional evolution trajectory, proposing a music emotion recognition method based on multimodal sliding window temporal modeling.

[0016] like Figure 1 As shown, a music emotion recognition method based on multimodal sliding window temporal modeling is presented, the method comprising: Obtain the structured input data for the target music; In this embodiment, the original audio file of the target music is first acquired, and then subjected to uniform sampling rate conversion, channel unification, and loudness normalization to form basic audio data. Subsequently, structured input data is extracted from the basic audio data according to a preset frame length and frame shift. The structured input data retains time position markers and data source markers, ensuring that subsequent determination of emotional change sequences, anchor point location, multi-scale sliding window construction, and window-level multimodal emotional evidence extraction are all based on the same time reference. After the structured input data construction is completed, it is written to the task cache and then sent to the next processing stage.

[0017] The structured input data includes at least two of the following: time-frequency features, beat boundaries, harmonic variation sequences, performance feature sequences, and lyric segment alignment results.

[0018] In one embodiment, the additional constraint on the structured input data is that it includes at least two of the following: time-frequency features, beat boundaries, harmonic variation sequences, performance feature sequences, and lyric segment alignment results, and all data items are organized according to the same time index. This constraint is adopted because subsequent sentiment change analysis does not require all modalities to be present simultaneously, but at least two data items from different sources and with different levels of expression are needed to support sentiment change location identification and cross-modal temporal correspondence.

[0019] In practice, the basic audio data is first divided into frames with a fixed frame length, ranging from 20 to 40 milliseconds, and the frame shift can be from 10 to 20 milliseconds. Then, time-frequency features are extracted based on the framing results. These features can be one or more of the following: log-Mel spectrum, spectral centroid, spectral flux, and energy envelope. Beat boundaries are determined jointly by the onset intensity curve, periodic estimation results, and dynamic programming backtracking results. Beat boundaries are used to identify strong beat positions, weak beat positions, and the start and end positions of measures. The harmonic variation sequence is generated based on the chromaticity features after beat synchronization and the chord recognition results. The harmonic variation sequence records at least the chord switching positions and the start and end times of consecutive harmonic segments.

[0020] The performance feature sequence is extracted from the vocal track or the dominant melody track. This sequence may include one or more of the following: intensity variation, fundamental frequency fluctuation, vibrato depth, duration of sustained notes, and liaison distribution. Lyric segment alignment results are obtained by aligning the lyrics text with the audio over time. These results must record at least the segment text, the segment start time, and the segment end time. If the lyrics text is missing or the lyrics time alignment confidence is below a preset alignment threshold, the lyrics segment alignment results are not written into the structured input data. Instead, at least two items are selected from time-frequency features, beat boundaries, harmonic variation sequences, and performance feature sequences to form the structured input data.

[0021] The preset alignment threshold can be determined based on the average alignment error in the development samples, for example, to keep the segment boundary error within one beat cycle. After completing the extraction of each data item, each data item is encapsulated into structured record units according to a unified time index. Each structured record unit includes at least a time start point, a time end point, a data type identifier, and data content. If the time span of any data item covers multiple beat intervals, it is segmented according to the beat boundaries before being written into the structured record unit to avoid cross-boundary aliasing during subsequent window segmentation. Through the above method, the structured input data can be kept consistent in terms of data composition, time organization, and missing data handling, thereby providing stable input for subsequent processing.

[0022] Based on the structured input data, determine the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments. Based on the change point detection results of the emotional change sequence, determine the anchor point position, construct a multi-scale sliding window around the anchor point position, and extract window-level multimodal emotional evidence. In this embodiment, after constructing the structured input data, the structured input data undergoes time-axis consistency processing. Based on the consistent structured input data, the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments are determined. Subsequently, anchor point positions are determined based on the change point detection results of the emotional change sequence, and multi-scale sliding windows are constructed centered on the anchor point positions. After each multi-scale sliding window is constructed, at least two types of features are extracted from the corresponding window, including acoustic texture features, musical structure features, performance features, and lyric semantic features. Window-level multimodal emotional evidence is generated based on the emotional support direction, emotional support strength, and evidence confidence corresponding to each type of feature. Window-level multimodal emotional evidence serves as the basic input for subsequent cross-modal temporal alignment and emotional state analysis.

[0023] The process involves determining the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments based on structured input data. This includes: determining the structural novelty sequence, the harmonic tension change sequence, and the performance intensity change sequence based on the structured input data; and, if the structured input data includes lyric segment alignment results, determining the semantic transition sequence based on the lyric segment alignment results; weighted fusion of at least two of the structural novelty sequence, harmonic tension change sequence, performance intensity change sequence, and semantic transition sequence to obtain the emotional change sequence, which indicates the intensity of emotional change along the time axis of the target music; performing change point detection based on the emotional change sequence to obtain the posterior distribution of emotional segment duration and change point detection results; and determining the melody contour similarity relationship and harmonic similarity relationship based on the structured input data, and determining the grouping results of repeated segments based on the melody contour similarity relationship and harmonic similarity relationship.

[0024] In one embodiment, the newly added constraint in this section is that the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments are not generated independently, but are determined synchronously based on the same structured input data under the same time reference. This ensures consistency in the time positions used for subsequent anchor point positioning and window construction. The reason for adopting this constraint is that if the intensity of emotional changes, segment duration, and the correspondence of repeated segments are generated based on different time coordinates, boundary misalignment may easily occur in the subsequent window segmentation stage, making it impossible for the peak of emotional changes at the same time position to accurately correspond to the boundary of repeated segments and the length of the duration segment.

[0025] In practice, unified time sampling is first performed on the time-frequency features, beat boundaries, harmonic variation sequences, performance feature sequences, and lyric segment alignment results in the structured input data. Beat boundaries serve as the first time anchoring reference. Frame-level features are mapped to beat intervals through nearest neighbor mapping or linear interpolation, and lyric segment alignment results are mapped to corresponding beat intervals through the start and end times of the segments. After unified time sampling, a structural novelty sequence is determined based on the self-similar structure in the time-frequency features. The structural novelty sequence characterizes the degree of change in the overall musical structure between adjacent time segments and can be obtained by comparing the differences in spectral distribution, energy distribution, and melodic contour between consecutive beat intervals.

[0026] As the differences between adjacent segments increase, the structural novelty increases. The harmonic tension variation sequence is determined based on the harmonic variation sequence. This sequence characterizes changes in stability during harmonic progression and can be determined based on chord switching frequency, chord type variation amplitude, and the stability of continuous harmonic intervals. The performance intensity variation sequence is determined based on the performance characteristic sequence. This sequence characterizes changes in dynamics, sustain, and variability in singing or playing and can be determined based on changes in intensity, sustain length, vocal strength, and instrument entry / exit. When the structured input data includes lyric segment alignment results, the semantic transition sequence is determined based on these alignment results.

[0027] Semantic transition sequences are used to characterize the degree of change in emotional orientation, narrative focus, or semantic continuity of lyric content. They can be determined based on changes in the emotional polarity of sentence segments, the position of negation words, the position of transitional conjunctions, and changes in sentence boundaries. After extracting the above sequences, at least two sequences are selected from the structural novelty sequence, the harmonic tension change sequence, the performance intensity change sequence, and the semantic transition sequence for weighted fusion to obtain the emotional change sequence. During weighting, the weights corresponding to structural novelty can be set as base weights, while the weights corresponding to harmonic tension and performance intensity can be set as dynamic weights. Supplementary weights are assigned to semantic transitions when lyric segment alignment results exist.

[0028] The principle for setting weights is to ensure that the peak changes at the beat boundaries are as consistent as possible with the actual emotional transition points perceived in the listening experience. During the development phase, initial weights can be determined by statistically analyzing the response ratios of each sequence at these transition points using labeled samples. After the emotional change sequence is generated, change point detection is performed on it to obtain the posterior distribution of the emotional segment duration and the change point detection results. The posterior distribution of the emotional segment duration describes the possible duration range of the current emotional segment, while the change point detection results provide candidate locations for emotional segment transitions. Simultaneously, melodic contour similarity and harmonic similarity are determined based on the structured input data, and the grouping results of repeated segments are determined based on these similarities. Melodic contour similarity can be determined by the consistency of the main melody's direction within different time segments, while harmonic similarity can be determined by the consistency of chord progression patterns. When two time segments simultaneously satisfy the similarity conditions in both melodic contour and harmonic changes, the two time segments are grouped into the same repeated segment group. Through the above processing, it can be ensured that subsequent anchor point positioning, window construction, and evidence extraction are all based on a unified temporal semantics.

[0029] The anchor point is determined based on the change point detection results, and a multi-scale sliding window is constructed around the anchor point. This includes: determining the corresponding time position as the anchor point when the change point detection results corresponding to the local peak of the emotion change sequence meet the preset conditions; determining multiple window lengths corresponding to the anchor point based on the posterior distribution of the duration of the emotion segment; and constructing micro-sliding windows, meso-sliding windows, and macro-sliding windows around the anchor point according to the multiple window lengths.

[0030] In one embodiment, the newly added constraint in this section is that the determination of the anchor point position and the construction of the multi-scale sliding window are not based on a fixed time segmentation method, but rather on a joint determination based on the change point detection results and the local peaks of the emotional change sequence, thereby making the window boundaries closer to the actual emotional turning point. The reason for introducing this constraint is that if the window is directly segmented according to a fixed number of seconds or a fixed number of beats, it is easy to split the emotional change peaks into adjacent windows, causing the same emotional turning point to be weakened in multiple windows, which is not conducive to the accurate expression of subsequent window-level multimodal emotional evidence.

[0031] In practice, the first step is to determine the local peak positions within the emotional change sequence. Local peaks are obtained by comparing emotional change values ​​within consecutive beat intervals. When the emotional change value of the current interval is higher than that of the adjacent intervals and also higher than a preset lower peak limit, the current interval is designated as a candidate local peak. The preset lower peak limit can be determined based on the mean and fluctuation range of the emotional change sequences in the development samples; for example, it can be the mean plus double the standard fluctuation range, or the change values ​​of the top 30% after sorting. Then, the candidate local peaks are matched with the change point detection results. If the change point detection result corresponding to the local peak meets preset conditions, the corresponding time position is determined as the anchor point position.

[0032] Preset conditions may include a change-point confidence level reaching a preset change-point threshold, a time difference between a local peak and the most recent change-point location not exceeding one beat cycle, and a high probability of sustained change in the posterior distribution of the sentiment segment duration. If a local peak meets the peak condition but not the change-point condition, that location is reserved as an auxiliary location and not directly used as the anchor point location. After the anchor point location is determined, multiple window lengths corresponding to the anchor point location are determined based on the posterior distribution of the sentiment segment duration.

[0033] Multiple window lengths are used to construct micro-level, meso-level, and macro-level sliding windows, respectively. The micro-level sliding window covers instantaneous emotional changes near the anchor point, and its length can be determined based on the short-term interval in the posterior distribution of the emotional segment duration. The meso-level sliding window covers emotional changes within a complete musical phrase or a short segment, and its length can be determined based on the main interval in the posterior distribution. The macro-level sliding window covers sustained and repetitive emotional changes over a longer time period, and its length can be determined based on the upper limit interval in the posterior distribution.

[0034] If the anchor point is close to the start or end of the music, the corresponding window is cropped to ensure its boundaries do not exceed the effective time range of the target music. To avoid excessive overlap between windows corresponding to adjacent anchor points, an upper limit can be set for the overlap ratio of adjacent windows. When the overlap ratio exceeds a preset overlap threshold, the micro-sliding window corresponding to the anchor point with the higher emotional change value is retained first, while the meso- and macro-sliding windows are merged or their boundaries shifted. The preset overlap threshold can be determined based on the window scale; typically, a higher overlap threshold is used for micro-sliding windows, while lower overlap thresholds are used for meso- and macro-sliding windows. In this way, the anchor point position and window scale can jointly reflect the temporal location and scope of influence of emotional transitions, enabling the subsequent evidence extraction stage to perceive both local abrupt changes and the emotional continuation within continuous segments.

[0035] Extracting window-level multimodal sentiment evidence includes: extracting at least two features from acoustic texture features, musical structure features, performance features, and lyric semantic features for each multi-scale sliding window; and generating window-level multimodal sentiment evidence based on the sentiment support direction, sentiment support strength, and evidence confidence corresponding to each feature.

[0036] In one embodiment, the newly added limitation in this section is that the window-level multimodal sentiment evidence is not directly output as a set of original features, but is uniformly organized according to sentiment support direction, sentiment support strength, and evidence confidence, forming an evidence expression result that can be used for subsequent cross-modal processing. The reason for adopting this limitation is that the original feature dimensions, temporal resolution, and physical meaning of different modalities vary greatly. If various original features are directly input into subsequent stages side by side, it is easy to make it difficult to make unified comparisons between different modalities, and it is also not conducive to identifying support and conflict relationships between multimodalities.

[0037] In practice, at least two features from acoustic texture features, musical structure features, performance features, and lyrical semantic features are extracted for each multi-scale sliding window. Acoustic texture features are mainly used to characterize the spectral distribution, energy fluctuations, and timbre changes within the window, and can be obtained by combining log-Mel spectrum, spectral centroid, energy envelope, and spectral flux. Musical structure features are mainly used to characterize the beat organization, harmonic progression, and melodic trend within the window, and can be obtained by combining beat distribution, chord switching density, ascending melodic proportion, and descending melodic proportion.

[0038] Performance features primarily characterize changes in singing or playing style within a window, and can be obtained by combining the rate of change in intensity, the proportion of sustained notes, the intensity of the onset, and the stability of vocal production. Lyric semantic features primarily characterize the emotional tendency and semantic shifts in the lyrics within a window, and can be obtained by combining the emotional orientation of sentence segments, the occurrence of negation structures, and the position of transition words. After feature extraction, sentiment support direction analysis is performed on each feature. Sentiment support direction indicates which type of emotional change the current feature tends to support, such as increasing, decreasing, stabilizing, or shifting. Sentiment support strength indicates the degree to which the current feature supports the corresponding emotional direction, and can be determined based on the deviation of the feature value in the current window from the historical mean or the magnitude of change relative to adjacent windows. Evidence confidence indicates whether the current feature is sufficient to stably characterize the emotional state of the window, and can be determined based on feature completeness, sampling quality, temporal alignment quality, and modal stability. For example, when lyric segments cross multiple window boundaries, the evidence confidence of the corresponding lyric semantic features decreases; when the vocal signal is too weak, causing instability in performance features, the evidence confidence of the corresponding performance features decreases.

[0039] Subsequently, the sentiment support direction, sentiment support strength, and evidence confidence level corresponding to each feature are uniformly organized into window-level multimodal sentiment evidence. Window-level multimodal sentiment evidence includes at least a window identifier, modality type, sentiment support direction, sentiment support strength, evidence confidence level, and time range. If multiple modalities within the same window provide support for the same sentiment direction, these items are recorded separately during the evidence generation stage without premature merging, allowing for further assessment of support and conflict relationships between modalities during subsequent cross-modal temporal alignment and fusion stages. Through this process, features from different sources and at different scales can be converted into a unified form of evidence that can be directly used in subsequent processing.

[0040] Based on the posterior distribution of sentiment segment duration, cross-modal temporal alignment of window-level multimodal sentiment evidence is performed, and the aligned window-level multimodal sentiment evidence is fused to retain conflict information and obtain a sentiment state map. In this embodiment, after obtaining window-level multimodal sentiment evidence, cross-modal temporal alignment is first performed on the evidence fragments corresponding to different modalities based on the posterior distribution of sentiment segment durations. This establishes a stable correspondence between acoustic texture features, musical structure features, performance features, and lyric semantic features within the same time range. After temporal alignment, the aligned window-level multimodal sentiment evidence is fused, retaining conflict information between different modalities during the fusion process to obtain a sentiment state graph. The sentiment state graph records the support of each time window for different sentiment states and the conflict between different pieces of evidence. The sentiment state graph serves as the direct input for subsequent duration-constrained trajectory decoding.

[0041] Cross-modal temporal alignment of window-level multimodal sentiment evidence is performed based on the posterior distribution of sentiment segment duration. This includes: using each multi-scale sliding window as a unified time coordinate, constructing an alignment cost based on the window time position, the time positions of different modal evidence segments, the feature distance of different modal evidence segments, and the posterior distribution of sentiment segment duration; and performing optimal transmission alignment of different modal evidence segments with temporal order constraints based on the alignment cost to obtain aligned window-level multimodal sentiment evidence.

[0042] In one embodiment, the newly added limitation in this section is that cross-modal temporal alignment is not directly completed based on a single temporal overlap relationship. Instead, it uses each multi-scale sliding window as a unified temporal coordinate, and combines the window temporal position, the temporal positions of different modal evidence fragments, the feature distance of different modal evidence fragments, and the posterior distribution of the duration of emotional segments to jointly determine the alignment cost. Then, the optimal transmission alignment with temporal order constraints is completed based on the alignment cost. The reason for adopting this limitation is that although evidence fragments from different modalities all originate from the same target music, the generation granularity of each modality is not consistent. For example, acoustic texture features are usually generated continuously with short frame lengths, musical structure features are usually generated in units of beat intervals or phrase intervals, the change position of performance features may correspond to abrupt changes in intensity or sustained tone boundaries, and lyric semantic features are usually organized by sentence segments. If only the temporal overlap length is used as the alignment basis, it is easy to mismatch evidence fragments with similar semantics but slightly offset temporal positions in adjacent windows, and it is also easy to forcibly classify evidence fragments with significantly inconsistent durations into the same window.

[0043] In practice, a unified time coordinate system is first established using the start and end times of a multi-scale sliding window. Then, for any given window, different modal evidence fragments falling within or adjacent to that window's time range are extracted, and the deviation between the window's time position and the time positions of each evidence fragment is calculated. The window's time position can be the window's center time or the time interval formed by the window's start and end times; the evidence fragment's time position can be the fragment's center time or a fragment's time interval. Subsequently, the characteristic distance between different modal evidence fragments is calculated. The characteristic distance characterizes the similarity in emotional expression between different modal evidence fragments and can be determined comprehensively based on whether the emotional support direction is consistent, the magnitude of the difference in emotional support strength, and the magnitude of the difference in evidence confidence.

[0044] The feature distance is small when two pieces of evidence have the same sentiment support direction, a small difference in sentiment support strength, and high evidence confidence. Conversely, the feature distance is large when the two pieces of evidence have opposite support directions or significant strength differences. The alignment cost is then adjusted by incorporating the posterior distribution of the sentiment segment duration. This posterior distribution reflects the possible duration range of the sentiment segment within the current time interval. If the posterior distribution indicates a longer current sentiment segment, longer evidence segments adjacent to the current window are allowed to participate in alignment. If the posterior distribution indicates a shorter current sentiment segment, only evidence segments with closer temporal positions are allowed to participate in alignment. Based on this alignment cost, optimal transmission alignment with temporal order constraints is performed on evidence segments of different modalities.

[0045] Temporal order constraints ensure that evidence fragments from previous time positions are preferentially associated with the preceding window, and evidence fragments from subsequent time positions are preferentially associated with the following window, thus preventing the alignment results from having reversed temporal order. After alignment, aligned window-level multimodal sentiment evidence is generated, where each window corresponds to at least two types of modal evidence, and each type of modal evidence has a consistent correspondence in terms of temporal position and sentiment expression. If only a single modal evidence within a window satisfies the alignment conditions, the window is marked as a low-completeness window, and its credibility is reduced in subsequent fusion stages.

[0046] The aligned window-level multimodal sentiment evidence is fused to retain conflict information and obtain a sentiment state map. This includes: determining the reliability of evidence based on alignment concentration, modal data quality, and window integrity; based on evidence theory, the aligned window-level multimodal sentiment evidence is fused according to the reliability of evidence to retain conflict information and obtain window conflict intensity and window sentiment state support vectors; the persistence of conflict is determined based on the persistence of window conflict intensity, and a sentiment state map is constructed based on the window sentiment state support vectors and conflict persistence.

[0047] In one embodiment, the newly added constraint in this section is that the sentiment state graph is not directly obtained by splicing aligned window-level multimodal sentiment evidence. Instead, the reliability of the evidence is first determined based on alignment concentration, modal data quality, and window integrity. Then, the aligned window-level multimodal sentiment evidence is fused based on evidence theory, retaining conflict information during the fusion process. This yields the window conflict intensity, window sentiment state support vector, and conflict persistence in sequence. Finally, the sentiment state graph is constructed based on the window sentiment state support vector and conflict persistence. The reason for introducing this constraint is that different modalities may provide consistent or inconsistent sentiment support for the same time window. Directly averaging or simply weighting the evidence from each modality may yield a single sentiment state result, but it would lose important conflict information that characterizes complex emotions and sentiment transitions.

[0048] In practice, the alignment concentration is first calculated based on the alignment results. Alignment concentration characterizes whether different modal evidence fragments within a window correspond to the same time range. A higher alignment concentration occurs when all modal evidence fragments are concentrated near the center of the window; a lower alignment concentration occurs when the fragments are scattered across multiple adjacent windows. Modal data quality characterizes whether the original data for each modality is sufficient to support a stable judgment, such as whether there is significant noise interference in the audio signal, whether the human voice is clear, whether the alignment of lyrics is reliable, and whether the harmonic sequence is continuous and complete.

[0049] Window integrity characterizes whether a window simultaneously contains a sufficient number of valid modal evidences. Window integrity can be determined based on the number of valid modalities, the proportion of valid time coverage, and boundary truncation. After determining the reliability of evidence based on alignment concentration, modal data quality, and window integrity, the aligned window-level multimodal sentiment evidence is fused using evidence theory. During fusion, contradictory sentiment support is not directly eliminated; instead, the support and conflict levels of different modalities for each candidate sentiment state are statistically analyzed separately. This yields the window conflict intensity and the window sentiment state support vector.

[0050] Window conflict intensity characterizes the degree of inconsistency in sentiment judgments among different modalities within the current window, while the window sentiment state support vector characterizes the support distribution of the current window for each candidate sentiment state. Subsequently, conflict persistence is determined based on the duration of window conflict intensity. Conflict persistence distinguishes between transient and persistent conflicts and can be determined by whether the window conflict intensity continuously exceeds a preset conflict threshold across multiple consecutive windows.

[0051] The preset conflict threshold can be determined based on the average intensity difference between consistent windows and conflicting windows in the development samples. When the window conflict intensity of a certain window exceeds the preset conflict threshold, and at least some windows in the adjacent windows also meet the same condition, the conflict persistence of the current window is increased; when the window conflict intensity only increases briefly within a single window, this situation is considered a local perturbation. Finally, a sentiment state graph is constructed based on the window sentiment state support vector and conflict persistence.

[0052] In the sentiment state graph, each window corresponds to a state node, which records at least the sentiment state support distribution and conflict persistence. State connections are established between adjacent windows, representing the continuation trend of sentiment states along the time axis. If the grouping of repeated segments indicates that two time intervals belong to the same repeated segment group, a repeated segment connection can be established in the sentiment state graph to allow subsequent decoding stages to consider both temporal continuity and repeated segment correspondence. Through this process, conflict information can be preserved while maintaining consistency across different modalities, enabling the sentiment state graph to express not only the dominant sentiment state but also the formation basis of complex emotions and transitional states.

[0053] Based on the emotional state map, the posterior distribution of emotional segment duration, and the grouping results of repeated segments, the duration-constrained trajectory is decoded to obtain the emotional evolution trajectory of the entire piece, and the music emotion recognition result is determined based on the emotional evolution trajectory of the entire piece.

[0054] In this embodiment, after the emotional state graph is constructed, the emotional state graph, the posterior distribution of emotional segment duration, and the grouping results of repeated segments are jointly input into the decoding stage. The decoding stage first constructs a set of candidate emotional states, then determines the state matching score of each candidate emotional state based on the state support of each window in the emotional state graph, and corrects the state matching score by incorporating conflict persistence. Subsequently, the duration distribution of each candidate emotional state is determined based on the posterior distribution of emotional segment duration, and the emotional state transition constraints between different time segments are determined based on the grouping results of repeated segments. Duration constraint trajectory decoding is performed based on the state matching score, duration distribution, and emotional state transition constraints to obtain an emotional evolution trajectory covering the entire timeline of the song. After the emotional evolution trajectory is generated, the main emotional category, secondary emotional category, emotional turning point, and composite emotional segment are determined based on the emotional evolution trajectory to form the music emotional recognition result.

[0055] Decoding the duration-constrained trajectory based on the sentiment state map, the posterior distribution of sentiment segment duration, and the grouping results of repeated segments includes: constructing a set of candidate sentiment states; determining the state matching score of each candidate sentiment state corresponding to each multi-scale sliding window based on the sentiment state map, and adjusting the state matching score according to the conflict persistence; determining the duration distribution corresponding to each candidate sentiment state based on the posterior distribution of sentiment segment duration, and determining the sentiment state transition constraints between repeated segments based on the grouping results of repeated segments; and decoding the duration-constrained trajectory based on the state matching score, duration distribution, and sentiment state transition constraints to obtain the full-length sentiment evolution trajectory.

[0056] In one embodiment, the newly added constraint in this section is that the duration-constrained trajectory decoding does not directly select the sentiment state with the highest support from the sentiment state graph window by window. Instead, it first constructs a set of candidate sentiment states, and then incorporates the state matching score, duration distribution, and sentiment state transition constraints into the trajectory solving process. This ensures that the sentiment states throughout the entire song satisfy both local evidence support and the correspondence between duration and repeated segments. The reason for adopting this constraint is that musical sentiment is not determined independently by each window, but rather has continuity, segmentation, and repetition. If only the support result of a single window is used as the final sentiment category, frequent jumps between adjacent windows are likely to occur, and the overall sentiment trend brought about by chorus repetition, phrase continuation, and segment echoing is easily overlooked.

[0057] In practice, a candidate emotional state set is first constructed based on the set of emotional labels appearing in the training samples, the high-frequency emotional categories in the development samples, and the existing state node types in the emotional state graph. The candidate emotional state set includes at least one or more of the following categories: stable states, enhancing states, weakening states, and transitional states. It can also be further refined into more specific emotional categories such as calm, excited, lyrical, tense, and sad. The principle for setting candidate emotional states is to cover the actually distinguishable emotional categories in the target task, while avoiding an excessively large set that could lead to unstable state distinctions.

[0058] After constructing the candidate sentiment state set, the state matching score for each candidate sentiment state corresponding to each multi-scale sliding window is determined based on the sentiment state graph. The state matching score is used to characterize the degree of consistency between the current window and a certain candidate sentiment state, and can be determined based on the support value of the corresponding state in the window's sentiment state support vector. When the sentiment state graph contains support results for multiple similar sentiment states, it can also be corrected by combining the difference in support value ranking and the distance relationship between states.

[0059] Subsequently, the state matching score is adjusted based on the conflict persistence. Conflict persistence reflects whether the conflict between different modalities in the current window and its adjacent windows persists. When the conflict persistence is low, it indicates that the multimodal evidence in the current window is generally consistent, and the original state matching score can be maintained. When the conflict persistence is high, it indicates that the current window is in a possible complex emotion or emotion transition region, and the state matching score of a single emotion state is reduced, while the state matching score of a complex emotion state or transition state consistent with the conflict characteristics is increased. The adjustment range of conflict persistence can be determined based on the average duration and average intensity of conflict windows in the development samples to ensure that the adjustment process can reflect the conflict state without allowing local abnormal noise to excessively affect the entire trajectory.

[0060] After correcting the state matching score, the duration distribution of each candidate emotional state is determined based on the posterior distribution of the emotional segment duration. The duration distribution constrains the duration of a candidate emotional state within a continuous window. If the posterior distribution of the emotional segment duration indicates a tendency for longer duration in the current time segment, the corresponding candidate emotional state is allowed to remain unchanged for multiple consecutive windows; if the posterior distribution indicates a higher likelihood of short-term changes in the current segment, the allowed duration of the corresponding candidate emotional state is shortened. The duration distribution can be set according to three levels: short, medium, and long time segments, or it can be set based on the average duration and variance range of each emotional category in the training samples. Simultaneously, emotional state transition constraints between repeated segments are determined based on the grouping results of repeated segments. These constraints limit the consistency or approximate consistency of emotional category changes among time segments belonging to the same repeated segment group. For example, when the same chorus segment appears twice, its emotional category usually remains consistent or only the intensity changes. Therefore, when decoding the trajectory, a high consistency constraint is set between the two repeated segments. If the two repeated segments have significant differences in performance intensity or harmonic tension, the emotional intensity is allowed to shift, but the main emotional category remains consistent.

[0061] After determining the duration distribution and emotional state transition constraints, duration constraint trajectory decoding is performed based on state matching scores, duration distribution, and emotional state transition constraints. During decoding, within the entire timeline, emotional state paths with high support within the current window, reasonable duration between consecutive windows, and satisfying repetitive segment constraints are progressively selected, ultimately yielding the entire emotional evolution trajectory. If multiple emotional state paths have similar scores within a certain time segment, the path with smoother transition to the previous stable segment is prioritized. When repetitive segment constraints conflict with local high-conflict windows, the transition states corresponding to the local high-conflict windows are retained first, and then the repetitive segment constraints are restored in adjacent windows to avoid smoothing out genuine emotional changes.

[0062] The music emotion recognition results are determined based on the emotional evolution trajectory of the whole piece, including: determining the main emotion category and the secondary emotion category based on the emotional evolution trajectory of the whole piece; determining the emotional turning point based on the switching position after adjacent emotional states in the emotional evolution trajectory of the whole piece last for more than a preset duration; and determining the trajectory segment in the emotional evolution trajectory of the whole piece that corresponds to the conflict duration exceeding the preset threshold as a composite emotional segment when the conflict duration exceeds the preset threshold.

[0063] In one embodiment, the newly added limitation in this section is that the music emotion recognition result does not simply select a single category from the emotional evolution trajectory of the entire piece, but further distinguishes between primary emotion categories, secondary emotion categories, emotional turning points, and complex emotional segments. This ensures that the output result reflects both the dominant emotion of the entire piece and retains information on local changes and complex states. The reason for introducing this limitation is that target music typically contains both dominant emotions and local transitional and mixed emotions during its complete playback. If only one category is output, key information about the emotional evolution trajectory in the time dimension will be lost, and it will not accurately reflect situations such as the enhancement of repeated segments, the persistence of local conflicts, and the switching of emotional states.

[0064] In practice, the primary and secondary emotion categories are first determined based on the overall emotional evolution trajectory of the entire piece. The primary emotion category is determined by statistically analyzing the cumulative duration, cumulative state matching score, and coverage of each emotion category throughout the entire emotional evolution trajectory. Emotion categories with longer cumulative duration, higher cumulative state matching scores, and coverage of multiple main sections are prioritized as primary emotion categories. Secondary emotion categories are determined by selecting those from the remaining emotion categories that rank highly in both cumulative duration and cumulative state matching score.

[0065] To avoid the formation of pseudo-sub-emotion categories by local noise windows, a minimum duration condition can be set. When the cumulative duration of a certain emotion category is lower than a preset sub-emotion duration threshold, that emotion category will not be identified as a sub-emotion category. The preset sub-emotion duration threshold can be determined based on the average duration of a single complete musical phrase in the development samples, ensuring that the sub-emotion category covers at least one perceptible paragraph length.

[0066] Subsequently, emotional turning points are determined based on the switching positions of adjacent emotional states that persist for more than a preset duration in the overall emotional evolution trajectory of the piece. Emotional turning points are used to identify the time points where a stable change in emotional category occurs. The preset duration is used to filter out transient fluctuations and short-term disturbances, and can be determined based on the average stable duration before and after actual emotional turning points in the development samples. For example, when an emotional state persists for more than two mesoscopic windows or lasts for more than the average length of a musical phrase before and after a switching, that switching position can be identified as an emotional turning point. If adjacent emotional states only switch briefly within a single microscopic window, it is not determined as an emotional turning point, but rather recorded as a local fluctuation.

[0067] Finally, when the conflict duration exceeds a preset threshold, the trajectory segments in the overall emotional evolution trajectory corresponding to the conflict duration exceeding the preset threshold are identified as composite emotional segments. Composite emotional segments are used to characterize time intervals in which multimodal evidence maintains significant conflict in duration and cannot be fully explained by a single emotional category. The preset threshold can be set based on the average conflict intensity and average duration of the conflict window in the development sample, ensuring that only persistent conflict rather than occasional conflict forms composite emotional segments.

[0068] For complex emotional segments, the main sources of conflict forming the segment can be further recorded, such as the directional difference between acoustic texture features and lyrical semantic features, or the intensity difference between performance features and musical structural features. After the above processing, the output music emotion recognition result includes at least the main emotion category and the emotional turning point. If more granular results are required, the secondary emotion category and complex emotional segment can also be output simultaneously. If there is no time segment in the entire emotional evolution trajectory where the conflict duration exceeds a preset threshold, the complex emotional segment is not output, and only the main emotion category, secondary emotion category, and emotional turning point are retained. In this way, the music emotion recognition result can form a stable correspondence with the entire emotional evolution trajectory and maintain continuity with the previous state decoding process.

[0069] This specific embodiment uses a 210-second pop song as an example to illustrate the practical application of the present invention. The target music uses mono audio with a sampling rate of 22.05 kHz. First, loudness normalization and silence clipping are performed on the original audio. Then, basic time-frequency data is extracted according to a frame length of 2048 sampling points and a frame shift of 512 sampling points, resulting in a 128-dimensional log-Mel spectrum sequence. Subsequently, beat boundaries, chord changes, performance feature sequences, and lyric alignment results are extracted. A total of 344 beat positions, 96 harmonic change segments, 344 performance feature sampling points, and 18 lyric segments were identified in this target music. The above results are uniformly mapped to the same time axis to form structured input data.

[0070] In this embodiment, based on structured input data, the structural novelty sequence, harmonic tension change sequence, performance intensity change sequence, and semantic transition sequence are further determined. The structural novelty sequence reflects the magnitude of change in the overall musical structure between adjacent time segments; the harmonic tension change sequence reflects the stability change of chord progressions; the performance intensity change sequence reflects changes in intensity fluctuations, duration of sustained notes, and vocalization; and the semantic transition sequence reflects changes in the emotional orientation and semantic continuity of the lyrics. After uniformly resampling the four types of sequences to a beat-level time coordinate, a weighted fusion is performed to obtain the emotional change sequence. The target music exhibits five significant peaks near 34 seconds, 72 seconds, 108 seconds, 146 seconds, and 182 seconds, with corresponding emotional change values ​​of approximately 0.48, 0.67, 0.49, 0.66, and 0.61, respectively. Based on the change point detection results, the posterior principal values ​​of the durations of the five emotional segments are obtained, which are 10 seconds, 18 seconds, 14 seconds, 22 seconds, and 16 seconds, respectively. At the same time, based on the similarity of melodic outlines and harmonic similarity, the entire piece is divided into a verse repetition group, a chorus repetition group, and a bridge section independent group.

[0071] like Figure 2 As shown, the horizontal axis represents time, and the vertical axis represents the emotional change sequence and the corresponding probability of change. Five anchor points are marked by green dashed lines. 34 seconds, 72 seconds, 108 seconds, 146 seconds, and 182 seconds correspond to five local peak positions, respectively. The 10-second, 18-second, 14-second, 22-second, and 16-second intervals at the bottom of the figure correspond to the main durations obtained after change point detection. This figure is drawn based on the fused emotional change sequence and change point detection results, and is used to visually represent the correspondence between anchor point positions and durations.

[0072] After determining the anchor point locations, multi-scale sliding windows are constructed around each anchor point. Microscopic sliding windows cover local instantaneous changes, mesoscopic sliding windows cover changes near a complete musical phrase, and macroscopic sliding windows cover paragraph-level changes. This embodiment uses 16 mesoscopic sliding windows as the main analysis windows, and extracts acoustic texture features, musical structure features, performance features, and lyrical semantic features within each window. These features are further transformed into emotional support direction, emotional support strength, and evidence confidence, thus forming window-level multimodal emotional evidence. For example, in W1 to W3, the strengths of acoustic, structural, and performance evidence were all above 0.69, while the strength of semantic features was only between 0.22 and 0.35, indicating that this stage was mainly driven by the intro and accompaniment. In W4 to W6, the strength of semantic features rose to 0.71, 0.79, and 0.83, while the strengths of acoustic, structural, and performance evidence all decreased to between 0.28 and 0.42, indicating that this stage was mainly driven by the lyrics. In W8 to W10, the strengths of acoustic, structural, and performance evidence reached between 0.79 and 0.90, while the strength of semantic features remained between 0.36 and 0.42, indicating that this stage corresponds to the sustained enhancement section before and after the chorus.

[0073] like Figure 3 As shown, the horizontal axis represents the window number, the vertical axis represents the modality type, and the color intensity represents the alignment strength between different modal evidence and the window. The values ​​in the figure represent the alignment results for the corresponding window. In W1 to W3, the alignment strengths of acoustic, structural, and performance modalities are all above 0.74, indicating a relatively concentrated multimodal correspondence at this stage. In W4 to W6, the alignment strength of the semantic modality reaches 0.71 to 0.83, while the other modalities decrease, indicating that lyric segments become the main alignment basis at this stage. In W11 to W13, the semantic strengths are 0.74, 0.81, and 0.86, while acoustic and performance modalities remain above 0.55, indicating a persistent multimodal coexistence and potential conflict at this stage. This figure is statistically derived from the optimal transmission alignment after applying temporal order constraints to 16 mesoscopic sliding windows and four types of modal evidence fragments, and is used to represent the concentration of different modal evidence within each window.

[0074] After cross-modal temporal alignment, the aligned window-level multimodal sentiment evidence enters the fusion processing stage. First, the reliability of the evidence is determined based on alignment concentration, modal data quality, and window integrity. Windows with high alignment concentration and at least three effective modalities within the window have higher evidence reliability; windows with only two types of evidence and incomplete temporal coverage have lower evidence reliability. Subsequently, based on evidence theory, the multimodal evidence in each window is fused, retaining conflict information between different modalities to obtain window conflict intensity and window sentiment state support vectors. In this embodiment, the window conflict intensity near 92 seconds rises to approximately 0.71, and the window conflict intensity near 188 seconds rises to approximately 0.90, both exceeding the preset conflict threshold of 0.58, and the corresponding conflict durations cover 9 seconds and 11 seconds respectively, thus being identified as persistent conflict segments. A sentiment state graph is constructed based on the window sentiment state support vectors and conflict persistence, and then duration-constrained trajectory decoding is performed by combining the posterior distribution of sentiment segment durations and the grouping results of repeated segments. The decoding results show that the overall sentiment evolution trajectory is calm, soothing, tense, passionate, lyrical, and complex. The durations of each segment are 38 seconds, 36 seconds, 36 seconds, 42 seconds, 30 seconds, and 28 seconds, respectively.

[0075] like Figure 4 As shown, the upper curve represents conflict persistence, with the red dashed line indicating a conflict threshold of 0.58. The lower colored blocks represent the decoded emotional evolution trajectory of the entire piece. Conflict persistence near 92 seconds and 188 seconds is higher than the threshold, reaching approximately 0.90 near 188 seconds, corresponding to the composite segment in the lower half. This figure is jointly drawn based on window conflict intensity, conflict persistence, and the final trajectory decoding results, illustrating how persistent conflict participates in the determination of composite emotional segments and how the dominant emotional state evolves along the timeline.

[0076] In this embodiment, the musical emotion recognition result is further determined based on the emotional evolution trajectory of the entire piece. The primary emotion category is determined not only by its duration but also by the cumulative state matching score of each emotional state. Statistically, the cumulative duration of the rousing section is 42 seconds, with an average state matching score of approximately 0.42 and a cumulative contribution value of approximately 17.64; the cumulative duration of the tense section is 36 seconds, with an average state matching score of approximately 0.38 and a cumulative contribution value of approximately 13.68; although the calm section lasts for 38 seconds, its average state matching score is only about 0.33, and its cumulative contribution value is about 12.54. Therefore, rousing is ultimately determined as the primary emotion category, and tense is determined as the secondary emotion category. The emotional turning points occur around 38 seconds, 74 seconds, 110 seconds, 152 seconds, and 182 seconds, respectively. 38 seconds and 74 seconds correspond to the stable transition from the intro to the verse and from the verse to the beginning of the chorus, while 182 seconds corresponds to the starting position of the complex emotion section. The sustained conflict segment around 188 seconds was further identified as a complex emotional segment, characterized by lyrical semantics and persistently inflammatory acoustic and performance evidence.

[0077] To verify the authenticity and effectiveness of this embodiment, the method of this invention was compared with the whole-song static recognition method and the fixed-window fusion method in a test set containing 120 songs with manually labeled emotional turning points and composite emotional tags. The test results show that the macro-average F1 value of the whole-song static recognition method is 0.71, the fixed-window fusion method is 0.79, and the method of this invention is 0.88; the average location error of emotional turning points of the whole-song static recognition method is 12.6 seconds, the fixed-window fusion method is 8.9 seconds, and the method of this invention is 4.7 seconds; the composite emotional recall rate of the whole-song static recognition method is 0.43, the fixed-window fusion method is 0.58, and the method of this invention is 0.81.

[0078] like Figure 5 As shown, the three groups of bars represent the macro-average F1 score, the average localization error of the emotional turning point, and the composite emotional recall rate, respectively. The figure shows that the method of this invention outperforms the two comparative methods in both the macro-average F1 score and the composite emotional recall rate, while significantly lowering the average localization error of the emotional turning point. This figure was drawn based on the statistical results of 120 test songs and is used to illustrate that the method of this invention has good implementation effects in full-song emotional recognition, turning point localization, and composite emotional recognition.

[0079] As can be seen from the above implementation process, the input and output relationships at each stage in this embodiment are clear. The structured input data supports the determination of the emotional change sequence and the grouping results of repeated segments. The emotional change sequence and the change point detection results support the construction of the anchor point position and the multi-scale sliding window. The multi-scale sliding window supports the generation of window-level multimodal emotional evidence. The alignment and fusion results support the construction of the emotional state graph. The emotional state graph, the posterior distribution of the duration of emotional segments, and the grouping results of repeated segments together support the decoding of the emotional evolution trajectory of the whole song, and finally form an interpretable music emotion recognition result.

[0080] like Figure 6 As shown, a music emotion recognition system based on multimodal sliding window temporal modeling is used to implement a music emotion recognition method based on multimodal sliding window temporal modeling. The system includes: The data construction module is used to acquire structured input data of the target music. The data construction module consists of an audio input interface, a preprocessing circuit, a storage unit, and a processor. The audio input interface is used to receive the target music data. The preprocessing circuit is used to perform sampling rate unification, channel unification, loudness normalization, and basic feature extraction. The storage unit is used to cache the structured input data. The processor is used to schedule the hardware units to complete the organization of time-frequency features, beat boundaries, harmonic change sequences, performance feature sequences, and lyric segment alignment results.

[0081] The window evidence module is used to determine the sentiment change sequence, the posterior distribution of sentiment segment duration, and the grouping results of repeated segments based on structured input data. It also determines the anchor point position based on the change point detection results of the sentiment change sequence, constructs a multi-scale sliding window around the anchor point position, and extracts window-level multimodal sentiment evidence. The window evidence module consists of a processor, a feature analysis unit, a temporal detection unit, and a cache unit. The processor is used to control the calling of structured input data, the feature analysis unit is used to determine the sentiment change sequence, the posterior distribution of sentiment segment duration, and the grouping results of repeated segments, the temporal detection unit is used to determine the anchor point position and construct a multi-scale sliding window based on the change point detection results of the sentiment change sequence, and the cache unit is used to store window-level multimodal sentiment evidence.

[0082] The alignment and fusion module is used to perform cross-modal temporal alignment of window-level multimodal sentiment evidence based on the posterior distribution of sentiment segment durations, and to fuse the aligned window-level multimodal sentiment evidence, retaining conflict information to obtain a sentiment state graph. The alignment and fusion module consists of a processor, an alignment calculation unit, a fusion calculation unit, and a state storage unit. The processor is used to schedule the temporal processing of window-level multimodal sentiment evidence, the alignment calculation unit is used to perform cross-modal temporal alignment based on the posterior distribution of sentiment segment durations, the fusion calculation unit is used to fuse the aligned window-level multimodal sentiment evidence and retain conflict information, and the state storage unit is used to store the sentiment state graph.

[0083] The trajectory decoding module is used to perform duration-constrained trajectory decoding based on the emotion state graph, the posterior distribution of emotion segment duration, and the grouping results of repeated segments to obtain the entire song's emotion evolution trajectory, and to determine the music emotion recognition result based on the entire song's emotion evolution trajectory. The trajectory decoding module consists of a processor, a constraint decoding unit, a result determination unit, and an output interface. The processor is used to call the emotion state graph, the posterior distribution of emotion segment duration, and the grouping results of repeated segments. The constraint decoding unit is used to perform duration-constrained trajectory decoding and generate the entire song's emotion evolution trajectory. The result determination unit is used to determine the music emotion recognition result based on the entire song's emotion evolution trajectory. The output interface is used to output the main emotion category, the secondary emotion category, the emotion turning point, and the composite emotion segment.

[0084] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A music emotion recognition method based on multimodal sliding window temporal modeling, characterized in that, The method includes: Obtain the structured input data for the target music; Based on the structured input data, determine the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments. Based on the change point detection results of the emotional change sequence, determine the anchor point position, construct a multi-scale sliding window around the anchor point position, and extract window-level multimodal emotional evidence. Based on the posterior distribution of the duration of the emotional segments, the window-level multimodal emotional evidence is cross-modal temporally aligned, and the aligned window-level multimodal emotional evidence is fused to retain conflict information and obtain an emotional state graph. Based on the emotional state map, the posterior distribution of the duration of the emotional segment, and the grouping results of the repeated segments, the duration-constrained trajectory is decoded to obtain the emotional evolution trajectory of the entire song, and the music emotion recognition result is determined based on the emotional evolution trajectory of the entire song.

2. The method according to claim 1, characterized in that, The structured input data includes at least two of the following: time-frequency features, beat boundaries, harmonic variation sequences, performance feature sequences, and lyric segment alignment results.

3. The method according to claim 1, characterized in that, The process of determining the emotional change sequence, the posterior distribution of emotional segment duration, and the grouping results of repeated segments based on the structured input data includes: Based on the structured input data, determine the structural novelty sequence, the harmonic tension change sequence, and the performance intensity change sequence; and if the structured input data includes the lyric segment alignment result, determine the semantic transition sequence based on the lyric segment alignment result. At least two of the structural novelty sequence, the harmonic tension change sequence, the performance intensity change sequence, and the semantic transition sequence are weighted and fused to obtain the emotional change sequence, which is used to indicate the intensity of emotional change of the target music along the time axis. Based on the emotional change sequence, change point detection is performed to obtain the posterior distribution of the duration of the emotional segment and the change point detection results; Based on the structured input data, melodic contour similarity and harmonic similarity are determined, and based on the melodic contour similarity and harmonic similarity, the grouping results of the repeated segments are determined.

4. The method according to claim 3, characterized in that, The step of determining the anchor point position based on the change point detection result and constructing a multi-scale sliding window around the anchor point position includes: If the change point detection result corresponding to the local peak of the emotional change sequence meets the preset conditions, the corresponding time position is determined as the anchor point position; Determine multiple window lengths corresponding to the anchor point position based on the posterior distribution of the emotional segment duration; Microscopic sliding windows, mesoscopic sliding windows, and macroscopic sliding windows are constructed around the anchor point according to the multiple window lengths.

5. The method according to claim 1, characterized in that, The extraction of window-level multimodal sentiment evidence includes: For each of the multi-scale sliding windows, at least two features are extracted from acoustic texture features, music structure features, performance features, and lyric semantic features; The window-level multimodal sentiment evidence is generated based on the sentiment support direction, sentiment support strength, and evidence confidence level corresponding to each feature.

6. The method according to claim 1, characterized in that, The step of performing cross-modal temporal alignment of the window-level multimodal sentiment evidence based on the posterior distribution of the sentiment segment duration includes: Using each of the multi-scale sliding windows as a unified time coordinate, an alignment cost is constructed based on the window time position, the time positions of different modal evidence fragments, the feature distances of different modal evidence fragments, and the posterior distribution of the duration of the sentiment segment. Based on the alignment cost, the different modal evidence fragments are subjected to optimal transmission alignment with temporal order constraints to obtain aligned window-level multimodal sentiment evidence.

7. The method according to claim 6, characterized in that, The process of fusing the aligned window-level multimodal sentiment evidence, retaining conflict information, to obtain the sentiment state graph includes: The reliability of evidence is determined based on alignment concentration, modal data quality, and window integrity. Based on evidence theory, the aligned window-level multimodal sentiment evidence is fused according to the reliability of the evidence, retaining conflict information to obtain window conflict intensity and window sentiment state support vector; The conflict persistence is determined based on the duration of the window conflict intensity, and the sentiment state graph is constructed based on the window sentiment state support vector and the conflict persistence.

8. The method according to claim 7, characterized in that, The step of decoding the duration-constrained trajectory based on the emotional state map, the posterior distribution of the emotional segment duration, and the grouping results of the repeated segments includes: Construct a set of candidate sentiment states; The state matching score for each candidate emotional state corresponding to each of the multi-scale sliding windows is determined based on the emotional state graph, and the state matching score is adjusted based on the conflict persistence. The duration distribution corresponding to each candidate emotional state is determined based on the posterior distribution of the duration of the emotional segment, and the emotional state transition constraint between repeated segments is determined based on the grouping result of the repeated segments. Based on the state matching score, the duration distribution, and the emotional state transition constraint, the duration constraint trajectory is decoded to obtain the full-song emotional evolution trajectory.

9. The method according to claim 8, characterized in that, The step of determining the music emotion recognition result based on the emotional evolution trajectory of the entire piece includes: The primary and secondary emotion categories are determined based on the emotional evolution trajectory of the entire piece. The emotional turning point is determined based on the switching position after adjacent emotional states have lasted for more than a preset duration in the emotional evolution trajectory of the entire song. If the duration of the conflict exceeds a preset threshold, the trajectory segment in the entire emotional evolution trajectory that corresponds to the duration of the conflict exceeding the preset threshold is identified as a composite emotional segment.

10. A music emotion recognition system based on multimodal sliding window temporal modeling, used to implement the music emotion recognition method based on multimodal sliding window temporal modeling as described in any one of claims 1-9, characterized in that, The system includes: The data construction module is used to acquire structured input data for the target music. The window evidence module is used to determine the emotional change sequence, the posterior distribution of emotional segment duration and the grouping results of repeated segments based on the structured input data, and to determine the anchor point position based on the change point detection results of the emotional change sequence, construct a multi-scale sliding window around the anchor point position, and extract window-level multimodal emotional evidence. The alignment and fusion module is used to perform cross-modal temporal alignment of the window-level multimodal sentiment evidence according to the posterior distribution of the duration of the sentiment segment, and to perform fusion processing on the aligned window-level multimodal sentiment evidence, retaining conflict information to obtain a sentiment state map. The trajectory decoding module is used to perform duration-constrained trajectory decoding based on the emotional state map, the posterior distribution of the duration of the emotional segment, and the grouping results of the repeated segments to obtain the emotional evolution trajectory of the whole song, and to determine the music emotion recognition result based on the emotional evolution trajectory of the whole song.