Flute ethnic music library construction method and system
By carefully dividing and analyzing the audio data of flute songs, identifying the tone and melody structure, and comparing the national style samples, a flute ethnic music library was constructed, solving the problem of failing to effectively capture the tone changes and melody structure in the existing technology, and improving the retrieval and recognition efficiency of the music database.
Patent Information
- Application Number
- CN202510884572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing technology fails to effectively capture and analyze timbre changes and melody structures in the construction of flute ethnic music library, resulting in difficulties for users to find repertoires with specific expressions or stylistic characteristics, affecting the depth of music education and academic research.
By carefully dividing the audio data of flute songs, counting the changes in tone waveforms, identifying the order of scale rise and fall and continuous changes in the melody, calling ethnic style sample tracks for structural comparison, generating a style proportion label set, recording the position of style feature changes, constructing the style segment time feature mapping results, and generating a style library field combination.
It improves the accuracy and meticulousness of music information management, enhances the classification efficiency and accuracy of the database, and improves the convenience and efficiency of users to retrieve and identify specific styles of tracks.
Smart Images

Figure CN120388550A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of music information management, and particularly to a method and system for constructing a national flute repertoire. Background Art
[0002] The technical field of music information management aims to efficiently process, store, and retrieve music-related information. The core content of this field involves the collection, classification, management, and display of music data. By using various data annotation and processing techniques, it ensures the systematic integration and convenient access to music resources. The music information management technology focuses on optimizing the access process of music materials, enhancing the functionality of music databases and the interactivity of user interfaces to meet the retrieval and usage needs of different users.
[0003] Among them, the method for constructing a national flute repertoire refers to establishing a national repertoire related to flute performance through certain technical means. This technical theme involves the audio collection, transcription, and storage of music works, as well as the precise classification and annotation of the collected music data to ensure that all data is organized according to attributes such as ethnic characteristics, style types, and performance techniques. This method also includes track data management to support the effective retrieval and display of tracks, and user interface design to allow performers to quickly filter tracks according to specific styles or difficulties.
[0004] The existing technology focuses on basic data collection and classification when dealing with music information, but it is insufficient in capturing subtle musical expressions and deep structural features. This insufficiency leads to limited functionality of music databases in supporting complex queries and providing refined music analysis. The failure to capture and analyze timbre changes and melody structures in detail makes it difficult for users to find tracks with specific expressive forms or style characteristics, affecting the depth of music education and academic research. Traditional information management technologies have not effectively integrated and utilized the structural characteristics of ethnic music styles, resulting in insufficient dissemination and utilization of specific cultural music resources in global music education and cultural inheritance, and affecting the protection and promotion of music diversity. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art and to propose a method and system for constructing a national flute repertoire.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions. A method for constructing a national flute repertoire includes the following steps: S1: Obtain the audio data of flute tracks, divide the audio content into continuous segments along the time axis, extract the continuous change trajectory of the timbre waveform of each segment, and cumulatively count the number of turns in the timbre state in each waveform segment to generate a time index table of concentrated timbre change segments; S2: Based on the time periods marked in the time index table of the concentrated section of timbre changes, intercept the main melody information within the corresponding segments, identify the ascending and descending order and continuous changes of the scales in the melody, divide nodes according to the directionality of the scale trend, and generate a table of the order of melody trend nodes; S3: According to the path structure in the table of the order of melody trend nodes, call the melody structure of the ethnic style sample pieces, conduct structure comparison according to the node arrangement order, mark the segments with the same structure in the sample, and generate an ethnic style proportion label set; S4: Adopt the ethnic style labels in the ethnic style proportion label set, retrieve the audio samples of the ethnic style, compare the timbre fluctuation segments in the target piece, record the time period intervals on the timeline that are of the same type as the sample fluctuation segments, analyze the changing positions of the style characteristics in the piece, generate a style section time feature mapping result, and identify and classify the style of the flute pieces.
[0007] As a further solution of the present invention, the time index table of the concentrated section of timbre changes includes the turning point density, the start time of the timbre concentrated section, and the end time of the timbre concentrated section. The table of the order of melody trend nodes includes the scale change order, the melody direction, and the arrangement order of path nodes. The ethnic style proportion label set includes the number of structurally similar segments, the proportion information of style categories, and the annotation of matching segments. The style section time feature mapping result includes the style type mark, the start position of the style period, and the end position of the style period.
[0008] As a further solution of the present invention, the specific steps for obtaining the time index table of the concentrated section of timbre changes are as follows: S111: Obtain the audio data of the flute pieces, evenly divide the audio into equal time intervals per second, collect the timbre envelope trajectory data within each time period, and record the audio sampling values at the start time point and the end time point to generate a timbre waveform change sequence; S112: Based on the continuous change sequence in the timbre waveform change sequence, calculate the number of slope sign changes between adjacent sampling points, count the number of slope direction changes in each audio segment, and obtain the waveform turning point quantity; S113: According to the time periods marked as dense distribution features in the waveform turning point quantity, extract the start time and end time nodes of the corresponding paragraphs, evaluate the corresponding relationship between the paragraph numbers and the time position intervals, and obtain the time index table of the concentrated section of timbre changes.
[0009] As a further solution of the present invention, the specific steps for obtaining the table of the order of melody trend nodes are as follows: S211: Based on the time period marked in the time index table of the timbre change concentration segment, intercept the main melody information within the corresponding segment, arrange the pitch values of each note in the main melody in sequence, judge the changing trend of the scale according to the relationship of the pitch values between adjacent notes, generate the rising and falling trend sequence of the melody scale within each time period, and obtain the melody rising and falling trend sequence; S212: Invoke the melody rising and falling trend sequence, divide nodes according to the directionality of the scale trend, extract the paragraphs with continuous directional changes, identify the notes at the reversal of the directional change as the division nodes, number and mark them according to the arrangement order of the nodes in the melody, and obtain the directional division node sequence; S213: Invoke the directional division node sequence, identify the scale span, time span and direction switching frequency between adjacent nodes in the node sequence, calculate the fluctuation degree of the melody structure according to the change relationship between the indicators, and obtain the melody trend node sequence table.
[0010] As a further solution of the present invention, the acquisition steps of the ethnic style proportion tag set are specifically as follows: S311: Based on the path structure in the melody trend node sequence table, invoke the melody structure of the ethnic style sample track, record the start node and end node intervals connected by each path, delimit the melody structure segments in the sample track with the same number of nodes as the path interval nodes as the comparison reference segments, extract the comparison reference segments for each path structure respectively, and generate the path structure comparison segment set; S312: According to the path structure comparison segment set, compare the melody structures in the ethnic style sample tracks, screen out the melody segments that are consistent with the melody structure in terms of the number of nodes, pitch flow, and rhythm distribution, label each matching segment with the ethnic style label, and count the numerical values of the matching segments corresponding to the different ethnic style labels to obtain the ethnic style segment statistical result; S313: Invoke the ethnic style segment statistical result, calculate the ethnic style distribution difference value according to the number of matching segments corresponding to the ethnic style label, combined with the number of matching segments under each ethnic style label, identify the deviation between the difference value corresponding to the ethnic style and the number of segments, and obtain the ethnic style proportion tag set.
[0011] As a further solution of the present invention, the acquisition steps of the style segment time feature mapping result are specifically as follows: S411: Based on the ethnic style proportion tag set, obtain the timbre fluctuation segment in the audio sample, detect the start time and end time on the time axis, calculate the amplitude difference value and time span difference value between adjacent peaks, and obtain the timbre fluctuation rhythm sequence; S412: Invoke the timbre fluctuation rhythm sequence, compare the timbre fluctuation segments in the target track, select the fluctuation segments with the same rhythm distribution, and based on the cross-matching degree between the timbre spectrum change value and the rhythm density value, calculate the rhythm fitting intensity value, screen the target segments that are consistent with the rhythm sequence of the sample fluctuation segment, and obtain the style matching time interval; S413: According to the style matching time interval, record the start and end positions of the interval on the time axis, mark the corresponding ethnic style label, analyze the distribution structure of the ethnic style on the time axis, and generate the style section time feature mapping result.
[0012] As a further solution of the present invention, the method further includes step S5: S5: Based on the style and time period information in the style section time feature mapping result, combined with the track number information, perform format normalization processing on the time period and style identifier, and combine them as the library entry field items of the music library to generate the style music library field combination; The style music library field combination includes the track number, the normalized style identifier, and the style section index information.
[0013] As a further solution of the present invention, the obtaining steps of the style music library field combination are specifically as follows: S511: Based on the style and time period information in the style section time feature mapping result, extract the style identifier and the corresponding time period number, evaluate the corresponding relationship between the style number and the time section number by matching the sequence number of the style field and the start and end marks of the time period, perform numerical recombination on the time period number, and merge the consecutive time periods corresponding to the same style for numbering to generate the style section number sequence; S512: Invoke the style section number sequence and the track number information in the original music library, and based on the position of the track in the time period in the sequence, judge the correspondence between the track number and the style number, and classify each track into the corresponding style identifier according to the time period it belongs to to generate the style track matching group; S513: According to the track number and style identifier information recorded in the style track matching group, standardize the field content in the format of the combination of the style number and the time period number, and uniformly encode the style and time period information of each group of tracks to generate the style music library field combination.
[0014] The ethnic flute music library construction system is used to execute the above-mentioned ethnic flute music library construction method, and the system includes: The audio feature recognition module obtains the complete audio data of the ethnic flute performance track, divides the audio into continuous segments, detects the change trajectory of the timbre waveform in each segment of audio, counts the turning point values between adjacent wave peaks and wave valleys, and records the time position interval on the corresponding time axis to obtain the timbre change concentration section time index table; The melody structure recognition module extracts the main melody note sequence within the time period recorded in the time index table of the timbre change concentration segment, identifies the frequency direction change, and marks and numbers the nodes where the directionality changes to obtain the melody trend node sequence table; The style segment comparison module calls the node arrangement sequence recorded in the melody trend node sequence table, compares the node positions and sequences with the melody structure path in the preset ethnic style sample, and counts the number of matches with the structure segments in the ethnic style sample to obtain the ethnic style proportion label set; The style feature mapping module filters the timbre fluctuation samples under the style label according to the style labels in the ethnic style proportion label set, calls the timbre data of the corresponding paragraph, compares the fitting degrees of the fluctuation period, amplitude, and energy distribution density with the sample fluctuation segment, marks the same type of time period and records the corresponding style to obtain the style section time feature mapping result; The music library construction module calls the style labels and time period information recorded in the style section time feature mapping result, matches the unique number of the real-time track, and combines the style label, the corresponding time period range, and the track number into the standard field format to obtain the style music library field combination.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, by carefully dividing the audio data of the flute tracks and counting the timbre waveform changes, the music information management is significantly improved in terms of accuracy and meticulousness. Indexing the feature segments with dense distribution of waveform turns helps to accurately capture the key timbre changes in the tracks, providing richer and more specific data support for subsequent melody structure analysis. The ascending and descending order and continuous changes in the melody can be carefully identified, and further node division is carried out through the directionality of the scale trend, which not only increases the multi-dimensional visualization of information but also makes the structural features of the melody clearer. Calling the ethnic style sample tracks for structure comparison and counting the consistent style segments can effectively mark the music segments that conform to a specific ethnic style, enhancing the classification efficiency and accuracy of the database. Through style marking and time feature mapping, a highly organized and standardized framework is provided for the integration and retrieval of music library information, greatly improving the convenience and efficiency for users to retrieve and identify tracks of specific styles. Brief Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the working process of the present invention; Figure 2 It is a flowchart for obtaining the time index table of the timbre change concentration segment in the present invention; Figure 3 It is a flowchart for obtaining the melody trend node sequence table in the present invention; Figure 4 It is the flowchart for obtaining the ethnic style proportion tag set in the present invention; Figure 5 It is the flowchart for obtaining the acquisition result of the style section time feature mapping in the present invention; Figure 6 It is the flowchart for obtaining the field combination of the style music library in the present invention. Detailed implementation manners
[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.
[0019] Please refer to Figure 1 , the present invention provides a technical solution, a method for constructing a folk flute music library, including the following steps: S1: Obtain the audio data of the flute music, divide the audio content into continuous segments according to the time axis, extract the continuous change trajectory of the timbre waveform of the segments, cumulatively count the number of turns in the timbre state in each waveform segment, record the time position in the segment where the number of turns reaches the dense distribution feature, and generate the time index table of the timbre change concentration segment; S2: Based on the time period marked in the time index table of the timbre change concentration segment, intercept the main melody information in the corresponding segment, identify the ascending and descending order and continuous change of the scales in the melody, divide the nodes according to the directionality of the scale trend, define the reverse behavior after the same direction change as the path node, record the continuous arrangement order of the path nodes, and identify the fluctuation law of the melody structure to generate the melody trend node sequence table; S3: According to the path structure in the melody trend node sequence table, call the melody structure of the ethnic style sample music, perform structure comparison according to the node arrangement order, mark the segments with the same structure in the sample, and count the number of matching segments in the ethnic style category to generate the ethnic style proportion tag set; S4: Adopt the ethnic style labels in the ethnic style proportion label set, retrieve the audio samples of the ethnic style, compare the timbre fluctuation segments in the target track, record the time period intervals on the time axis that are of the same type as the sample fluctuation segments, assign style marks to the time periods, analyze the change positions of the style characteristics in the track, and generate the style section time feature mapping result; S5: Based on the style and time period information in the style section time feature mapping result, combined with the track number information, perform format normalization processing on the time period and style identifier, and combine them as the library entry field items of the music library to generate the style music library field combination; The time index table of the concentrated timbre change segment includes the turning point density, the start time of the timbre concentrated segment, and the end time of the timbre concentrated segment. The melody trend node sequence table includes the scale change sequence, the melody direction, and the arrangement order of the path nodes. The ethnic style proportion label set includes the number of structurally similar segments, the proportion information of the style categories, and the matching segment annotations. The style section time feature mapping result includes the style type mark, the start position of the style time period, and the end position of the style time period. The style music library field combination includes the track number, the normalized style identifier, and the style section index information.
[0020] Please refer to Figure 2 , and the specific steps for obtaining the time index table of the concentrated timbre change segment are as follows: S111: Obtain the audio data of the flute track, equally divide the audio by second, collect the timbre envelope track data in each time period, and record the audio sampling values at the start time point and the end time point to generate the timbre waveform change sequence; Obtain the audio data of the flute track. The audio is stored in the standard WAV format, with a sampling rate of 44100Hz, recorded in 16-bit mono, 44100 audio sample points are collected per second, divided into continuous time periods in segments of 5 seconds, the total number of sampling points for each segment of data is set to 220500 points, each segment of data is divided into 432 frames, the size of each frame is 512 sampling points, and the frames are slid and processed in an overlapping manner of 256 points. By reading the sampling value array of each frame, calculate the maximum amplitude value of its envelope line, and use this maximum amplitude as the representative value of each frame. Sequentially construct the envelope line change trajectory formed by 432 representative values in the entire time period. For the data with paragraph number 1, record its start time as 0 seconds and the end time as 5 seconds. The corresponding audio is 220500 sampling points, and its envelope trajectory is a change sequence containing 432 amplitude nodes. In actual processing, construct the timbre waveform change sequence of this time period. Each value in this sequence represents the overall energy intensity of the audio signal within this time frame, forming a complete change trajectory array for subsequent turning point judgment and quantity statistics. This processing operation will be executed sequentially for all time periods, serving as the basic array for subsequent judgment of the timbre change area, and generating the timbre waveform change sequence.
[0021] S112: Based on the continuous change sequences in the timbre waveform change sequence, calculate the number of slope sign changes between adjacent sampling points, count the number of slope direction changes in each audio segment, and obtain the waveform turning quantity. Extract the average amplitude value sequence frame by frame, calculate the sign difference of the amplitude values of any two adjacent frames. If the increase and decrease direction of the amplitude value changes between positive and negative, that is, there is a transition from positive to negative or from negative to positive between the average values of the original frame and the next frame, it is recognized as a waveform direction turning. Count the number of positive and negative changes during this time period as the waveform turning quantity of this segment. In actual data, taking paragraph number 1 as an example, it contains 432 frames. Compare 431 groups of adjacent frames in sequence. If there is a slope change, that is, the sign of the difference changes from positive to negative or from negative to positive, the number of turning points is incremented by 1. The total number of change points is counted as 12. 18 change points are identified in paragraph number 2, 9 in paragraph number 3, 22 in paragraph number 4, and 15 in paragraph number 5. Set the dense distribution threshold to 16. Paragraphs with values lower than this value are regarded as general fluctuations, and those higher than or equal to this value are regarded as dense fluctuation segments. According to this judgment condition, select two time periods of paragraph numbers 2 and 4, which meet the characteristics of frequent timbre state transitions, as the subsequent extraction targets. The statistics of the waveform turning quantity are all generated based on the frame difference direction judgment mechanism and recorded in an array to obtain the waveform turning quantity.
[0022] S113: According to the time periods marked with dense distribution characteristics in the waveform turning quantity, extract the start time and end time nodes of the corresponding paragraphs, evaluate the corresponding relationship between the paragraph number and the time position interval, and obtain the time index table of the concentrated timbre change segment. After screening the paragraph numbers whose turning quantity meets the timbre turning density threshold, extract the start time and end time corresponding to this paragraph number, perform a one-to-one mapping process on the time period and the number, and construct a two-way index table of the paragraph number and the time position interval. Taking paragraph number 2 as an example, its start time is 5 seconds and the end time is 10 seconds. The start time of paragraph number 4 is 15 seconds and the end time is 20 seconds. This type of time interval directly comes from the index identification and sampling duration setting in the audio segmenting operation. Therefore, the index value is the original sampling division identification rather than the result of secondary extraction. At the same time, retain the timbre change trajectory sequence number information corresponding to this segment, and mark that the data in this time period can be used for the comparison of timbre feature concentration. It includes identification fields such as paragraph number, start time, end time, and whether it belongs to the high-density change section. This data table will be used as the benchmark source for the time interval screening in the subsequent main melody extraction operation to obtain the time index table of the concentrated timbre change segment.
[0023] Please refer to Figure 3 , and the specific steps for obtaining the melody trend node sequence table are as follows: S211: Based on the time periods marked in the timbre change concentration section time index table, intercept the melodic main line information within the corresponding segments, arrange the pitch values of each note in the melodic main line in sequence, judge the changing trend of the scale according to the relationship between the pitch values of adjacent notes, generate the rising and falling trend sequence of the melody scale within each time period, and obtain the melody rising and falling trend sequence; Clarify the start and end time nodes corresponding to each time period in this index table. Set a certain index range to be from 1.5 seconds to 3.8 seconds. Then intercept the original melody audio data within this range, and extract the melodic main line from this audio segment. Mainly use the pitch detection method for each frame to obtain the frequency curve of the main melody. Uniformly convert the fundamental frequency value of each frame into MIDI value to represent the pitch level. For example, 440Hz is mapped to MIDI value 69 (i.e., A4). Arrange the frames in chronological order to form a pitch sequence, and judge the size relationship between the pitch values of adjacent notes to identify their rising or falling trends. For example, if the pitch of the 3rd frame is 62 and the 4th frame is 64, it is a rising trend. A rising and falling mark is formed between every two frames. Process the entire paragraph in this way to form a complete trend sequence. If there are consecutive frames with no pitch change (set the MIDI value to 66 for 3 consecutive frames), it is marked as a parallel trend. This trend can be removed or classified through post-processing. Through the above operations, the entire melody segment is divided into continuous rising and falling direction segments, set as rising - rising - falling - falling - rising. Further, it is uniformly sorted into a structured data sequence and numbered and marked for subsequent analysis. Among them, when judging the rising and falling trend, the error tolerance range should be considered. Within ±0.5 MIDI values is regarded as the same level of pitch. Therefore, it is necessary to set this error tolerance threshold. In this example, it is set to ±0.3. When the pitch value changes within this range, it is marked as no change. Through this method, the slight vibrato or pitch fluctuation in the melody can be effectively processed, and the melody rising and falling trend sequence can be obtained.
[0024] S212: Invoke the melody rising and falling trend sequence, divide nodes according to the directionality of the scale trend, extract the paragraphs with continuous directional changes, identify the notes at the reversal of the directional change as the division nodes, and number and mark them according to the arrangement order of the nodes in the melody to obtain the directional division node sequence; According to the continuity of the trend direction in each segment, identify the nodes with direction changes as switching points, read the marks in the rising and falling trend sequences. Set a sequence as "rising, rising, falling, falling, rising", then the positions where "rising" turns to "falling" (between the 2nd and 3rd positions) and where "falling" turns to "rising" (between the 4th and 5th positions) are used as directional switching nodes. Locate the corresponding note index positions, MIDI values, and time points of the nodes, and further perform numbering processing. Starting from the first direction switching point, mark them as P1, P2, P3, etc. in sequence to reflect their arrangement order in the melody. In the example, if the first switching point is 2.4 seconds and the second switching point is 3.1 seconds, they are marked as P1 and P2 respectively, and record their original pitch MIDI values (such as 65 and 67 respectively). This numbering is used for subsequent path construction and melody structure modeling. In specific operations, it is necessary to exclude the slight movements caused by non-directional changes, such as the up and down floating trends caused by detection errors. When "rising - falling - rising" appears in the trend sequence but the middle falling amplitude is less than 1 MIDI value, it is determined as an invalid reverse trend and this node is excluded. It is necessary to set an amplitude threshold for judgment, and the value refers to 10% of the average value of the melody span. If the average span is 8, the threshold is set to 0.8. This setting is determined based on the natural change law of the melody and combined with statistical empirical values to identify effective reverse nodes. Its content includes the time point, pitch value, trend turning category, and number index in the overall sequence of each node, obtaining a directional division node sequence.
[0025] S213: Call the directional division node sequence, identify the scale span, time span, and direction switching frequency between adjacent nodes in the node sequence. According to the change relationship between the indicators, use the formula: ; Calculate the fluctuation degree of the melody structure and obtain the melody trend node order table; Among them, represents the fluctuation degree of the melody structure, represents the th scale span value between node pairs, represents the th time span value between node pairs, represents the th direction switching frequency value of node pairs, represents the th same-direction continuous continuation value of node pairs, is the total number of node pairs; Extract the relative positions and directional changes between nodes. Extract the pitch span values between adjacent nodes according to the node index sequence. Obtain the pitch difference by numerically comparing the pitch values of two consecutive nodes. Set the pitch of node 1 to C4 (MIDI value 60) and the pitch of node 2 to A4 (MIDI value 69), then the pitch span value is 9. After obtaining the pitch span value, calculate the time span value by combining the time indices of the two nodes. Set the time of node 1 to 2.1 seconds and the time of node 2 to 3.0 seconds, then the time span value is 0.9 seconds. The direction switching frequency value is calculated by identifying the number of direction changes in the continuous node sequence as a proportion of the total number of node pairs. For example, if there are 3 direction reversals in 5 node pairs, the frequency is 0.6. Collect the direction continuation length as the same-direction continuation value, which is achieved by counting the number of consecutive pitch changes in the same direction between two nodes. Set that there are two notes, E4 and F4, both in the ascending direction between C4 and G4, then the same-direction continuation value is 3. Combine the above parameters and substitute them into the formula: To illustrate the process of parameter setting, the parameters of three node pairs are extracted as follows and presented in an array form: Pitch span value array ; Time span value array ; Direction reversal frequency value array ; Same-direction continuation value array ; Perform the formula calculation based on the above arrays as follows: The numerator part: ; The denominator part: ; Substitute the parameters into the formula: ; The result shows that the degree of fluctuation under the current melody node structure is 0.457, which represents the average change amplitude of the overall melody direction and can further be used as the basis for generating the melody direction node sequence table.
[0026] Please refer to Figure 4 , and the specific steps for obtaining the ethnic style proportion tag set are as follows: S311: Based on the path structure in the melody direction node sequence table, call the melody structure of the ethnic style sample repertoire, record the start node and end node intervals connected by each path, delimit the melody structure segments in the sample repertoire with the same number of nodes as the path interval nodes as the comparison reference sections, extract the comparison reference sections for each path structure respectively, and generate the path structure comparison section set; The path structure in the melodic contour node sequence table corresponds to the sequence combination of multiple melodic structures. In practice, it can be sampled from the transition nodes of the tune segments in the traditional music of a certain region. Taking the melodic nodes of ethnic minority traditional tunes as an example, path 1 can be recorded as [1, 3, 5, 7], path 2 as [2, 4, 6], and path 3 as [1, 4, 6, 8]. When calling the melodic structure of the ethnic style sample tune, by establishing a mapping table, the path node sequence is corresponding to the specific note sequence in the melodic segment of the sample tune. The extraction process takes the starting node and the ending node connected by each path as a reference, locates the corresponding note positions in the sample melody, and delimits the continuous note segments with the same number of path nodes in the sample as the comparison reference section. If path 1 corresponds to 4 nodes, then search for the melodic segment between node numbers 1 and 7 in the sample and extract the note sequence with a corresponding length of 4. Then perform the same processing on path 2 and path 3, unify the set of melodic segments extracted by the path structure into a comparison reference section set, evaluate the sequence correspondence between the node sequence and the note structure, and generate a path structure comparison section set.
[0027] S312: According to the path structure comparison section set, compare the melodic structure in the ethnic style sample tune, screen out the melodic segments that are consistent with the melodic structure in terms of the number of nodes, pitch flow, and rhythm distribution, label each matching segment with an ethnic style label, and count the numerical values of the matching segments corresponding to different ethnic style labels to obtain the ethnic style segment statistical result; When comparing the sample melodic structure, three parameters need to be determined: the number of nodes, pitch flow, and rhythm distribution. Record the node values for each comparison section and directly compare them with the number of segment nodes in the sample to determine whether they match. Secondly, for the determination of pitch flow, it is determined whether the flow is consistent by whether the continuous pitch difference is within the same sign range. Set the pitch sequence of the comparison section as [200, 210, 220], and the sample melodic segment as [198, 208, 218]. Then the pitch differences of both are positive values, and the flow is consistent. If it is [220, 210, 200], it means a descending flow, which needs to be consistent with the direction of the comparison section. The rhythm distribution takes the standard deviation of the adjacent note durations not exceeding 5 ms as the matching standard. Set the section rhythm as [250 ms, 260 ms, 255 ms], and the sample segment rhythm as [248 ms, 262 ms, 257 ms], then it is determined that the rhythm is consistent. Perform the three-parameter judgment on each sample segment respectively. The segments that pass are marked as matching segments, and count the matching segments in the sample, classify them according to their corresponding ethnic style labels to generate a quantity statistics. For example, there are 5 matching segments under style A, 7 under style B, and 6 under style C to obtain the ethnic style segment statistical result.
[0028] S313: Invoke the statistical results of ethnic style segments. According to the number of matching segments corresponding to the ethnic style labels, combined with the number of matching segments under each ethnic style label, use the formula: ; Calculate the ethnic style distribution difference value, identify the deviation between the difference value corresponding to the ethnic style and the number of segments, and obtain the ethnic style proportion label set; Among them, represents the ethnic style distribution difference value of the th ethnic style, represents the node distribution quantity of the th path segment in the th type of ethnic style, represents the average pitch of the th path segment in the th type of ethnic style, represents the average pitch of the comparison reference section in the th path structure, represents the total number of rhythms in the reference section of the th path structure, is the number of paths; Record the number of matching segments, the node density of each path segment, the total number of rhythms, and the pitch distribution characteristic values included in each ethnic style. After extracting the segment data corresponding to each style label, perform numerical calculations on the distribution intensity difference in the sample sequence, using the formula: Substitute the parameter values for Style A as follows: Number of path segments: 5; Node distribution quantity: [9, 10, 8, 9, 9]; Average pitch: [210.0, 212.5, 211.0, 209.5, 210.5]; Average pitch of the comparison section: [200, 200, 200, 200, 200]; Total number of rhythms: [12, 12, 12, 12, 12]; The process of substituting into the formula is as follows: Numerator: ; Denominator: ; Substitute into the formula for calculation: ; The results show that the difference value of ethnic style distribution is 0.388, which represents the degree of deviation between the number of segments of each style label and the expected number in the overall dataset, and can help identify which ethnic styles are over-represented or under-represented in the dataset. The benefit of the formula lies in introducing the product relationship between the square term of node density and the pitch deviation amount, reflecting the structural stability in the difference measurement, analyzing the quantitative relationship between the difference value and the segment distribution density, and obtaining the ethnic style proportion label set.
[0029] Please refer to Figure 5 , and the steps for obtaining the mapping result of the time characteristics of the style section are specifically as follows: S411: Based on the ethnic style proportion label set, obtain the timbre fluctuation segments in the audio sample, detect the start time and end time on the time axis, calculate the amplitude difference value and time span difference value between adjacent peaks, and obtain the timbre fluctuation rhythm sequence; Extract the original audio signal sequence from the sample, linearly segment it according to the time axis, take each 1 second as the analysis window. Assume a track with a total length of 180 seconds is divided into 180 segments. Collect the amplitude signals within each window, sample the waveform, extract the continuously appearing rising and falling wave peaks and valleys as the boundary points of the fluctuation segments, measure the amplitude change value between each wave peak and the adjacent wave valley. If the change exceeds 3dB, it is considered a significant fluctuation. At the same time, record the time span from the wave peak to the wave valley. If the span is less than 0.5 seconds, it is marked as a short-period fluctuation segment. Furthermore, uniformly normalize the amplitude difference and period data to generate a unified rhythm profile vector sequence. Assume that the rhythm profile corresponding to a certain sample at the 60th second is [3.1, 0.6], where 3.1 is the maximum amplitude difference of this segment and 0.6 is the fluctuation period, which can be used for rhythm matching with the target track to obtain the timbre fluctuation rhythm sequence.
[0030] S412: Call the timbre fluctuation rhythm sequence, compare the timbre fluctuation segments in the target track, select the fluctuation segments with the same rhythm distribution, and according to the cross-matching degree of the timbre spectrum change value and the rhythm density value, use the formula: ; Calculate the rhythm fitting intensity value, screen the target segments consistent with the rhythm sequence of the sample fluctuation segments, and obtain the style matching time interval; Among them, represents the rhythm fitting intensity value, represents the rhythm amplitude difference value of the target segment under the specified ethnic style, represents the timbre frequency domain gradient value, represents the number of pitch jumps, represents the rhythm center of gravity offset, represents the th record in the sample for the fitting degree, For the number of records; According to the time distribution characteristics of the sample sequence, perform a sliding window matching process with second-by-second alignment on the target track, scan the target track at a rate of 1 segment per second, extract the amplitude difference value and frequency gradient change value of the current segment. Let the amplitude difference of the th segment of the target segment be , and the frequency change be . If its fluctuation performance is relatively intense, set the pitch jump number according to experience, and detect the offset of its rhythm center of gravity position . For the th to the th style matching paragraphs in the sample, calculate the timbre similarity metric value between them and the target segment, denoted as , , . After averaging, obtain approximately , which serves as the subsequent comparison benchmark; ; Substitute the data for calculation as follows: First step, calculate the part under the square root: ; Second step, multiply by and divide by : ; Third step, find the average similarity of the samples: ; Fourth step, substitute to find the absolute value difference: ; The result shows that the rhythm fitting intensity value is 7.2933, which is used to measure the matching degree between the rhythm structure of a music segment and the rhythm structure of another target music segment. Through this value, the similarity and fitting degree of the music segment in terms of rhythm can be evaluated. According to the preset judgment benchmark, if this value exceeds 5, it is considered that there is a matching association between the target segment and the sample segment under this ethnic style. Therefore, the current segment is included in the candidate matching interval, and multiple intervals that meet this intensity threshold are extracted from the target track as candidate segments to obtain the style matching time interval.
[0031] S413: According to the style matching time interval, record the start and end positions of the interval on the time axis, mark the corresponding ethnic style labels, analyze the distribution structure of the ethnic style on the time axis, and generate the style section time feature mapping result; Extract the start and end times of each matching interval, mark them in seconds, and combine with the scale information of the target track timeline to match the absolute time information corresponding to the interval. Set a currently extracted matching interval as from the 32nd second to the 34th second, then this interval is marked as [32s, 34s]. According to the ethnic style label pointed to in its source sample, assign the "Tibetan" style identifier to this interval. Another example is that the segment from the 45th second to the 47th second corresponds to the "Yi" style, and so on. Process all matching intervals in sequence to form a mapping set of multiple time periods and ethnic styles. To avoid marking conflicts, for overlapping segments of time periods, it is necessary to sort their rhythm fitting intensities, and preferentially retain the paragraph with a higher intensity value. At the same time, perform interval merging processing on the overlapping segments. Set when both the 60th second to the 62nd second and the 61st second to the 63rd second meet the style matching and are marked as "Miao" and "Zhuang" respectively, compare their fitting intensity values. If the Miao paragraph is 6.8 and the Zhuang paragraph is 7.2, then retain the latter and reorganize the entire paragraph into [61s, 63s], delete the item with a lower fitting degree in the overlapping part, and store all paragraph information in a key-value pair structure, where the key is the time period and the value is the ethnic style label, for the subsequent visual output of the style transformation atlas to obtain the style section time feature mapping result.
[0032] Please refer to Figure 6 , the specific steps for obtaining the style library field combination are as follows: S511: Based on the style and time period information in the style section time feature mapping result, extract the style identifier and the corresponding time period number. By matching the sequential number of the style field and the start and end marks of the time period, evaluate the correspondence between the style number and the time section number, and reorganize the numerical values of the time period numbers. Merge the consecutive time periods corresponding to the same style and generate a style section number sequence; It is necessary to clarify the extraction methods of the "style identifier" and the "time period number". By disassembling each element in the two-dimensional array structure in the feature mapping result into a mapping pair in the form of , where represents the style identifier, represents the corresponding time period number. Each pair of data can be regarded as the main style shown in a certain time period in the library. Set in the sample library, the style appears in the time period , , , the style appears in the time period , . Aggregate the time period numbers of the same style into a set, such as , , perform number merging processing on the time period number sequence in each style set in turn, and the merging rule is: if the adjacent time period numbers are continuous, they are merged into one number segment, such as the sequence Merge into segment 1; if the time period is not continuous, it will be split into multiple style segment numbers. In this process, it is necessary to judge the continuity of the number, that is, to judge any satisfy , then merge, otherwise start a new numbering, the merged result can form the following style segment number sequence: , In actual implementation, the input music library is set to include 10 time periods, each of which is 30 days long. The styles are respectively , the execution process is as follows: After extracting the style-time period pairs, build a list , judge the continuity of the number, 1 to 3 are continuous time periods, merged into a number segment , 5 to 6 are another continuous segment, merged into a numbered segment ,Here the number reorganization is done by traversing the time segment number list, setting up a ,temporary pointer to record the difference between the previous number and the current number, ,and when the difference is 1, recording the length of the continuous sequence, and creating a new ,numbering segment when the difference is greater than 1 and marking the sequence count ,increment by one. A one-to-one mapping relationship between the style number and ,the time segment number is used to assist the recording.
[0033] S512: The style segment number sequence and the track number information in the original music library are called, and the track number and style number are matched according to the position of the time period of the track in the sequence. Each track is classified into a corresponding style identifier according to the time period, and a style track matching group is generated; It is necessary to establish a time period positioning mechanism for tracks, and set a corresponding timestamp field for each track in the original music library. , the timestamp needs to be converted into the aforementioned time period number The specific operation is to set the time window length Days, taking the start time of the music library as the benchmark value , calculate the time period number of each track as , where the time unit is "day", if the track The timestamp of is the 62nd day, so the time period number is After obtaining all the track numbers and their corresponding time period numbers, compare the corresponding time period number range in the style segment number sequence, such as the track In time period , and the style section Coverage time period , then the track is classified into the style , perform the above classification process on the tracks, that is, it is necessary to traverse the track set, compare and judge the time period numbers of each track with the range of style section numbers. If the track time period number and the style to which the set belongs , then the track number is stored corresponding to the style identifier . The specific judgment operation is to set a loop variable to read the track number and time period number, and then perform a condition match to determine whether the time period is in the time period set of a certain style number. If it holds, add the mapping pair of the track and the corresponding style to the matching group to form a set of style-track matching groups.
[0034] S513: According to the track numbers and style identifier information recorded in the style-track matching group, standardize the field content in the format of the combination of style number and time period number, and uniformly encode the style and time period information of each group of tracks to generate a combined style library field; When standardizing the fields according to the track numbers and style identifier information recorded in the style-track matching group, it is necessary to construct a unified encoding mechanism, that is, assign a combined label of style number and time period number to each group of tracks. The format can be set in the form of "F number - T number". For example, if the track belongs to the style and belongs to the time period , then its field identifier is "F1 - T3". It is necessary to clarify the style identifier to which each track belongs and its time period number. The style identifier has been obtained from the matching group, and the time period number is obtained by reverse inference of the timestamp. Call its number, style identifier and time period number for each track, and form a unified identification field through string splicing. This process needs to traverse the track set, generate the corresponding combined field for each track and write it into the standard field set. If the track is of the style , and the time period , then its combined field is "F2 - T5". The combined fields form a new field set. During the execution of this operation, the situation of duplicate identifiers needs to be avoided. Therefore, a uniqueness verification mechanism for combined fields needs to be established, that is, perform deduplication processing on the field list after generating the combined fields. The judgment logic is whether the field already exists in the identifier set. If it exists, skip writing; if it does not exist, add it to the column. It not only includes the track number, but also completes the mapping with its style and time period attributes. Standardize the classification of each track and mark a unique style-time period combination code to form a combined style library field.
[0035] The system for constructing the ethnic flute library is used to execute the above method for constructing the ethnic flute library. The system includes: The audio feature recognition module obtains the complete audio data of the ethnic flute performance repertoire, divides the audio into continuous segments, detects the change trajectory of the timbre waveform in each segment of audio, counts the turning point values between adjacent wave peaks and wave valleys, records the time position interval on the corresponding time axis, and obtains the timbre change concentration segment time index table; The melody structure recognition module extracts the main melody note sequence within the time period recorded in the timbre change concentration segment time index table, recognizes the change in the frequency direction, and obtains the melody trend node sequence table by marking and numbering the nodes where the directionality changes; The style segment comparison module calls the node arrangement order recorded in the melody trend node sequence table, compares the node positions and orders with the melody structure path in the preset ethnic style sample, counts the number of matches with the structure segments in the ethnic style sample, and obtains the ethnic style proportion label set; The style feature mapping module filters the timbre fluctuation samples under the style label according to the style labels in the ethnic style proportion label set, calls the timbre data of the corresponding segment, compares the fitting degrees of the fluctuation period, amplitude, and energy distribution density with the sample fluctuation segment, marks the same type of time periods and records the corresponding styles, and obtains the style segment time feature mapping result; The music library construction module calls the style labels and time period information recorded in the style segment time feature mapping result, matches the unique number of the real-time repertoire, combines the style labels, the corresponding time period interval, and the repertoire number into the standard field format, and obtains the style music library field combination.
[0036] The above is only the preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical content of the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A method for constructing a national flute music library, characterized in that, It includes the following steps: S1: Obtain the audio data of the flute repertoire, divide the audio content into continuous segments along the time axis, extract the continuous change trajectory of the timbre waveform of the segments, cumulatively count the number of turns in the timbre state in each waveform segment, and generate a time index table of concentrated timbre change segments; S2: Based on the time periods marked in the time index table of concentrated timbre change segments, intercept the main melody information within the corresponding segments, identify the ascending and descending order and continuous changes of the scales in the melody, divide nodes according to the directionality of the scale trend, and generate a sequence table of melody trend node orders; S3: According to the path structure in the sequence table of melody trend node orders, call the melody structure of the ethnic style sample repertoire, conduct structural comparison according to the node arrangement order, mark the segments with the same structure in the sample, and generate an ethnic style proportion label set; S4: Adopt the ethnic style labels in the ethnic style proportion label set, retrieve the audio samples of the ethnic style, compare the timbre fluctuation segments in the target repertoire, record the time period intervals of the same type as the sample fluctuation segments on the time axis, analyze the change positions of the style characteristics in the repertoire, generate a style section time feature mapping result, and identify and classify the style of the flute repertoire.
2. The method for constructing a flute ethnic music library according to claim 1, wherein The time index table of concentrated timbre change segments includes the turning point density, the start time of the timbre concentrated segment, and the end time of the timbre concentrated segment. The sequence table of melody trend node orders includes the scale change order, the melody direction, and the arrangement order of path nodes. The ethnic style proportion label set includes the number of structurally similar segments, the proportion information of style categories, and the annotation of matching segments. The style section time feature mapping result includes the style type mark, the start position of the style period, and the end position of the style period.
3. The method for constructing a flute national music library according to claim 1, characterized in that The specific steps for obtaining the time index table of concentrated timbre change segments are as follows: S111: Obtain the audio data of the flute repertoire, equally divide the audio by second, collect the timbre envelope trajectory data within each time period, and record the audio sampling values at the start time point and the end time point to generate a timbre waveform change sequence; S112: Based on the continuous change sequence in the timbre waveform change sequence, calculate the number of slope sign changes between adjacent sampling points, count the number of slope direction changes in each audio segment, and obtain the waveform turning number; S113: According to the time periods marked as dense distribution features in the waveform turning number, extract the start time and end time nodes of the corresponding segments, evaluate the corresponding relationship between the segment numbers and the time position intervals, and obtain the time index table of concentrated timbre change segments.
4. The method for constructing a flute ethnic music library according to claim 3, wherein The specific steps for obtaining the sequence table of melody trend node orders are as follows: S211: Based on the time periods marked in the time index table of concentrated timbre change segments, intercept the main melody information within the corresponding segments, arrange the values of each note in the main melody in sequence, judge the change trend of the scale according to the relationship between the pitch values of adjacent notes, generate a sequence of the ascending and descending trends of the melody scale within each time period, and obtain a sequence of melody ascending and descending trends; S212: Call the melody rising and falling trend sequence, divide nodes according to the scale directionality, extract paragraphs with continuous directional changes, identify the notes at the reversal of the directional change as division nodes, and number and mark them according to the arrangement order of the nodes in the melody to obtain a directional division node sequence; S213: Call the directional division node sequence, identify the scale span, time span, and direction switching frequency between adjacent nodes in the node sequence, calculate the degree of melody structure fluctuation based on the change relationship between the indicators, and obtain a melody trend node sequence table.
5. The method for constructing a flute ethnic music library according to claim 4, wherein The steps for obtaining the ethnic style proportion label set are specifically as follows: S311: Based on the path structure in the melody trend node sequence table, call the melody structure of the ethnic style sample tracks, record the start node and end node intervals connected by each path, delimit the melody structure segments in the sample tracks that are consistent with the number of nodes in the path interval as comparison reference segments, extract the comparison reference segments for each path structure respectively, and generate a path structure comparison segment set; S312: According to the path structure comparison segment set, compare the melody structures in the ethnic style sample tracks, screen out the melody segments that are consistent with the sample tracks in terms of the number of nodes, pitch flow, and rhythm distribution, label each matching segment with an ethnic style label, and count the numerical values of the matching segments corresponding to different ethnic style labels to obtain an ethnic style segment statistical result; S313: Call the ethnic style segment statistical result, calculate the ethnic style distribution difference value according to the number of matching segments corresponding to the ethnic style label, combined with the number of matching segments under each ethnic style label, identify the deviation between the difference value corresponding to the ethnic style and the number of segments, and obtain the ethnic style proportion label set.
6. The method for constructing a flute ethnic music library according to claim 5, characterized in that, The steps for obtaining the time feature mapping result of the style segment are specifically as follows: S411: Based on the ethnic style proportion label set, obtain the timbre fluctuation segments in the audio sample, detect the start time and end time on the time axis, calculate the amplitude difference value and time span difference value between adjacent peaks to obtain a timbre fluctuation rhythm sequence; S412: Call the timbre fluctuation rhythm sequence, compare the timbre fluctuation segments in the target track, select the fluctuation segments with the same rhythm distribution, and calculate the rhythm fitting intensity value according to the cross-matching degree of the timbre spectrum change value and the rhythm density value, screen out the target segments that are consistent with the sample fluctuation segment rhythm sequence to obtain a style matching time interval; S413: According to the style matching time interval, record the start and end positions of the interval on the time axis, mark the corresponding ethnic style label, analyze the distribution structure of the ethnic style on the time axis, and generate a time feature mapping result of the style segment.
7. The method for constructing a flute ethnic music library according to claim 1, characterized in that The method further includes step S5: S5: Based on the style and time period information in the time feature mapping result of the style segment, combined with the track number information, perform format normalization processing on the time period and style identifier, and combine them as the library entry field items of the music library to generate a style music library field combination; The style music library field combination includes the track number, normalized style identifier, and style segment index information.
8. The method for constructing a flute ethnic music library according to claim 7, wherein The steps for obtaining the style music library field combination are specifically as follows: S511: Based on the style and time period information in the style section time feature mapping result, extract the style identifier and the corresponding time period number. By matching the sequential number of the style field with the start and end markers of the time period, evaluate the correspondence between the style number and the time section number, perform numerical recombination on the time period number, and merge the consecutive time periods corresponding to the same style for numbering to generate a style section number sequence; S512: Invoke the style section number sequence and the track number information in the original music library. According to the position of the track in the sequence based on the time period it belongs to, judge the correspondence between the track number and the style number, and classify each track into the corresponding style identifier according to the time period it belongs to to generate a style-track matching group; S513: According to the track number and style identifier information recorded in the style-track matching group, standardize the field content in the format of the combination of the style number and the time period number, uniformly encode the style and time period information of each group of tracks, and generate a style music library field combination.
9. A construction system for a national flute music library, characterized in that, The system is used to implement the method for constructing a flute ethnic music library according to any one of claims 1-8. The system includes: The audio feature recognition module obtains the complete audio data of the flute ethnic performance track, divides the audio into continuous segments, detects the change trajectory of the timbre waveform in each segment of audio, counts the turning point values between adjacent wave peaks and wave valleys, and records the time position interval on the corresponding time axis to obtain a timbre change concentration section time index table; The melody structure recognition module extracts the main melody note sequence within the time period according to the time period recorded in the timbre change concentration section time index table, recognizes the frequency direction change, and marks and numbers the nodes where the directionality changes to obtain a melody trend node sequence table; The style segment comparison module invokes the node arrangement order recorded in the melody trend node sequence table, compares the node positions and orders with the melody structure path in the preset ethnic style sample, and counts the number of structure segments matching the ethnic style sample to obtain an ethnic style proportion label set; The style feature mapping module filters the timbre fluctuation samples under the style label according to the style label in the ethnic style proportion label set, invokes the timbre data of the corresponding segment, compares the fitting degrees of the fluctuation period, amplitude, and energy distribution density with the sample fluctuation segment, marks the same type of time periods and records the corresponding style to obtain a style section time feature mapping result; The music library construction module invokes the style label and time period information recorded in the style section time feature mapping result, matches the unique number of the real-time track, and combines the style label, the corresponding time period interval, and the track number into a standard field format to obtain a style music library field combination.
Citation Information
Patent Citations
Audio assisted creation method based on big data
CN109754773A
Method for automatically compiling accompaniment chords
CN111739491A
Song generation method and device, equipment and storage medium
CN112951184A
Music melody identification method and device for online education
CN119479590A
Melody retrieval system
US20070163425A1
Cited By
College music education creative course visualization method and system
CN121070228A
A high school music education creative course visualization method and system
CN121070228B