A flute national style piece library construction method and system
By carefully dividing and analyzing the audio data of flute repertoire, generating timbre change and melody structure indexes, calling ethnic style samples for comparison, and generating style feature mapping results, the problem of the existing technology that fails to effectively capture timbre and melody structure is solved, the classification and retrieval efficiency of music databases is improved, and the protection and dissemination of ethnic music is promoted.
Patent Information
- Application Number
- CN202510884572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-30
AI Technical Summary
When processing music information, existing technologies fail to effectively capture and analyze timbre changes and melodic structure, causing users to encounter difficulties in finding tracks with specific expressions or stylistic characteristics, affecting the depth of music education and academic research, and failing to effectively integrate and utilize the structural characteristics of national music styles, affecting cultural dissemination and protection.
By carefully dividing the audio data of the flute repertoire and counting the timbre waveform changes, a time indexing method for concentrated segments of timbre changes is generated, the concentrated segment indexing method for timbre changes in the corresponding segments is intercepted, a time index table of timbre changes is generated, the melody direction node indexing method is identified, a melody direction node sequence table is generated, the melody structure of the ethnic style sample repertoire is called, the ethnic style proportion label set is generated, the style feature change position is analyzed, the style segment time feature mapping result is generated, and the format is normalized in combination with the repertoire number information to generate a style music library field combination.
It achieves accurate capture and detailed analysis of the timbre and melodic structure of flute repertoire, improves the classification efficiency and accuracy of the music database, enhances the convenience and efficiency of users in retrieving and identifying repertoires of specific styles, and provides a highly organized and standardized music library information integration and retrieval framework.
Smart Images

Figure CN120388550B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of music information management, and in particular to a flute national music library construction method and system. BACKGROUND
[0002] The technical field of music information management aims to efficiently process, store and retrieve music-related information. The core content of this field involves the collection, classification, management and display of music data. By using various data labeling and processing techniques, the systematic integration and convenient access of music resources are ensured. Music information management technology focuses on optimizing the access process of music materials, improving the functionality of music databases and the interactivity of user interfaces to meet the retrieval and usage needs of different users.
[0003] Among them, the flute national music library construction method refers to the establishment of a national music library related to flute performance through certain technical means. This technical theme involves audio collection, transcription and storage of music works, as well as accurate classification and labeling of collected music data to ensure that each piece of data is organized according to attributes such as national characteristics, style types and performance techniques. This method also includes music data management, supporting effective retrieval and display of music, and user interface design, allowing performers to quickly filter music according to specific styles or difficulty.
[0004] The existing technology focuses on basic data collection and classification when dealing with music information, but lacks in capturing subtle musical expressions and deep structural features. This deficiency limits the functionality of music databases in supporting complex queries and providing detailed music analysis. The inability to capture and analyze timbre changes and melody structures makes it difficult for users to find music with specific expression methods or style characteristics, affecting the depth of music education and academic research. Traditional information management technology fails to effectively integrate and utilize the structural characteristics of national music styles, resulting in insufficient dissemination and utilization of specific cultural music resources in the context of global music education and cultural heritage, affecting the protection and promotion of music diversity. SUMMARY
[0005] The purpose of the present application is to solve the shortcomings of the prior art and to provide a flute national music library construction method and system.
[0006] In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme, a flute national music library construction method, comprising the following steps:
[0007] S1: Obtain the audio data of the flute music, divide the audio content into continuous paragraphs according to the time axis, extract the continuous change trajectory of the timbre waveform of the paragraph, and accumulate the number of timbre state transitions in each waveform to generate a timbre change concentrated segment time index table;
[0008] S2: Based on the time period marked in the timbre change concentrated period time index table, the main melody information in the corresponding segment is intercepted, the ascending and descending order and continuous change of the scale in the melody are identified, the nodes are divided according to the directionality of the scale direction, and a melody direction node sequence table is generated;
[0009] S3: According to the path structure in the melody direction node sequence table, the melody structure of the national style sample music is called, the structure is compared according to the node arrangement order, the segments with consistent structure in the sample are marked, and a national style proportion label set is generated;
[0010] S4: The national style label in the national style proportion label set is used to call the national style audio sample, the timbre fluctuation segment in the target music is compared, the same time period interval as the sample fluctuation segment is recorded on the time axis, the style feature change position in the music is analyzed, a style section time feature mapping result is generated, and the style of the flute music is identified and classified.
[0011] As a further scheme of the present application, the timbre change concentrated period time index table includes the turning point density, the timbre concentrated period start time, and the timbre concentrated period end time, the melody direction node sequence table includes the scale change order, the melody direction, and the path node arrangement order, the national style proportion label set includes the structure similar segment quantity, the style category proportion information, and the matching segment label, and the style section time feature mapping result includes the style type label, the style time period start position, and the style time period end position.
[0012] As a further scheme of the present application, the obtaining step of the timbre change concentrated period time index table is specifically:
[0013] S111: Obtain the audio data of the flute music, equally divide the audio according to each second, collect the timbre envelope line track data in each time period, record the audio sampling values of the start time point and the end time point, and generate a timbre waveform change sequence;
[0014] S112: Based on the continuous change sequence in the timbre waveform change sequence, the number of times of slope sign change between adjacent sampling points is calculated, the number of times of slope direction change in each segment of audio is counted, and the number of waveform turning points is obtained.
[0015] S113: According to the time period marked as dense distribution feature in the waveform turning point number, the start time and end time node of the corresponding paragraph are extracted, the corresponding relationship between the paragraph number and the time position interval is evaluated, and a timbre change concentrated period time index table is obtained.
[0016] As a further scheme of the present application, the obtaining step of the melody direction node sequence table is specifically:
[0017] S211: Based on the time period marked in the timbre change set time index table, the melody main line information in the corresponding segment is intercepted, each note value in the melody main line is sequentially arranged, the change trend of the scale is judged according to the relationship between adjacent notes, the rising and falling trend sequence of the melody scale in each time period is generated, and the melody rising and falling trend sequence is obtained;
[0018] S212: The melody rising and falling trend sequence is called, the node division is carried out according to the directionality of the scale, the paragraph with continuous change of directionality is extracted, the note at the time of reversing the change of directionality is identified as a division node, the node is numbered and marked according to the arrangement order in the melody, and the directionality division node sequence is obtained.
[0019] S213: The directionality division node sequence is called, the scale span, time span and direction switching frequency between adjacent nodes in the node sequence are identified, the fluctuation degree of the melody structure is calculated according to the change relationship between the indexes, and the melody trend node order table is obtained.
[0020] As a further scheme of the present application, the obtaining step of the ethnic style proportion label set is specifically:
[0021] S311: Based on the path structure in the melody trend node order table, the melody structure of the ethnic style sample music is called, the starting node and the ending node interval connected by each path are recorded, the melody structure segment in the sample music consistent with the number of path interval nodes is demarcated as a comparison reference section, the comparison reference section is extracted from the path structure respectively, and the path structure comparison section set is generated.
[0022] S312: According to the path structure comparison section set, the melody structure in the ethnic style sample music is compared, the melody segment consistent with the number of nodes, the pitch flow direction and the rhythm distribution is screened, each consistent segment corresponding to the ethnic style label is labeled, the number of consistent segments corresponding to the differential ethnic style label is counted, and the ethnic style segment statistical result is obtained.
[0023] S313: The ethnic style segment statistical result is called, the number of consistent segments corresponding to the ethnic style label is calculated according to the number of consistent segments corresponding to the ethnic style label, the number of consistent segments under each ethnic style label is combined, the difference value of ethnic style distribution is calculated, the deviation of the difference value and the number of segments corresponding to the ethnic style is identified, and the ethnic style proportion label set is obtained.
[0024] As a further scheme of the present application, the obtaining step of the style section time feature mapping result is specifically:
[0025] S411: Based on the ethnic style proportion label set, obtaining the timbre fluctuation segment in the audio sample, detecting the start time and the end time on the time axis, calculating the amplitude difference and time span difference between adjacent peaks, and obtaining the timbre fluctuation rhythm sequence;
[0026] S412: Call the timbre fluctuation rhythm sequence, compare the timbre fluctuation segments in the target song, select the fluctuation segments with the same rhythm distribution, and calculate the cross-matching degree between the timbre spectrum change value and the rhythm density value.
[0027] Calculate the rhythm fit strength value, select the target segment that is consistent with the rhythm sequence of the sample fluctuation segment, and obtain the style matching time interval;
[0028] S413: According to the style matching time interval, the start and end positions of the interval in the time axis are recorded, the corresponding ethnic style labels are marked, the distribution structure of the ethnic styles on the time axis is analyzed, and a style segment time feature mapping result is generated.
[0029] As a further solution of the present invention, the method further includes step S5:
[0030] S5: Based on the style and time period information in the style segment time feature mapping result, combined with the track number information, the time segment and style identifier are formatted and normalized, and the combination is used as a music library entry field item to generate a style music library field combination;
[0031] The style music library field combination includes a track number, a standardized style identifier, and style section index information.
[0032] As a further solution of the present invention, the steps for obtaining the style music library field combination are specifically as follows:
[0033] S511: Based on the style and time period information in the style segment time feature mapping result, extract the style identifier and the corresponding time period number, evaluate the correspondence between the style number and the time segment number by matching the sequential number of the style field with the start and end marks of the time period, reorganize the time period numbers, merge the numbers of consecutive time periods corresponding to the same style, and generate a style segment number sequence;
[0034] S512: The style segment number sequence and the number information of the tracks in the original music library are called, and the track number and style number are matched according to the position of the time period of the track in the sequence. Each track is classified into a corresponding style identifier according to the time period, and a style track matching group is generated;
[0035] S513: Based on the track number and style identification information recorded in the style track matching group, the field content is normalized according to the format of the style number and time period number combination, the style and time period information of each group of tracks are uniformly encoded, and a style music library field combination is generated.
[0036] The system for constructing a flute music library of national characteristics is used to execute the method for constructing a flute music library of national characteristics, and the system comprises:
[0037] The audio feature recognition module obtains the complete audio data of the flute folk music, divides the audio into continuous segments, detects the timbre waveform change trajectory in each audio segment, counts the turning point values between adjacent peaks and troughs, records the time position intervals on the corresponding time axis, and obtains a time index table of the timbre change concentrated segments;
[0038] The melody structure recognition module extracts the main melody note sequence within the time period recorded in the timbre change concentrated time index table, identifies the frequency direction change, and obtains a melody direction node sequence table by marking and numbering the nodes where the directionality changes;
[0039] The style segment comparison module calls the node arrangement order recorded in the melody trend node sequence table, compares the node position and order with the melody structure path in the preset ethnic style sample, counts the number of structural segments matching with the ethnic style sample, and obtains an ethnic style proportion label set;
[0040] The style feature mapping module selects the timbre fluctuation samples under the style labels in the ethnic style proportion label set, calls the timbre data of the corresponding paragraph, compares the fit with the fluctuation period, amplitude, and energy distribution density of the sample fluctuation segment, marks the same time period and records the corresponding style, and obtains the style segment time feature mapping result;
[0041] The music library construction module calls the style label and time period information recorded in the style segment time feature mapping result, matches the unique number of the real-time track, combines the style label, corresponding time period interval and track number into a standard field format, and obtains the style music library field combination.
[0042] Compared with the prior art, the advantages and positive effects of the present invention are:
[0043] In the present invention, by carefully dividing the audio data of the flute repertoire and counting the changes in timbre waveform, the accuracy and meticulousness of music information management are significantly improved. The characteristic segments with dense distribution of waveform transitions are indexed, which helps to accurately capture the key timbre changes in the repertoire and provide richer and more specific data support for subsequent melody structure analysis. The ascending and descending order and continuous changes in the melody can be carefully identified, and the nodes can be further divided according to the direction of the scale trend, which not only increases the multi-dimensional visualization of the information, but also makes the structural characteristics of the melody clearer. Calling ethnic style sample repertoires for structural comparison and counting the fragments with consistent styles can effectively mark out music fragments that conform to specific ethnic styles, thereby enhancing the classification efficiency and accuracy of the database. Through style marking and time feature mapping, a highly organized and standardized framework is provided for the integration and retrieval of music library information, greatly improving the convenience and efficiency of users in retrieving and identifying specific style repertoires. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0045] Figure 2 This is a flow chart for obtaining a time index table for a concentrated period of timbre change in the present invention;
[0046] Figure 3 This is a flow chart for obtaining a melody trend node sequence table in the present invention;
[0047] Figure 4 This is a flow chart for obtaining the ethnic style proportion label set in the present invention;
[0048] Figure 5 This is a flow chart for obtaining the temporal feature mapping results of the style segments in the present invention;
[0049] Figure 6 This is a flow chart for obtaining the style music library field combination in the present invention. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0051] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0052] See also Figure 1 The present invention provides a technical solution, a method for constructing a flute folk music library, comprising the following steps:
[0053] S1: Acquire audio data of a flute piece, divide the audio content into continuous segments along the time axis, extract the continuous change trajectory of the timbre waveform of each segment, accumulate and count the number of timbre state transitions in each waveform segment, record the time position in the segment where the number of transitions reaches a dense distribution feature, and generate a time index table of concentrated timbre change segments;
[0054] S2: Based on the time periods marked in the time index table of the concentrated segments of timbre changes, the main melody line information in the corresponding segments is intercepted, the ascending and descending order of the scale in the melody and the continuous changes are identified, nodes are divided according to the directionality of the scale trend, and the reversal behavior after the same direction change is defined as a path node. The continuous arrangement order of the path nodes is recorded, the fluctuation pattern of the melody structure is identified, and a melody trend node sequence table is generated;
[0055] S3: Based on the path structure in the melody trend node sequence table, the melody structure of the ethnic style sample music is called, and the structure is compared according to the node arrangement order. The fragments with consistent structure in the sample are marked, and the number of matching fragments under the ethnic style category is counted to generate the ethnic style ratio label set;
[0056] S4: Using ethnic style labels from the ethnic style proportion label set, retrieve ethnic style audio samples, compare the timbre fluctuation segments in the target track, record the time intervals on the time axis that are similar to the sample fluctuation segments, assign style tags to the time intervals, analyze the location of style feature changes in the track, and generate style segment time feature mapping results;
[0057] S5: Based on the style and time period information in the style segment time feature mapping result, combined with the track number information, the time segment and style identifier are formatted and normalized, and the combination is used as a music library entry field item to generate a style music library field combination;
[0058] The timbre change concentrated segment time index table includes the turning point density, the timbre concentrated segment start time, and the timbre concentrated segment end time, the melody trend node order table includes the scale change order, the melody direction, and the path node arrangement order, the national style proportion label set includes the structure similar segment quantity, the style category proportion information, and the matching segment label, the style segment time feature mapping result includes the style type label, the style time period start position, and the style time period end position, and the style music library field combination includes the music number, the standardized style identifier, and the style segment index information.
[0059] Referring to Figure 2 The obtaining step of the timbre change concentrated segment time index table is specifically as follows:
[0060] S111: Obtain the audio data of the flute music piece, equally divide the audio according to each second, collect the timbre envelope line trajectory data in each time period, record the audio sample values of the start time point and the end time point, and generate a timbre waveform change sequence;
[0061] Obtain the audio data of the flute music piece, the audio is stored in a standard WAV format, the sampling rate is 44100 Hz, 16-bit single-channel recording is adopted, 44100 audio sample points are collected per second, the data is divided into continuous time periods at an interval of 5 seconds, the total number of sample points of each period of data is set to 220500 points, each period of data is divided into 432 frames, each frame has a size of 512 sample points, the frames are processed in a sliding manner with an overlap of 256 points between frames, the maximum amplitude value of the envelope line is calculated by reading the sample value array of each frame, the maximum amplitude is used as the representative value of each frame, and the envelope line change trajectory formed by the 432 representative values in the entire time period is sequentially constructed. For the data with a paragraph number of 1, the start time is recorded as 0 seconds and the end time is recorded as 5 seconds, and the corresponding audio has 220500 sample points, and the envelope trajectory is a change sequence containing 432 amplitude nodes. In actual processing, the timbre waveform change sequence of the time period is constructed, each value in the sequence represents the overall energy intensity of the audio signal in the time frame, and a complete change trajectory array is formed, so that subsequent turning point judgment and quantity statistics can be performed. The processing operation is sequentially performed on all time periods, and is a basic array for subsequent judgment of the timbre change region, and the timbre waveform change sequence is generated.
[0062] S112: Based on the continuous change sequence in the timbre waveform change sequence, the number of times of slope sign change between adjacent sample points is calculated, the number of times of slope direction change in each period of audio is counted, and the number of waveform turning points is obtained.
[0063] The average amplitude value sequence is extracted by frame, and the sign difference of the amplitude values of any two adjacent frames is calculated. If the amplitude value increases or decreases in a positive or negative direction, that is, there is a transition from positive to negative or negative to positive between the average value of the original frame and the next frame, it is considered a waveform direction turning point. The number of positive and negative changes in this time period is the number of waveform turning points in this section. In the actual data, taking paragraph number 1 as an example, it contains 432 frames. 431 groups of adjacent frames are compared in sequence. If a slope change occurs, that is, the difference sign changes from positive to negative or negative to positive, the number of turning points is increased by 1, and the total change is counted. The number of change points is 12, 18 change points are identified in paragraph number 2, 9 in paragraph number 3, 22 in paragraph number 4, and 15 in paragraph number 5. The dense distribution threshold is set to 16. Paragraphs below this value are regarded as general fluctuations, and those above or equal to this value are regarded as intensive fluctuation segments. Based on this judgment condition, the two time periods of paragraph numbers 2 and 4 are selected to meet the characteristics of frequent changes in timbre state. They are used as subsequent extraction targets. The statistics of the number of waveform turning points are generated according to the inter-frame difference direction judgment mechanism and recorded in the array to obtain the number of waveform turning points.
[0064] S113: extracting the start and end time nodes of the corresponding paragraphs based on the time periods identified as densely distributed features in the waveform transition count, evaluating the correspondence between the paragraph numbers and the time position intervals, and obtaining a time index table of the concentrated segments of timbre changes;
[0065] After screening the paragraph numbers whose number of transitions meets the timbre transition density threshold, extract the start time and end time corresponding to the paragraph number, map the time period and number one by one, and construct a bidirectional index table of paragraph numbers and time position intervals. Taking paragraph number 2 as an example, its start time is 5 seconds and end time is 10 seconds. Paragraph number 4 starts at 15 seconds and ends at 20 seconds. This type of time interval comes directly from the index identifier and sampling duration setting in the audio segmentation operation. Therefore, the index value is the original sampling division identifier rather than the secondary extraction. At the same time, the timbre change trajectory sequence number information corresponding to the segment is retained, indicating that the data in this time period can be used for timbre feature concentration comparison, including paragraph number, start time, end time, and identification field of whether it belongs to a high-density change segment. This data table will be used as the screening benchmark source for time intervals in subsequent melody main line extraction operations to obtain a timbre change concentration segment time index table.
[0066] See also Figure 3 ,The specific steps for obtaining the melody direction node sequence table are:
[0067] S211: Based on the time period marked in the timbre change centralized section time index table, the main melody information in the corresponding segment is intercepted, each note value in the main melody is sequentially arranged, the change trend of the scale is judged according to the relationship between adjacent notes, the rising and falling trend sequence of the melody scale in each time period is generated, and the melody rising and falling trend sequence is obtained;
[0068] The start and end time nodes corresponding to each time period in the index table are determined, the index range of a certain segment is set to 1.5 seconds to 3.8 seconds, the original melody audio data in the range is intercepted, the main melody of the audio segment is extracted, the frequency curve of the main melody is obtained by mainly using the pitch detection method of each frame, the fundamental frequency value of each frame is uniformly converted into MIDI value to represent the pitch level, such as 440Hz mapped to MIDI value 69 (i.e. A4), the frames are arranged in time sequence to form a pitch sequence, and the size relationship between adjacent note pitch values is judged to identify the rising or falling trend, such as the pitch of the 3rd frame is 62 and the pitch of the 4th frame is 64, which is an upward trend, a rising and falling mark is formed between every two frames, the entire paragraph is processed in this way to form a complete trend sequence, if there is no change in the continuous frame pitch (set the MIDI value to 66 for 3 consecutive frames), it is marked as parallel trend, which can be removed or classified by post-processing, through the above operation, the entire melody segment is divided into continuous rising and falling direction segments, set up-rise-up-down-down-up, and further unified and arranged into a structured data sequence and numbered and marked for subsequent analysis, wherein the judgment of the rising and falling trend should consider the error tolerance range, which is ±0.5 MIDI value within the same level pitch, therefore, the error tolerance threshold needs to be set, which is ±0.3 in this example, when the pitch value varies within this range, it is marked as no change, and through this way, slight tremolo or pitch floating in the melody can be effectively processed to obtain the melody rising and falling trend sequence.
[0069] S212: Call the melody rising and falling trend sequence, divide the nodes according to the directionality of the scale, extract the paragraphs with continuous directional changes, identify the notes when the directional change reverses as the division nodes, number and mark the nodes according to their arrangement order in the melody, and obtain the directional division node sequence;
[0070] According to the continuity of the direction of each trend, identify the nodes where the direction changes as switching points, read the marks in the rising and falling trend sequence, set a sequence as "rising, rising, falling, falling, rising", and the position where "rising" turns to "falling" (between the 2nd and 3rd positions) and the position where "falling" turns to "rising" (between the 4th and 5th positions) are used as directional switching nodes. The note index position corresponding to the node and its corresponding MIDI value and time point are located, and further numbering is performed. Starting from the first directional switching point, they are marked as P1, P2, P3, etc. to reflect their arrangement order in the melody. In the example, if the first switching point is 2.4 seconds and the second switching point is 3.1 seconds, they are marked as P1 and P2 respectively, and their original pitch M is recorded. IDI value (such as 65 and 67 respectively) is used for subsequent path construction and melody structure modeling. In specific operations, it is necessary to exclude micro-movements caused by non-directional changes, such as the up and down floating trend caused by detection errors. In the trend sequence, "up-down-up" appears, but the intermediate decline amplitude is less than 1 MIDI value. It is judged as a non-valid reversal trend and the node is excluded. It is necessary to set an amplitude threshold judgment. The value refers to 10% of the average value of the melody span. If the average span is 8, the threshold is set to 0.8. This setting is based on the natural change law of the melody and is determined in combination with statistical experience values to identify valid reversal nodes. Its content includes the time point, pitch value, trend turning category and number index of each node in the overall sequence to obtain a directional division node sequence.
[0071] S213: Call the directionality division node sequence to identify the scale span, time span, and direction switching frequency between adjacent nodes in the node sequence. Based on the change relationship between the indicators, the formula is used:
[0072] ;
[0073] Calculate the degree of fluctuation of the melody structure and obtain the sequence table of melody trend nodes;
[0074] in, Represents the degree of fluctuation of the melody structure, Representative The scale span value between node pairs, Representative The time span value between node pairs, Representative The switching frequency value of the node pair direction, Representative The continuous continuation value of the node pairs in the same direction, is the total number of node pairs;
[0075] Extract the relative position and direction changes between nodes, extract the scale span value between adjacent nodes according to the node index sequence, and obtain the scale difference by numerically comparing the pitch values of two consecutive nodes. Set the pitch of node 1 to C4 (MIDI value 60) and node 2 to A4 (MIDI value 69), then its scale span value is 9. After obtaining the scale span value, calculate the time span value based on the time index of the two nodes. Set the time of node 1 to 2.1 seconds and the time of node 2 to 3.0 seconds, and the time span value is 0.9 seconds. The direction switching frequency value is calculated by identifying the number of direction changes in the continuous node sequence as a percentage of the total number of node pairs. For example, if there are 3 direction reversals in 5 node pairs, the frequency is 0.6. The direction continuation length is collected as the same direction continuation value, which is achieved by counting the number of consecutive scale changes with the same direction between two nodes. Set that there are two notes E4 and F4 between C4 and G4, both in the ascending direction, then the same direction continuation value is 3. Combined with the above parameters, substitute into the formula:
[0076] To illustrate the parameter setting process, the following three node pair parameters are extracted and described in array form:
[0077] Array of scale span values ;
[0078] Array of timespan values ;
[0079] Direction reversal frequency array ;
[0080] Array of continuation values in the same direction ;
[0081] The formula calculation based on the above array is as follows:
[0082] Molecular part:
[0083] ;
[0084] Denominator:
[0085] ;
[0086] Substitute the parameters into the formula:
[0087] ;
[0088] The result shows that the degree of fluctuation under the current melody node structure is 0.457, which represents the average variation of the overall melody trend and can be further used as a basis for generating a melody trend node sequence table.
[0089] See also Figure 4,The specific steps for obtaining the ethnic style proportion label set are:
[0090] S311: Based on the path structure in the melody trend node sequence table, the melody structure of the ethnic style sample music is called, the start node and end node interval connected by each path is recorded, and the melody structure segments in the sample music with the same number of nodes as the path interval are delineated as comparison reference segments. The comparison reference segments are respectively extracted from the path structure to generate a path structure comparison segment set;
[0091] The path structure in the melody trend node sequence table corresponds to a sequence combination of multiple melody structures. In practice, it can be sampled from the transition nodes of the melody paragraphs in traditional music of a certain region. Taking the melody nodes of traditional ethnic minority music as an example, path 1 can be recorded as [1, 3, 5, 7], path 2 as [2, 4, 6], and path 3 as [1, 4, 6, 8]. When calling the melody structure of the ethnic style sample music, a mapping table is established to map the path node sequence to the specific note sequence in the melody paragraph of the sample music. The extraction process uses the starting and ending nodes connected by each path as a reference to locate the corresponding note position in the sample melody. The continuous note segments with the same number of corresponding path nodes in the sample are delineated as the comparison reference segments. Assuming that path 1 corresponds to 4 nodes, the melody segments with node numbers between 1 and 7 are searched in the sample and the corresponding note sequences of length 4 are extracted. The same process is then performed on path 2 and path 3. The melody segment sets extracted from the path structures are unified into a comparison reference segment set. The sequence correspondence between the node sequence and the note structure is evaluated to generate a path structure comparison segment set.
[0092] S312: Comparing the segment set based on the path structure, comparing the melody structure of the ethnic style sample songs, selecting melody segments that are consistent with the melody structure in terms of the number of nodes, pitch flow, and rhythm distribution, labeling each matching segment with an ethnic style label, and counting the values of the matching segments corresponding to the differentiated ethnic style labels to obtain ethnic style segment statistics;
[0093] When comparing the sample melody structures, it is necessary to determine the three parameters according to the number of nodes, pitch flow and rhythm distribution. The node value of each comparison segment is recorded and directly compared with the number of nodes in the sample segment to determine whether it matches. Secondly, the pitch flow direction is determined by whether the continuous pitch difference is within the same sign range to determine whether the flow direction is consistent. Set the comparison segment pitch sequence to [200, 210, 220] and the sample melody segment to [198, 208, 218]. If the pitch difference of the two is positive, the flow direction is consistent. If it is [220, 210, 200], it means a descending flow direction, which needs to be determined. Consistent with the direction of the comparison segment, the rhythm distribution uses the standard deviation of adjacent note values not exceeding 5ms as the matching standard. The segment rhythm is set to [250ms, 260ms, 255ms], and the sample segment rhythm is [248ms, 262ms, 257ms] to determine that the rhythm is consistent. Three parameter judgments are performed on each sample segment, and those that pass are marked as matching segments. The matching segments in the sample are counted and classified according to their corresponding ethnic style labels to generate quantitative statistics, such as 5 matching segments under style A, 7 segments under style B, and 6 segments under style C, and the statistical results of ethnic style segments are obtained.
[0094] S313: Call the statistical results of ethnic style segments, and use the formula:
[0095] ;
[0096] Calculate the difference value of ethnic style distribution, identify the deviation between the difference value and the number of fragments corresponding to the ethnic style, and obtain the ethnic style proportion label set;
[0097] in, Representative The difference value of the ethnic style distribution of the ethnic styles, Indicates the The first among the ethnic styles The number of node distributions of the path segments, Indicates the The first among the ethnic styles The average pitch of the path segments, Indicates the The average pitch of the reference segment in the path structure, Indicates the The total number of rhythms in the reference section of the path structure, is the number of paths;
[0098] The number of matching segments, node density of each path segment, total number of rhythms, and pitch distribution characteristic values of each ethnic style were recorded. After extracting the segment data corresponding to each style label, the distribution intensity difference in the sample sequence was numerically calculated using the formula:
[0099] Substitute the following parameter values for style A:
[0100] Number of path segments: 5;
[0101] Node distribution number: [9, 10, 8, 9, 9];
[0102] Average pitch: [210.0, 212.5, 211.0, 209.5, 210.5];
[0103] The mean pitch of the comparison segment is: [200, 200, 200, 200, 200];
[0104] Total number of rhythms: [12, 12, 12, 12, 12];
[0105] The substitution process is as follows:
[0106] molecular:
[0107] ;
[0108] Denominator:
[0109] ;
[0110] Substitute into the formula for calculation:
[0111] ;
[0112] The results show that the difference value of the ethnic style distribution is 0.388, which represents the degree of deviation between the number of segments of each style label and the expected number in the overall data set. It can help identify which ethnic styles are over-represented or under-represented in the data set. The benefit of the formula is that by introducing the product relationship between the square of the node density and the pitch deviation, the structural stability reflection is introduced into the difference measurement, and the quantitative relationship between the difference value and the segment distribution density is analyzed to obtain the ethnic style proportion label set.
[0113] See also Figure 5 ,The steps for obtaining the style segment time feature mapping results are as follows:
[0114] S411: Based on the ethnic style proportion label set, obtain the timbre fluctuation segment in the audio sample, detect the start time and the end time on the time axis, calculate the amplitude difference and time span difference between adjacent peaks, and obtain the timbre fluctuation rhythm sequence;
[0115] The original audio signal sequence is extracted from the sample and linearly segmented along the time axis. Every 1 second is taken as the analysis window. A track with a total length of 180 seconds is divided into 180 segments. The amplitude signal is collected in each window, and the waveform is sampled. The continuously rising and falling peaks and troughs are extracted as the boundary points of the fluctuation segment. The amplitude change value between each peak and the adjacent trough is measured. If the change exceeds 3dB, it is considered a significant fluctuation. At the same time, the time span between the peak and the trough is recorded. If the span is less than 0.5 seconds, it is marked as a short-period fluctuation segment. Furthermore, the amplitude difference and period data are uniformly normalized to generate a unified rhythm contour vector sequence. The rhythm contour corresponding to a sample at the 60th second is set to [3.1, 0.6], where 3.1 is the maximum amplitude difference of this segment and 0.6 is the fluctuation period. It can be used for rhythm matching with the target track to obtain a timbre fluctuation rhythm sequence.
[0116] S412: Call the timbre fluctuation rhythm sequence, compare the timbre fluctuation segments in the target song, select the fluctuation segments with the same rhythm distribution, and use the formula based on the cross-matching degree between the timbre spectrum change value and the rhythm density value:
[0117] ;
[0118] Calculate the rhythm fit strength value, select the target segment that is consistent with the rhythm sequence of the sample fluctuation segment, and obtain the style matching time interval;
[0119] in, Represents the rhythm fit strength value, Represents the rhythm amplitude difference value of the target segment in the specified national style, Represents the timbre frequency domain gradient value, represents the number of pitch jumps, Represents the rhythm center of gravity offset, Representative sample The fitting degree of the records, is the number of records;
[0120] According to the time distribution characteristics of the sample sequence, the target track is aligned second by second and the sliding window matching process is performed. The target track is scanned at 1 segment per second, and the amplitude difference and frequency gradient change value of the current segment are extracted. The target segment is set to The amplitude difference of the segment is , the frequency changes to , the fluctuation is more intense, and the number of pitch jumps is set based on experience. , and detect the rhythm center position offset , for the first To style matching paragraphs, and calculate the timbre similarity between them and the target paragraph, recorded as 、 、 , after averaging, we get approximately , which is the benchmark for subsequent comparison;
[0121] ;
[0122] Substitute the data and calculate as follows:
[0123] The first step is to calculate the square root part:
[0124] ;
[0125] Step 2: Multiply and divided by :
[0126] ;
[0127] The third step is to find the average similarity of samples:
[0128] ;
[0129] The fourth step is to substitute and find the absolute value difference:
[0130] ;
[0131] The results show that the rhythm fit strength value is 7.2933, which is used to measure the degree of match between the rhythm structure of a music clip and the rhythm structure of another target music clip. This value can be used to evaluate the rhythmic similarity and fit of the music clips. According to the preset judgment benchmark, if the value exceeds 5, it is considered that the target clip and the sample paragraph have a matching association in the national style. Therefore, the current paragraph is included in the candidate matching interval. Multiple intervals that meet the strength threshold are extracted from the target track as candidate paragraphs to obtain the style matching time interval.
[0132] S413: Matching the time interval according to the style, recording the start and end positions of the interval in the time axis, marking the corresponding ethnic style label, analyzing the distribution structure of the ethnic style on the time axis, and generating a style segment time feature mapping result;
[0133] The start and end time of each matching interval is extracted and labeled in seconds. The absolute time information of the interval is matched by combining the scale information of the target track timeline. Set the current extracted matching interval as 32s to 34s, then the interval is marked as [32s, 34s]. According to the ethnic style label pointed to in the source sample, assign this interval to the "Tibetan" style identifier. For example, the 45s to 47s segment corresponds to the "Yi" style. Process all matching intervals in turn to form a mapping set of multiple time periods and ethnic styles. To avoid label conflicts, sort the rhythm fit strength of overlapping segments, and prefer to keep the segment with higher strength value. At the same time, in this process, the overlapping segments are merged. When the 60s to 62s and the 61s to 63s both meet the style matching and are marked as "Miao" and "Zhuang" respectively, compare their fit strength values. If the Miao paragraph is 6.8 and the Zhuang is 7.2, keep the latter and recombine the entire paragraph as [61s, 63s]. Delete the lower fit degree item in the overlapping part. Store all paragraph information as a key-value pair structure, where the key is the time period and the value is the ethnic style label. This is used for subsequent visualization output of the style transformation graph to get the style segment time feature mapping result.
[0134] Please refer to Figure 6 The style library field combination acquisition step is as follows:
[0135] S511: Based on the style and time period information in the style segment time feature mapping result, extract the style identifier and corresponding time period number. Match the order number of the style field with the start and end markers of the time period to evaluate the correspondence between the style number and the time period number. Recombine the time period number value, merge the continuous time period numbers corresponding to the same style, and generate a style segment number sequence.
[0136] It should be clear that the extraction method of "style identifier" and "time period number" is to disassemble each element in the two-dimensional array structure in the feature mapping result into a mapping pair in the form of , where represents the style identifier, represents the corresponding time period number. Each pair of data can be regarded as the main style exhibited in a certain time period in the library. Set in the sample library, style appears in time period , , , style appears in time period , , and the time period numbers of the same style are aggregated into a set, such as , , perform number merging processing on the time period number sequence in each style set in turn, and the merging rule is: if the adjacent time period numbers are continuous, they are merged into one number segment, such as the sequence Merge into segment 1; if the time period is not continuous, it will be split into multiple style segment numbers. In this process, it is necessary to judge the continuity of the number, that is, to judge any satisfy , then merge, otherwise start a new numbering, the merged result can form the following style segment number sequence: , In actual implementation, the input music library is set to include 10 time periods, each of which is 30 days long. The styles are respectively , the execution process is as follows: After extracting the style-time period pairs, build a list , judge the continuity of the number, 1 to 3 are continuous time periods, merged into a number segment , 5 to 6 are another continuous segment, merged into a numbered segment ,Here the number reorganization is done by traversing the time segment number list, setting up a ,temporary pointer to record the difference between the previous number and the current number, ,and when the difference is 1, recording the length of the continuous sequence, and creating a new ,numbering segment when the difference is greater than 1 and marking the sequence count ,increment by one. A one-to-one mapping relationship between the style number and ,the time segment number is used to assist the recording.
[0137] S512: The style segment number sequence and the track number information in the original music library are called, and the track number and style number are matched according to the position of the time period of the track in the sequence. Each track is classified into a corresponding style identifier according to the time period, and a style track matching group is generated;
[0138] It is necessary to establish a time period positioning mechanism for tracks, and set a corresponding timestamp field for each track in the original music library. , the timestamp needs to be converted into the aforementioned time period number The specific operation is to set the time window length Days, taking the start time of the music library as the benchmark value , calculate the time period number of each track as , where the time unit is "day", if the track The timestamp of is the 62nd day, so the time period number is After obtaining all the track numbers and their corresponding time period numbers, compare the corresponding time period number range in the style segment number sequence, such as the track In time period , and the style section Coverage time period , then the track is classified into the style The above classification process is performed on the music pieces, i.e. the time period number of each music piece is compared with the range of the style section number, and if the time period number of the music piece belongs to the style and the set belongs to the time period , the music piece number and the style identifier are stored correspondingly. The judgment operation is specifically setting a loop variable to read the music piece number and the time period number, and then performing conditional matching to determine whether the time period is located in the time period set of a style number. If yes, the mapping pair of the music piece and the corresponding style is added to the matching group to form a set of style-music piece matching groups.
[0139] S513: According to the music piece number and the style identifier information recorded in the style-music piece matching group, the style and time period information of each group of music pieces are uniformly encoded according to the format specification of the combination of the style number and the time period number to generate a combination of style music library fields;
[0140] When field standardization is performed according to the music piece number and the style identifier information recorded in the style-music piece matching group, a uniform encoding mechanism needs to be constructed, i.e. a combination tag of the style number and the time period number is given to each group of music pieces. The format can be set as "F number-T number", such as the music piece belongs to the style and belongs to the time period , the field identifier is "F1-T3". The style identifier of each music piece needs to be determined and the time period number is obtained by reverse calculation of the timestamp. The number, style identifier and time period number of each music piece are called to form a uniform identification field through string splicing. This process needs to traverse the music piece set, generate a corresponding combination field for each music piece and write it into a standard field set. If the music piece is of the style and the time period , the combination field is "F2-T5". The combination field constitutes a new field set. During the execution of this operation process, duplicate identification needs to be avoided, so a combination field uniqueness verification mechanism needs to be established, i.e. after generating the combination field, a de-duplication process is performed on the field list to determine whether the field already exists in the identification set. If yes, the writing is skipped. If not, it is added to the list. It not only contains the music piece number, but also maps the style and time period attributes. Each music piece is standardized and labeled with a unique style-time period combination code to form a combination of style music library fields.
[0141] The flute national music library construction system is used to execute the above flute national music library construction method. The system comprises:
[0142] The audio feature recognition module obtains complete audio data of a flute national performance piece, divides the audio into continuous paragraphs, detects the timbre waveform change track in each audio, counts the turning point value between adjacent wave crests and troughs, records the time position interval on the corresponding time axis, and obtains a timbre change concentrated paragraph time index table;
[0143] The melody structure recognition module extracts the main melody note sequence in the time period according to the time period recorded in the timbre change concentrated paragraph time index table, recognizes the frequency direction change, marks and numbers the nodes where the directionality changes, and obtains a melody direction node sequence table;
[0144] The style segment comparison module calls the node arrangement sequence recorded in the melody direction node sequence table, compares the node position and sequence with the melody structure path in the preset national style sample, counts the number of structure segments matched with the national style sample, and obtains a national style proportion label set.
[0145] The style feature mapping module filters the timbre fluctuation samples under the style label according to the style label in the national style proportion label set, calls the timbre data of the corresponding paragraph, compares the fitting degree of the fluctuation period, the amplitude amplitude, and the energy distribution density with the sample fluctuation segment, marks the same time period and records the corresponding style, and obtains a style section time feature mapping result.
[0146] The music library construction module calls the style label and time period information recorded in the style section time feature mapping result, matches the unique number of the real-time piece, combines the style label, the corresponding time period interval, and the piece number into a standard field format, and obtains a style music library field combination.
[0147] The above is only a preferred embodiment of the present application, and does not limit the form of the present application, any skilled person in the art can use the disclosed technical content to make changes or modifications into equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made according to the technical essence of the present application to the above embodiments without departing from the technical solution content of the present application still belongs to the protection scope of the present application technical solution.
Claims
1. A method for constructing a flute folk music library, characterized in that: The following steps are involved: S1: Acquire audio data of a flute piece, divide the audio content into continuous segments along the time axis, extract the continuous change trajectory of the timbre waveform of each segment, accumulate and count the number of timbre state transitions in each waveform segment, and generate a time index table of the timbre change concentration segment; S2: Based on the time periods marked in the time index table of the concentrated segments of the timbre change, intercept the melody main line information in the corresponding segments, identify the ascending and descending order and continuous changes of the scale in the melody, divide the nodes according to the directionality of the scale trend, and generate a melody trend node sequence table; S3: Based on the path structure in the melody trend node sequence table, the melody structure of the ethnic style sample music is called, the structure is compared according to the node arrangement order, and the segments with the same structure in the sample are marked to generate an ethnic style proportion label set; S4: Using the ethnic style labels in the ethnic style proportion label set, retrieving ethnic style audio samples, comparing the timbre fluctuation segments in the target piece, recording the time intervals similar to the sample fluctuation segments on the time axis, analyzing the change positions of the style characteristics in the piece, generating style segment time feature mapping results, and identifying and classifying the style of the flute piece; S5: Based on the style and time period information in the style segment time feature mapping result, combined with the track number information, the time segment and style identifier are formatted and normalized, and the combination is used as a music library entry field item to generate a style music library field combination; S511: Based on the style and time period information in the style segment time feature mapping result, extract the style identifier and the corresponding time period number, evaluate the correspondence between the style number and the time segment number by matching the sequential number of the style field with the start and end marks of the time period, reorganize the time period numbers, merge the numbers of consecutive time periods corresponding to the same style, and generate a style segment number sequence; S512: The style segment number sequence and the number information of the tracks in the original music library are called, and the track number and style number are matched according to the position of the time period of the track in the sequence. Each track is classified into a corresponding style identifier according to the time period, and a style track matching group is generated; S513: Based on the track number and style identification information recorded in the style track matching group, the field content is normalized according to the format of the style number and time period number combination, the style and time period information of each group of tracks are uniformly encoded, and a style music library field combination is generated.
2. The method for constructing a flute folk music library according to claim 1, wherein: The timbre change concentrated segment time index table includes turning point density, timbre concentrated segment start time, and timbre concentrated segment end time; the melody direction node sequence table includes scale change order, melody direction, and path node arrangement order; the ethnic style proportion label set includes the number of structurally similar segments, style category proportion information, and matching segment annotations; the style segment time feature mapping result includes style type mark, style period start position, and style period end position.
3. The method for constructing a flute folk music library according to claim 1, wherein: The steps for obtaining the time index table of the timbre change concentration section are specifically as follows: S111: Acquire audio data of a flute piece, divide the audio into equal time intervals per second, collect timbre envelope trajectory data within each time interval, and record audio sampling values at the start and end time points to generate a timbre waveform change sequence; S112: Calculating the number of slope sign changes between adjacent sampling points based on the continuous change sequence in the timbre waveform change sequence, and counting the number of slope direction changes in each audio segment to obtain the number of waveform turning points; S113: Extract the start and end time nodes of the corresponding paragraphs according to the time periods marked as densely distributed features in the waveform transition number, evaluate the correspondence between the paragraph numbers and the time position intervals, and obtain a time index table of the concentrated segments of the timbre changes.
4. The method for constructing a flute folk music library according to claim 3, wherein: The steps for obtaining the melody trend node sequence table are specifically as follows: S211: Based on the time periods marked in the time index table of the timbre change concentration, extract the melody main line information in the corresponding segment, sequentially arrange the values of each note in the melody main line, determine the change trend of the scale based on the relationship between the values of adjacent notes, and generate a rising and falling trend sequence of the melody scale in each time period to obtain a melody rising and falling trend sequence; S212: Calling the melody ascending and descending trend sequence, dividing nodes according to the scale direction, extracting sections with continuous direction changes, identifying notes when the direction changes reverse as division nodes, and numbering and marking the nodes according to their arrangement order in the melody to obtain a directional division node sequence; S213: Calling the directional division node sequence, identifying the scale span, time span and direction switching frequency between adjacent nodes in the node sequence, calculating the degree of melody structure fluctuation based on the change relationship between the indicators, and obtaining the melody direction node sequence table.
5. The method for constructing a flute folk music library according to claim 4, wherein: The steps for obtaining the ethnic style proportion label set are specifically as follows: S311: Based on the path structure in the melody trend node sequence table, the melody structure of the ethnic style sample music is called, the start node and end node interval connected by each path is recorded, and the melody structure segments in the sample music with the same number of nodes as the path interval are delineated as comparison reference segments. The comparison reference segments are respectively extracted from the path structure to generate a path structure comparison segment set; S312: Comparing the melody structure of the ethnic style sample music with the path structure comparison segment set, selecting melody segments that are consistent with the melody structure in terms of the number of nodes, pitch flow, and rhythm distribution, labeling each matching segment with an ethnic style label, and counting the matching segment values corresponding to the differentiated ethnic style labels to obtain ethnic style segment statistics; S313: Call the statistical results of the ethnic style fragments, calculate the ethnic style distribution difference value based on the number of matching fragments corresponding to the ethnic style label and the number of matching fragments under each ethnic style label, identify the deviation between the difference value corresponding to the ethnic style and the number of fragments, and obtain the ethnic style proportion label set.
6. The method for constructing a flute folk music library according to claim 5, wherein: The steps for obtaining the style segment time feature mapping result are specifically as follows: S411: Based on the ethnic style proportion label set, obtaining the timbre fluctuation segment in the audio sample, detecting the start time and the end time on the time axis, calculating the amplitude difference and time span difference between adjacent peaks, and obtaining the timbre fluctuation rhythm sequence; S412: Calling the timbre fluctuation rhythm sequence, comparing the timbre fluctuation segments in the target song, selecting fluctuation segments with the same rhythm distribution, calculating the rhythm fit strength value based on the cross-matching degree between the timbre spectrum change value and the rhythm density value, selecting the target segment that is consistent with the rhythm sequence of the sample fluctuation segment, and obtaining the style matching time interval; S413: According to the style matching time interval, the start and end positions of the interval in the time axis are recorded, the corresponding ethnic style labels are marked, the distribution structure of the ethnic styles on the time axis is analyzed, and a style segment time feature mapping result is generated.
7. A system for constructing a flute folk music library, characterized in that: The system is used to implement the method for constructing a flute folk music library according to any one of claims 1 to 6, and the system comprises: The audio feature recognition module obtains the complete audio data of the flute folk music, divides the audio into continuous segments, detects the timbre waveform change trajectory in each audio segment, counts the turning point values between adjacent peaks and troughs, records the time position intervals on the corresponding time axis, and obtains a time index table of the timbre change concentrated segments; The melody structure recognition module extracts the main melody note sequence within the time period recorded in the timbre change concentrated time index table, identifies the frequency direction change, and obtains a melody direction node sequence table by marking and numbering the nodes where the directionality changes; The style segment comparison module calls the node arrangement order recorded in the melody trend node sequence table, compares the node position and order with the melody structure path in the preset ethnic style sample, counts the number of structural segments matching with the ethnic style sample, and obtains an ethnic style proportion label set; The style feature mapping module selects the timbre fluctuation samples under the style labels in the ethnic style proportion label set, calls the timbre data of the corresponding paragraph, compares the fit with the fluctuation period, amplitude, and energy distribution density of the sample fluctuation segment, marks the same time period and records the corresponding style, and obtains the style segment time feature mapping result; The music library construction module calls the style label and time period information recorded in the style segment time feature mapping result, matches the unique number of the real-time track, combines the style label, corresponding time period interval and track number into a standard field format, and obtains the style music library field combination.
Citation Information
Patent Citations
Audio assisted creation method based on big data
CN109754773A
Method of music information retrieval and classification using continuity information
WO2005010865A2