A music time series data image encoding method based on space filling curve
Patent Information
- Application Number
- CN202610851653.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-25
AI Technical Summary
这类方法虽然直观,但波形图难以反映音乐的长期结构,频谱图则丢失了相位信息和时序细节
[0008]本发明提供的有益效果是:本发明通过引入音乐结构感知的自适应预处理、动态曲线选择、误差图像记录、多尺度金字塔编码以及保真度闭环优化等技术手段,有效解决了上述问题。与现有技术相比,本发明至少具有以下有益效果:第一,通过节拍同步的帧长调整和音乐语义特征提取,生成的编码图像具有良好的音乐可解释性,能够直观反映音乐的节奏模式、旋律轮廓和段落结构;第二,通过误差图像记录与补偿机制,实现了编码图像到原始音乐数据的高保真可逆还原,既可输出有损压缩的预览图像,也可输出无损编码的完整数据;第三,通过多尺度图像金字塔,能够同时呈现音乐从全局结构到局部细节的多层次信息,便于音乐分析;第四,通过邻近性保持保真度指标的自适应优化,使得编码参数能够根据不同音乐内容动态调整,提高了方法的通用性和鲁棒性。本发明的方案不仅适用于音乐数据的可视化存储与分析,也可推广到其他一维时序数据的结构化图像编码中。
Smart Images

Figure CN122824910A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of signal processing and image coding, and in particular to a method for encoding music temporal data images based on space-filling curves. Background Technology
[0002] Music time-series data is a typical one-dimensional time series signal, containing rich structured information such as rhythm, melody, and harmony. In tasks such as music information retrieval, music sentiment analysis, and automatic music annotation, how to effectively represent and encode the features of music time-series data is a fundamental problem. Space-filling curves (SFCs) are mathematical tools that can continuously map high-dimensional spaces to one-dimensional spaces, possessing excellent properties of preserving data locality. Classic space-filling curves such as Hilbert curves, Peano curves, and Z-order curves can, to the greatest extent possible, preserve the proximity relationships between adjacent points in the original data when reducing two-dimensional or higher-dimensional data to a one-dimensional sequence. In recent years, space-filling curves have been applied to image coding, genome visualization, and DNA data storage, and their "proximity preservation" ability provides new ideas for the structured representation of time-series data.
[0003] In existing technologies, encoding and visualization methods for music temporal data mainly fall into two categories. The first category is based on traditional signal processing methods, such as directly displaying the music waveform as an amplitude-time curve or generating a spectrogram through short-time Fourier transform. While intuitive, these methods struggle to reflect the long-term structure of the music in waveform graphs, and spectrograms lose phase information and temporal details. The second category is based on deep learning-based feature learning methods, such as inputting music segments into convolutional neural networks or recurrent neural networks to extract high-level features. These methods typically require large amounts of labeled data and have poor interpretability. Of particular note is the application of space-filling curves to the graphical representation of audio data. For example, the arXiv paper "A Novel Audio Representation using Space Filling Curves" proposes mapping the original audio waveform sampling points to a two-dimensional image using space-filling curves, and then using this for keyword recognition tasks. Essentially, this method directly fills a two-dimensional grid with a one-dimensional audio waveform, using the space-filling curve to preserve the local temporal structure of the audio signal. However, this method has the following drawbacks: First, it directly uses the original audio waveform sampling points without any preprocessing for the semantic features of the music, such as rhythm and pitch, resulting in a lack of interpretability at the musical level in the encoded image; Second, it uses a fixed order of space-filling curves, which cannot adapt to the encoding requirements of music data of different lengths or complexities; Third, this method only implements unidirectional mapping, does not provide a complete reversible decoding scheme, and lacks an effective compensation mechanism for quantization errors generated during the mapping process; Fourth, its output is a single-scale grayscale image, which cannot simultaneously present the global structure and local details of the music. Summary of the Invention
[0004] The technical problem to be solved by this invention is: how to provide a high-fidelity reversible image coding method for music temporal data. This method can make full use of the rhythmic structure and semantic features of music itself, adaptively select coding parameters, maximize the temporal proximity in the process of converting a one-dimensional music sequence into a two-dimensional image, and support multi-scale analysis and fully reversible decoding and restoration.
[0005] Specifically, the present invention provides a music temporal data image encoding method based on space-filling curves, the method comprising the following steps: S1. Obtain the original music timing data, preprocess the original music timing data, and obtain the preprocessed music sequence; S2. Determine the type of the space filling curve and the target image size, and calculate the order of the space filling curve based on the target image size to generate the corresponding space filling curve path index matrix; S3. Based on the space filling curve path index matrix, map each data value in the preprocessed music sequence to the corresponding pixel position of a blank two-dimensional image according to the path order, and quantize and encode the mapped pixel values to obtain the initial encoded image. S4. Record the quantization error at each pixel position during the quantization encoding process and generate an error image; S5. The initial encoded image and the error image are output together as the final music timing data encoded image; The space-filling curve is used to maintain the temporal proximity in a one-dimensional music sequence as the spatial proximity in a two-dimensional image, and the error image is used to compensate for quantization errors during decoding, thereby achieving reversible restoration of the music sequence.
[0006] A storage medium storing instructions and data for implementing a music timing data image encoding method based on a space-filling curve.
[0007] A music temporal data image encoding device based on a space-filling curve includes: a processor, a storage medium, and an edge computing acceleration module; the processor loads and executes instructions and data in the storage medium to implement a music temporal data image encoding method based on a space-filling curve.
[0008] The beneficial effects provided by this invention are as follows: By introducing adaptive preprocessing with music structure awareness, dynamic curve selection, error image recording, multi-scale pyramid coding, and fidelity closed-loop optimization, this invention effectively solves the aforementioned problems. Compared with existing technologies, this invention has at least the following beneficial effects: First, through frame length adjustment with beat synchronization and extraction of music semantic features, the generated coded image has good music interpretability and can intuitively reflect the rhythmic pattern, melody outline, and paragraph structure of the music; Second, through error image recording and compensation mechanisms, high-fidelity reversible restoration from the coded image to the original music data is achieved, which can output both lossy compressed preview images and lossless coded complete data; Third, through multi-scale image pyramids, multi-level information of music from global structure to local details can be presented simultaneously, facilitating music analysis; Fourth, through adaptive optimization of the proximity-preserving fidelity index, the coding parameters can be dynamically adjusted according to different music content, improving the versatility and robustness of the method. The solution of this invention is not only applicable to the visualization, storage, and analysis of music data, but can also be extended to the structured image coding of other one-dimensional time-series data. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the hardware device operation according to an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0011] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.
[0012] Example 1 Please refer to Figure 1 The present invention provides a method for encoding music temporal data images based on space-filling curves, comprising the following steps: S1. Obtain the original music timing data, preprocess the original music timing data, and obtain the preprocessed music sequence; It should be noted that step S1 further includes: S11. Perform sampling rate unification, channel merging and amplitude normalization on the original music timing data to obtain preliminary cleaned data; As a specific embodiment of the present invention, it is assumed that the original music file is a stereo 44.1kHz, 16-bit WAV format. First, the stereo is merged into a mono channel (taking the average of the left and right channels). Then, it is resampled to a uniform 16kHz (to reduce computational load). Finally, amplitude normalization is performed: the maximum absolute value max_abs of the entire waveform is calculated, and each sample point is divided by max_abs to ensure that all sample values fall within the interval [-1, 1]. The preliminary cleaned data Y is obtained.
[0013] S12. Detect the instantaneous beat period of the preliminary cleaning data, and adaptively adjust the frame length of the frame processing according to the instantaneous beat period so that each frame contains an integer multiple of the beat period, and set the frame shift to half a beat period. In one specific embodiment of the present invention, instantaneous beat period detection is performed on Y. The specific method involves calculating a short-time energy sequence, then using the autocorrelation function method to determine the fundamental frequency of the energy sequence. The reciprocal of this fundamental frequency is the beat period (unit: seconds). For example, the beat period of the current music segment is detected to be 0.5 seconds (corresponding to 120 BPM). The target is that each frame contains two complete beat periods, so the frame length = 2 × 0.5 = 1 second. The frame shift is set to half a beat period = 0.25 seconds. Based on a sampling rate of 16kHz, the frame length corresponds to 16000 sampling points, and the frame shift corresponds to 4000 sampling points. Thus, each frame exactly covers two complete measures (4 beats), and there is a 0.75-second overlap between frames (3 beats), ensuring beat boundary alignment.
[0014] S13. Extract music feature parameters from each frame. The music feature parameters include at least one of the following: short-time energy, zero-crossing rate, Mel frequency cepstral coefficients (MFCC) and their first-order difference, chromaticity features, beat intensity sequence, and pitch profile. Concatenate the feature parameters of each frame in chronological order to form a one-dimensional feature sequence. This feature sequence is used as the object of subsequent encoding. During decoding, the corresponding reconstructed feature sequence is output. The reconstructed feature sequence is used for music information retrieval or music structure analysis, but not for reconstructing the original audio waveform. As a specific embodiment of the present invention, the following features are extracted for each frame of data (taking the t-th frame as an example): Short-time energy: E t = sum(s i 2 ), where s i For intra-frame sampling points. Zero-crossing rate: ZCR t = 1 / 2 * sum(|sign(s i ) - sign(s i-1 MFCC and its first-order difference: First, the energy of the Mel filter bank is calculated, then the 12-dimensional MFCC coefficients are obtained through discrete cosine transform, and then the difference values of the MFCCs of adjacent frames are calculated, for a total of 24 dimensions. Chroma: The spectral energy is mapped onto 12 semitone levels to obtain a 12-dimensional vector. Beat intensity sequence: The beat intensity values (0~1) are obtained by calculating the peak position of the inter-frame autocorrelation function. Pitch profile: The fundamental frequency (F0) is extracted using the YIN algorithm, and non-voiced segments are filled with 0s.
[0015] All the above features are concatenated into a multi-dimensional feature vector F. t In this embodiment, the feature dimensions of each frame are: 1 (energy) + 1 (ZCR) + 24 (MFCC + differential) + 12 (Chroma) + 1 (beat intensity) + 1 (pitch) = 40 dimensions.
[0016] S14. The music feature parameters extracted from all frames are concatenated in chronological order to form a one-dimensional preprocessed music sequence.
[0017] In one specific embodiment of the present invention, the feature vectors of all frames are arranged in chronological order to form a two-dimensional matrix F, with a size of frame number × 40. Then, it is expanded row-wise (row-major priority) into a one-dimensional sequence X. feat The sequence length is 40 frames. This sequence is the "preprocessed music sequence" that is finally output in step S1. It retains the temporal structure of the original music and condenses the semantic information of the music.
[0018] A concrete example: For the aforementioned 5-second piano piece, after steps S11~S14, approximately 20 frames are obtained (frame length 1 second, frame shift 0.25 seconds, approximately 20 frames in total for 5 seconds), each frame has 40-dimensional features, therefore X feat The length is 800. Subsequent steps S2~S5 will perform image encoding on this 800-length feature sequence, instead of the original waveform. Since the feature sequence has a low dimension, the required image size can be smaller (e.g., 32×32=1024 pixels when n=5, slightly larger than 800), thus significantly reducing the computational load.
[0019] S2. Determine the type of the space filling curve and the target image size, and calculate the order of the space filling curve based on the target image size to generate the corresponding space filling curve path index matrix; It should be noted that step S2 further includes: S21. Extract the type features of the original music time series data, including rhythm regularity index, pitch variation degree and spectrum distribution, and construct a music type vector; As a specific embodiment of the present invention, step S21 specifically includes: Before step S2 begins, type features are extracted from the original music time series data: Rhythm regularity index: Calculated by taking the autocorrelation function of a short-time energy sequence and taking the ratio of the largest peak (excluding zero-delay) to the second largest peak. The larger the ratio, the more regular the rhythm. For example, the regularity index of electronic dance music can reach above 0.9, while that of free jazz may only be 0.3.
[0020] Pitch variation: Extract the fundamental frequency sequence and calculate the standard deviation of the fundamental frequency difference between adjacent frames. A larger standard deviation indicates a more drastic pitch variation. Spectral distribution: Calculate the average spectral centroid and bandwidth of the entire music.
[0021] After normalizing these three features, a three-dimensional music type vector V = (rhythm) is constructed. reg pitch var ,spec cent ).
[0022] S22. Based on the music type vector, dynamically select the optimal curve type from the set of space-filling curves: when the rhythm regularity index is greater than the first threshold, select the Hilbert curve as the space-filling curve; when the rhythm regularity index is less than the second threshold, select the Z-order curve as the space-filling curve. As a specific embodiment of the present invention, step S22 specifically includes: Set the first threshold Th1 = 0.7 and the second threshold Th2 = 0.4. When the rhythm...reg When the value is greater than 0.7, the music rhythm is considered to be very regular (such as dance music or marches). In this case, the Hilbert curve is chosen because it is optimal in preserving locality and is suitable for data with strong regularity. When the rhythm... reg When the value is less than 0.4, the music is considered to have free rhythm (such as classical piano solos or atonal music). In this case, the Z-order curve is chosen because it is simpler to calculate and adapts better to irregular data. When 0.4 ≤ rhythm... reg When the value is ≤ 0.7, the Hilbert curve can be selected by default, or further subdivided according to the degree of pitch change: if the degree of pitch change is also high, then select the Z-order curve; otherwise, select the Hilbert curve.
[0023] S23. Based on the selected space fill curve type and target image size 2 n ×2 n The corresponding space-filling curve path index matrix is generated recursively, where n is the order of the space-filling curve and n is an integer greater than or equal to 1.
[0024] As a specific embodiment of the present invention, step S23 specifically includes: Assuming a Hilbert curve is chosen according to the above rules, and the order n=5 is determined based on the data length L=800 (because 2 2*5 =1024≥800). Then, the standard Hilbert curve recursive generation algorithm is used: The first-order Hilbert curve is a " A shape that covers a 2×2 grid.
[0025] Recursive rule: Copy the current order curve four times, rotate it by 0°, 90°, 180° and 270° respectively, and then connect them with three line segments.
[0026] For n=5, it is necessary to recursively call the algorithm 5 times, eventually obtaining a 32×32 grid. The curve passes through each grid point exactly once, generating a path index matrix M with a size of 32×32 and matrix element values from 1 to 1024.
[0027] If a Z-order curve is selected, the index is generated by interleaving binary bits: for grid coordinates (x, y) (binary representation), the Z-order index is the value obtained by interleaving the bits of x and y.
[0028] For example: For an electronic dance music track (rhythm regularity index = 0.95), step S22 selects a Hilbert curve with n=5 to generate a 32×32 index matrix. For a free jazz track (rhythm regularity index = 0.35), a Z-order curve is selected, which also generates a 32×32 index matrix. Dynamic curve selection allows the encoded image to better adapt to the inherent structure of different music genres.
[0029] S3, according to the space-filling curve path index matrix, mapping each data value in the preprocessed music sequence to corresponding pixel positions of a blank two-dimensional image according to the path order, and performing quantization encoding on the mapped pixel values to obtain an initial encoded image; Step S3 further comprises: S31, comparing the length L of the preprocessed music sequence with the total number of pixels P=2 corresponding to the space-filling curve path index matrix 2n ; As a specific embodiment of the present invention, step S31 specifically comprises: Assuming that after step S1, the length of the obtained preprocessed music sequence is L=800, and the order n of the curve selected in step S2 is 5, then the total number of pixels P=2 10 =1024. Obviously, L<P, so padding is required.
[0030] S32, if L<P, a cyclic padding method is used to repeatedly fill the preprocessed music sequence from the beginning until the length P is reached; or according to user settings, no padding is performed and only the corresponding positions in the image are left black (pixel values set to 0), and the information of unfilled blank positions is recorded; if L>P, according to the distribution of the sum of squares of the amplitudes of each sampling point in the music sequence, a data segment containing consecutive P sampling points and having the highest energy, that is, the highest sum of squares of amplitudes, is intercepted, and the information of the truncated position is recorded; As a specific embodiment of the present invention, the padding strategy in step S32 specifically comprises: This embodiment provides two padding strategies, which can be selected according to application requirements. Strategy 1 (default): Circular padding. Repeat the preprocessed music sequence from the beginning until P points are filled. For example, if the length of the original sequence is 800 and the target length is 1024, then the padded sequence is the original sequence repeated twice (800+800=1600, which exceeds 1024, so the first 1024 points are taken). This padding method is simple and does not introduce textures inconsistent with the original material. Strategy 2: No padding and leaving blank. If L < P, only the first L data are mapped to the corresponding L pixel positions in the image, and the remaining (P-L) pixel positions are directly set to 0 (or a fixed background value), and the blank position indices are recorded in the metadata, and these positions do not participate in feature recovery during decoding. This strategy avoids the interference of padded data on image texture and is suitable for subsequent tasks that rely on proximity analysis.
[0031] In another case, if L > P (for example, the original sequence is very long), truncation is performed according to the energy distribution. Calculate the square of the amplitude (energy) of each sampling point, intercept the data segment containing consecutive P sampling points with the highest sum of amplitude squares (i.e., energy), and discard the rest. The start index of the truncated position is also recorded and saved as metadata, so that the general structure can be recovered during decoding.
[0032] S33, according to the space-filling curve path index matrix, assign the k-th data value in the music sequence after length matching to the pixel position with the path index value k in the blank two-dimensional image, where k=1,2,…,P; As a specific embodiment of the present invention, step S33 specifically includes: Using the index matrix M (32×32, with values from 1 to 1024) generated in step S2. Create a 32×32 blank image I. For k from 1 to 1024, find the position (i,j) with value k in M, and assign X filled (k) to I(i,j). After completion, I is the mapped image, and the pixel values are music feature values.
[0033] S34, perform discretization encoding on the assigned pixel values by using an adaptive quantization method, the adaptive quantization method is: for each pixel, take the pixel values in its 3×3 neighborhood window, calculate the standard deviation σ of the neighborhood, set the quantization step Δ as Δ = α·σ, where α is a preset proportionality coefficient, and then take the rounded integer value obtained by dividing the current pixel value by Δ as the quantization encoded value, so as to obtain the initial encoded image.
[0034] As a specific embodiment of the present invention, the adaptive quantization in step S34 specifically includes: Perform adaptive quantization on the mapped image I, set α=1.0. Traverse each pixel (i,j): Take the 3×3 neighborhood of the pixel (only the actual existing pixels are taken at the boundary, such as the neighborhood of the top left corner (1,1) contains four pixels (1,1), (1,2), (2,1), and (2,2)).
[0035] Calculate the standard deviation σ of these 9 (or fewer) pixel values. The quantization step size Δ = α·σ = σ. For the current pixel value v, the quantized encoded value q = round(v / Δ). The quantized restored value v recon = q * Δ. Error e = v - v recon The error image was recorded.
[0036] After adaptive quantization, the quantization step size of locally flat regions in the image is smaller (σ small), preserving more details; while the quantization step size of edge or abrupt regions is larger (σ large), using fewer bits for encoding, thereby achieving rate-distortion optimization.
[0037] For example, in a 32×32 mapped image, the central region represents the duration of a musical note, where pixel values change gradually, with a local σ≈0.05 and Δ=0.05, resulting in a small range of quantized encoded values. However, at the start of the note, pixel values change abruptly, with a local σ≈0.3 and Δ=0.3, expanding the range of quantized encoded values. The final quantized image I... q With error image I err Save them together.
[0038] S4. Record the quantization error at each pixel position during the quantization encoding process and generate an error image; It should be noted that step S4 further includes: S41. For each pixel position k, calculate the quantization error ε. k = Original value - Quantized restored value; As a specific embodiment of the present invention, step S41 specifically includes: In step S34, for each pixel (i,j) quantized, its error ε(i,j) is calculated. In this embodiment, these error values are stored in a matrix E of the same size as I, initialized to 0. During the quantization loop, E(i,j) = ε(i,j) is directly updated.
[0039] S42. Arrange all error values in the same spatial filling curve path order as the original data to form a one-dimensional error sequence; As a specific embodiment of the present invention, step S42 specifically includes: Since subsequent steps require the error sequence for decoding, and the one-dimensional sequence needs to be reconstructed according to the same spatial filling curve path order during decoding, the error matrix E is rearranged into a one-dimensional sequence E according to the path order. seqThe specific procedure is as follows: For k=1~P, find the position (i,j) with value k in the index matrix M, take out E(i,j), and put it into E in sequence. seq (k). Thus, E seq The k-th error value in the data corresponds exactly to the quantization error at the k-th position (after padding) in the original data.
[0040] S43. Encode the one-dimensional error sequence into an error image with the same size as the initial encoded image. Each pixel position of the error image records the quantization error of the main image at the corresponding position. As a specific embodiment of the present invention, step S43 specifically includes: The one-dimensional error sequence E_seq is encoded into an error image. Since the error values are typically small (between -0.5Δ and 0.5Δ), they need to be scaled and offset for easier storage and display. For example, the error value is multiplied by 100 and added to 128, then limited to the range of 0-255 to obtain a grayscale value. Then, following the same mapping method as the main image (using the same M matrix, but with "path-order fill"), the scaled error value is filled into a new image I. err middle. I err Size and I q same.
[0041] S44. Store the initial encoded image and the error image together as the final music timing data encoded image.
[0042] As a specific embodiment of the present invention, step S44 specifically includes: Main image I q And error image I err Shared storage. This embodiment uses a multi-page TIFF format, storing the two images as two separate pages. JSON-formatted metadata is written to the ImageDescription tag of the TIFF file, including: curve type, order n, α value, padding method (theme extension), original length L, and quantization parameters min / max (if uniform quantization is used). Finally, a single file "music_encoding.tif" is generated. This file is the final encoded image of the music timing data, supporting subsequent reversible decoding.
[0043] For example, for a piano piece excerpt, the main image I q It exhibits a distinct striped texture, corresponding to the musical note sequence; Error Image I errThe image is almost entirely black (with most errors close to 0), with only faint gray dots at the start and end points of the notes. The two images together constitute the complete encoding result. If the error image is lost, a lossy version close to the original music can be recovered using only the main image; retaining the error image allows for completely lossless restoration.
[0044] S5. The initial encoded image and the error image are output together as the final music timing data encoded image; The space-filling curve is used to maintain the temporal proximity in a one-dimensional music sequence as the spatial proximity in a two-dimensional image, and the error image is used to compensate for quantization errors during decoding, thereby achieving reversible restoration of the music sequence.
[0045] As a specific embodiment of the present invention, step S5 specifically includes: The initial encoded image I q And error image I err The output is a pair of images. It can be saved as two separate PNG files, or the error image can be embedded into the main image file as an alpha channel or additional layer. Simultaneously, the encoding parameters (curve type = Hilbert curve, order n = 9, quantization parameters min / max, padding mode = zero padding, etc.) are written as metadata into the image file's annotation field. The final output file is the encoded image representation of the music time-series data, which can be used for subsequent analysis, storage, or transmission.
[0046] For example, consider a 5-second excerpt from the piano piece "Für Elise" (sampling rate 44.1kHz, mono). After processing according to steps S1-S5, a 512×512 grayscale main image and an error image of the same size are obtained. In the main image, the sampled values at adjacent time points are spatially mapped to adjacent or nearby pixel positions, thus the image exhibits a regular texture pattern, reflecting the waveform envelope and rhythmic changes of the music. In the error image, the error value in most areas is close to 0 (the quantization of the main image is more accurate), but the error value is larger at abrupt changes in the waveform (such as the starting point of a note), appearing as sparse bright spots. Storing both images together allows for the reversible recovery of the music sequence data.
[0047] It should be noted that the method also includes a step of multi-scale encoding of the music temporal data: S6. Use multiple different space-filling curve orders n1, n2, …, n respectively. k , where n1 <n2<…<n k Repeat steps S2 to S5 to generate music temporal data encoded image pairs corresponding to multiple scales. The image pairs include the main image I. m With error image E mThis forms a multi-scale image pyramid pair; S7. Scale the master images of each layer in the multi-scale image pyramid to the maximum size 2 using bilinear interpolation. nk ×2 nk Then, they are stacked into a three-dimensional tensor in order from low to high order. The error images of each layer are stacked in the same way to generate two multi-channel encoded data: the main tensor and the error tensor.
[0048] As a specific embodiment of the present invention, step S6 specifically includes: This embodiment selects three different orders: n1=4 (16×16=256 pixels), n2=5 (32×32=1024 pixels), and n3=6 (64×64=4096 pixels). For each order, steps S2 to S5 are repeated (note: the curve type and fill strategy in step S2 remain the same to ensure multi-scale consistency). This yields three encoded image pairs at different scales: (I4, E4), (I5, E5), and (I6, E6). These three image pairs form a multi-scale image pyramid, where the lower scale (I4) reflects the global structure of the music (e.g., paragraph outlines), and the higher scale (I6) reflects local details (e.g., note beginnings and endings).
[0049] As a specific embodiment of the present invention, step S7 specifically includes: Since the three images have different sizes (16×16, 32×32, 64×64), they need to be unified to the maximum size of 64×64. Bilinear interpolation is used to upsample I4 and I5: for I4, each original pixel is expanded into a 4×4 block, and the pixel values within the block are obtained through bilinear interpolation; I5 is processed similarly. After upsampling, I is obtained. 4_up and I 5_up All of them are 64×64 in size. Then I 4_up I 5_up I6 are stacked in ascending order of order to form a 64×64×3 three-dimensional tensor, Tensor_main. The third dimension is the channel dimension: channel 1 corresponds to global information (n=4), channel 2 to medium-scale information (n=5), and channel 3 to detail information (n=6). Similarly, the same upsampling and stacking are performed on error images E4, E5, and E6 to obtain the error tensor. err .
[0050] For example, for a 2-minute symphony, a low-scale image (16×16) can only roughly show the fluctuations of the movement (such as forte and pianissimo passages), a medium-scale image (32×32) can show the changes in each musical phrase, and a high-scale image (64×64) can clearly present the envelope of each note. The stacked 3D tensor can be regarded as a "musical cube" to analyze the musical structure at different scales. This tensor can be saved in NIfTI format (commonly used for medical images) or directly stored as a multi-channel PNG (such as using three channels in RGBA).
[0051] It should be noted that the method also includes steps for encoding performance evaluation and adaptive optimization: Calculate the proximity preservation fidelity index F = 1 - (D) for the coded image. loss / D total ), where D loss To address the proximity loss caused by the mismatch between data length and total number of pixels during space-filling curve mapping, D total The total proximity energy of the original sequence; As a specific embodiment of the present invention, "proximity energy" is first defined. For an original one-dimensional music sequence X (length L), the sum of squares of the differences between its adjacent points is calculated: D total = This value reflects the overall volatility of the sequence. After mapping to an image, for any two adjacent points (i and i+1) in the original sequence, their Euclidean distance d(i, i+1) in the image may be greater than 1 (if the curved path maps adjacent points to non-adjacent pixels). The "proximity loss" is defined as: In other words, if neighboring points with large fluctuations are mapped to pixels that are far away, the loss is significant; conversely, neighboring points with small fluctuations suffer less loss even if they are relatively far apart. Therefore, the fidelity index... The closer F is to 1, the better the proximity is maintained.
[0052] If the fidelity index F is lower than the preset threshold, the encoding parameter reselection is triggered, and the process returns to step S2 to reselect the space filling curve order or type. As a specific embodiment of the present invention, a preset threshold Th is set. F=0.85. After the initial encoding (e.g., using n=5, Hilbert curve), calculate the F value. Assuming the calculated F=0.78, which is lower than 0.85, trigger parameter reselection. Reselection strategies include: trying to increase the order n (e.g., from 5 to 6), because a larger image size can accommodate more pixel spacing, making it easier for adjacent points to be mapped to closer positions in the image. Trying to change the curve type (from Hilbert curve to Z-order curve), because different curves have different locality preservation capabilities under different data distributions. Return to step S2, re-execute S2~S5 with the new parameters, and recalculate the F value.
[0053] If the fidelity index is still below the threshold after repeated reselection more than the preset number of times, a warning message will be output and the combination of encoding parameters with the highest fidelity will be used.
[0054] In one specific embodiment of the present invention, the maximum number of reselections is set to 3. If the F value is still lower than 0.85 after 3 reselections, the process stops, and a warning message "Warning: Unable to reach the preset fidelity threshold, current optimal F=0.82" is output. Then, the set of parameters with the highest F value from the 3 reselections and the initial encoding is selected as the final encoding parameters, and the corresponding encoded image is output.
[0055] For example, for a piece of highly random, noisy music, even using a high-order Hilbert curve, the F-value might only be 0.65. In this case, the system will try a Z-order curve (F=0.68), then a Peano curve (F=0.66), neither of which exceeds 0.85. Finally, the system selects the Z-order curve with the highest F-value and outputs a warning. The user can then decide whether to accept lossy proximity preservation or use a longer music segment for encoding based on the warning.
[0056] It should be noted that the method also includes the step of decoding and restoring the generated final music timing data encoded image: S81. Obtain the main encoded image and error image to be decoded. If it is a compressed format, decompress it first and parse out the space filling curve type, order n, adaptive quantization scaling factor α and filling or truncation method. As a specific embodiment of the present invention, step S81 specifically includes: Suppose a user receives a file "music_encoding.tif" containing two pages of images (main image and error image) and metadata. First, the file is read using a TIFF library and decompressed (if LZW compression is used). The JSON metadata is parsed from the ImageDescription tag, yielding: curve_type="hilbert", n=5, alpha=1.0, padding_method="theme_extension", original_length=800, quantization_params = {"min": -1.0, "max": 1.0} (if uniform quantization is used) or the Δ value for each pixel is directly stored (if adaptive quantization is used, this embodiment stores the σ value for each pixel as additional metadata to simplify decoding).
[0057] S82. Based on the parsed curve type and order n, regenerate the same space-filling curve path index matrix; As a specific embodiment of the present invention, step S82 specifically includes: Based on curve_type="hilbert" and n=5, generate a 32×32 Hilbert curve path index matrix M using the same recursive algorithm as the encoding (consistent with the encoding).
[0058] S83. When multi-scale coding is used, the main image layer and error image layer of the m-th channel are selected according to the decoding scale m specified by the user, and the image of the layer is inversely scaled to the original size of 2nm×2nm by bilinear interpolation. As a specific embodiment of the present invention, step S83 specifically includes: If multi-scale encoding is used (e.g., n1=4, n2=5, n3=6), the user specifies the decoding scale m=2 (medium scale). The main image (64×64) and error image (64×64) of the second channel are extracted from the multi-channel tensor. Since the original image size corresponding to n2=5 should be 32×32, the 64×64 main image needs to be inversely scaled back to 32×32 using bilinear interpolation. Note: Inverse scaling may introduce interpolation errors, but these errors can be compensated for since the error image is also saved. This embodiment recommends directly saving the original multi-scale images during encoding without scaling and stacking to avoid interpolation errors; however, if stacking is used to save storage space, inverse scaling must be performed during decoding.
[0059] S84. According to the inverse mapping relationship of the path index matrix, extract the quantization code value of each pixel from the main image pixel quantization value matrix, and reconstruct it into a one-dimensional quantization sequence according to the path order. As a specific embodiment of the present invention, step S84 specifically includes: Obtain the 32×32 main image pixel quantization value matrix I q According to the inverse mapping relationship, for k=1 to P (P=1024), find the position (i,j) with value k in matrix M, and retrieve I. q (i,j), put into a one-dimensional sequence Q seq (k). Q seq This is the sequence of quantized encoded values.
[0060] S85. Perform inverse quantization on the one-dimensional quantization sequence to obtain the main image restoration value. The restoration value = quantization value × Δ, where Δ = α·σ, and σ is the 3×3 neighborhood standard deviation of the corresponding pixel in the decoded image. As a specific embodiment of the present invention, step S85 specifically includes: Q seq Each quantization value q in k The original value needs to be restored. Since adaptive quantization (based on the 3×3 neighborhood standard deviation σ) is used during encoding, the σ value for each pixel needs to be known during decoding. In this embodiment, the σ value for each pixel (a one-dimensional array of length 1024) is stored in the encoded metadata. After reading, for the k-th position, its quantization step size Δk = α * σ k = 1.0 * σ k Then the restored value v of the main image recon_k = q_k * Δk. This yields the one-dimensional reconstructed sequence V. recon .
[0061] S86. Extract the quantization error value of each pixel position from the error image, reconstruct the one-dimensional error sequence in the same path order, and superimpose the restored value of the main image with the error value to obtain the compensated restored sequence. As a specific embodiment of the present invention, step S86 specifically includes: Similarly, the error value sequence E is extracted from the error image I_err in the same path order. seq (Note that the error image stores scaled integers, which need to be restored to the true error: e) k = (E seq (k) - 128) / 100). Then the restored value of the main image is superimposed with the error value: V comp_k = V recon_k + e k The obtained V comp That is, the restored sequence after compensation, which should theoretically be equal to the original padded sequence before encoding.
[0062] S87. Based on the padding or truncation method in the metadata, remove the redundant padding data to obtain the reconstructed one-dimensional feature sequence. The reconstructed one-dimensional feature sequence maintains the reversible restoration accuracy in numerical terms with the preprocessed music sequence output in step S1 before encoding, but it is not used to reconstruct the original audio waveform. Instead, it is directly used as the input feature for music information retrieval, music structure analysis, or music visualization tasks.
[0063] As a specific embodiment of the present invention, step S87 specifically includes: Based on the padding method in the metadata (e.g., circular padding or no padding with black) and the original length L=800, from V comp The first 800 points are extracted to obtain the reconstructed one-dimensional feature sequence V. reconstructed The sequence maintains high precision and reversible reconstruction with respect to the preprocessed music sequence output from step S1 before encoding (quantization error has been compensated by the error image).
[0064] It is important to note that since this method encodes the music feature parameters extracted in step S13 (such as MFCC, chroma features, beat intensity sequence, etc.), rather than the original audio waveform sampling points, the decoded result is a reconstructed feature sequence, not a playable audio waveform. Reverse engineering the original waveform from features such as MFCC is mathematically underdetermined and is not within the scope of this invention. The feature sequence decoded by this invention can be directly used for downstream music information retrieval tasks, for example: Calculate the similarity between the decoded features and the features in the library for cover song recognition or copyright infringement detection; The decoded features are input into a pre-trained machine learning model for music emotion classification or structural analysis. Visualize decoding features and generate music structure maps, etc.
[0065] The decoded output is typically saved as a feature vector sequence file (such as NumPy's .npy format, CSV format, or HDF5 format), along with metadata recorded during encoding (including original length, padding method, quantization parameters, etc.) for direct use in subsequent analysis tasks. This embodiment does not perform inverse waveform synthesis on the feature sequence, nor does it output a WAV audio file.
[0066] Specifically, if the encoded sequence is a feature sequence, the decoded sequence is also a reconstructed feature sequence. This sequence is numerically almost lossless compared to the original feature sequence (the error is compensated by the quantization error image). It can be directly used for subsequent tasks such as music information retrieval (e.g., similarity calculation, cover song recognition), music structure analysis (e.g., paragraph boundary detection), or music visualization, without needing to attempt to reverse engineer it back to the original audio waveform. The decoded output is usually not saved as a WAV file, but rather as a feature vector sequence file (e.g., .npy or .csv format) for interface with downstream tasks.
[0067] Example 2: Please see Figure 2 , Figure 2 This is a schematic diagram of the hardware device in operation according to an embodiment of the present invention. The hardware device specifically includes: a music timing data image encoding device 401 based on a space-filling curve, a processor 402, and a storage medium 403.
[0068] A music timing data image encoding device 401 based on a space-filling curve: The music timing data image encoding device 401 based on a space-filling curve implements the music timing data image encoding method based on a space-filling curve.
[0069] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the music timing data image encoding method based on space-filling curve.
[0070] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the music timing data image encoding method based on space-filling curve.
[0071] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for encoding music temporal data images based on space-filling curves, characterized in that: Includes the following steps: S1. Obtain the original music timing data, and preprocess the original music timing data to obtain the preprocessed music sequence; S2. Determine the type of the space filling curve and the target image size, and calculate the order of the space filling curve based on the target image size to generate the corresponding space filling curve path index matrix; S3. Based on the space filling curve path index matrix, map each data value in the preprocessed music sequence to the corresponding pixel position of a blank two-dimensional image according to the path order, and quantize and encode the mapped pixel values to obtain the initial encoded image. S4. Record the quantization error at each pixel position during the quantization encoding process and generate an error image; S5. The initial encoded image and the error image are output together as the final music timing data encoded image; The space-filling curve is used to maintain the temporal proximity in a one-dimensional music sequence as the spatial proximity in a two-dimensional image, and the error image is used to compensate for quantization errors during decoding, thereby achieving reversible restoration of the music sequence.
2. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that: Step S1 further includes: S11. Perform sampling rate unification, channel merging and amplitude normalization on the original music timing data to obtain preliminary cleaned data; S12. Detect the instantaneous beat period of the preliminary cleaning data, and adaptively adjust the frame length of the frame processing according to the instantaneous beat period so that each frame contains an integer multiple of the beat period, and set the frame shift to half a beat period. S13. Extract music feature parameters from each frame. The music feature parameters include at least one of the following: short-time energy, zero-crossing rate, Mel frequency cepstral coefficients (MFCC) and their first-order difference, chromaticity features, beat intensity sequence, and pitch profile. Concatenate the feature parameters of each frame in chronological order to form a one-dimensional feature sequence. This feature sequence is used as the object of subsequent encoding. During decoding, the corresponding reconstructed feature sequence is output. The reconstructed feature sequence is used for music information retrieval or music structure analysis, but not for reconstructing the original audio waveform. S14. The music feature parameters extracted from all frames are concatenated in chronological order to form a one-dimensional preprocessed music sequence.
3. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that, Step S2 further includes: S21. Extract the type features of the original music time series data, including rhythm regularity index, pitch variation degree and spectrum distribution, and construct a music type vector; S22. Based on the music type vector, dynamically select the optimal curve type from the set of space-filling curves: when the rhythm regularity index is greater than the first threshold, select the Hilbert curve as the space-filling curve; when the rhythm regularity index is less than the second threshold, select the Z-order curve as the space-filling curve. S23. Based on the selected space fill curve type and target image size 2 n ×2 n The corresponding space-filling curve path index matrix is generated recursively, where n is the order of the space-filling curve and n is an integer greater than or equal to 1.
4. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that, Step S3 further includes: S31. The length L of the preprocessed music sequence is equal to the total number of pixels P corresponding to the space-filling curve path index matrix. 2n Compare; S32. If L < P, the pre-processed music sequence is repeatedly filled in from the beginning by cyclic filling until the length P is reached; or according to user settings, no filling is performed, only the corresponding position in the image is left black (pixel value is set to 0), and the information of the unfilled blank position is recorded; if L > P, according to the distribution of the sum of squares of the amplitude of each sampling point in the music sequence, a data segment containing consecutive P sampling points with the highest energy, that is, the sum of squares of amplitude, is intercepted, and the information of the truncated position is recorded; S33. According to the space-filling curve path index matrix, the k-th data value in the music sequence after length matching is assigned to the pixel position with the path index value k in the blank two-dimensional image, where k=1, 2, …, P; S34. The adaptive quantization method is used to perform discretization coding on the assigned pixel values. The adaptive quantization method is: for each pixel, take the pixel values in its 3×3 neighborhood window, calculate the standard deviation σ of the neighborhood, set the quantization step size Δ as Δ = α·σ, where α is a preset proportionality coefficient, and then the rounded integer value obtained by dividing the current pixel value by Δ is used as the quantization coding value, so as to obtain the initial coded image.
5. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that: Step S4 further comprises: S41. For each pixel position k, calculate the quantization error ε. k = Original value - Quantized restored value; S42. Arrange all error values in the same space-filling curve path order as the original data mapping to form a one-dimensional error sequence; S43. Encode the one-dimensional error sequence into an error image with the same size as the initial coded image, and each pixel position of the error image records the main image quantization error at the corresponding position; S44. Store the initial coded image and the error image together as the final music time series data coded image.
6. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that: The method further comprises the step of performing multi-scale coding on the music time series data: S6. Use multiple different space-filling curve orders n1, n2, …, n respectively. k , where n1 <n2<…<n k Repeat steps S2 to S5 to generate music temporal data encoded image pairs corresponding to multiple scales. The image pairs include the main image I. m With error image E m This forms a multi-scale image pyramid pair; S7. Scale the master images of each layer in the multi-scale image pyramid to the maximum size 2 using bilinear interpolation. nk ×2 nk Then, they are stacked into a three-dimensional tensor in order from low to high order. The error images of each layer are stacked in the same way to generate two multi-channel encoded data: the main tensor and the error tensor.
7. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that, The method further comprises the step of coding effect evaluation and adaptive optimization: Calculate the proximity preservation fidelity index F = 1 - (D) of the coded image. loss / D total ), where D loss To address the proximity loss caused by the mismatch between data length and total number of pixels during space-filling curve mapping, D total The total proximity energy of the original sequence; If the fidelity index F is lower than the preset threshold, reselection of coding parameters is triggered, and the process returns to step S2 to reselect the order or type of the space-filling curve; If the fidelity index is still lower than the threshold after repeated reselection exceeds the preset number of times, a warning message is output and the coding parameter combination with the highest fidelity is adopted.
8. The music temporal data image encoding method based on space-filling curves as described in claim 1, characterized in that, The method further comprises the step of decoding and restoring the generated final music time series data coded image: S81. Obtain the main coded image and the error image to be decoded, if they are in compressed format, decompress them first, and parse out the type of space-filling curve, the order n, the adaptive quantization proportionality coefficient α, and the filling or truncation method therefrom; S82. Regenerate the same space-filling curve path index matrix according to the parsed curve type and order n; S83. When using multi-scale coding, select the main image layer and error image layer of the m-th channel according to the user-specified decoding scale m, and inversely scale the image of this layer to the original size 2 using bilinear interpolation. nm ×2 nm ; S84. According to the inverse mapping relationship of the path index matrix, extract the quantization coding value of each pixel from the main image pixel quantization value matrix, and reconstruct a one-dimensional quantization sequence according to the path order; S85. Perform inverse quantization on the one-dimensional quantization sequence to obtain the restored value of the main image, where restored value = quantization value × Δ, and Δ = α·σ, σ is the 3×3 neighborhood standard deviation of the corresponding pixel in the decoded image; S86. Extract the quantization error value of each pixel position from the error image, reconstruct the one-dimensional error sequence in the same path order, and superimpose the restored value of the main image with the error value to obtain the compensated restored sequence. S87. Based on the padding or truncation method in the metadata, remove the redundant padding data to obtain the reconstructed one-dimensional feature sequence. The reconstructed one-dimensional feature sequence maintains the reversible restoration accuracy in numerical terms with the preprocessed music sequence output in step S1 before encoding, but it is not used to reconstruct the original audio waveform. Instead, it is directly used as the input feature for music information retrieval, music structure analysis, or music visualization tasks.
9. A storage medium, characterized in that: The storage medium stores instructions and data to implement the music timing data image encoding method based on space-filling curves as described in any one of claims 1 to 8.
10. A music temporal data image encoding device based on a space-filling curve, characterized in that: include: Processor, storage media, and edge computing acceleration modules; The processor loads and executes instructions and data in the storage medium to implement the music timing data image encoding method based on space-filling curves as described in any one of claims 1 to 8.