Music track matching vibration haptic feedback method, system, and related devices
By using a deep learning model to process music tracks, the energy proportion of the tracks is calculated and weighted to generate vibration signals, which solves the problem of inaccurate vibration feedback in existing technologies and improves the user's tactile feedback experience.
Patent Information
- Application Number
- CN202211283874.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-10-20
AI Technical Summary
Existing technologies are unable to effectively generate accurate vibration feedback based on the rhythm and beat of music, especially in music with slower tempos, resulting in limited user experience.
A deep learning model is used to process music tracks, and accurate vibration signals are generated by calculating the energy proportion and weighting rules of different tracks. This includes short-time Fourier transform and time-frequency spectrum analysis to generate vibration feedback that accurately matches the rhythm and beat of the music.
This achieves vibration output that more accurately matches the rhythm and beat of music, enhancing the user's tactile feedback experience.
Smart Images

Figure CN116185167B_ABST
Abstract
Description
Technical field
[0001] The present invention relates to the application field of deep learning technology, and in particular to a tactile feedback method, system and related equipment for music track matching vibration. [Background Technology]
[0002] Music can express emotions such as joy, sorrow, anger, and strength through varying rhythms, rhymes, and speeds. Haptic feedback technology that matches vibrations to the music's tempo, emphasis, and rhythmicity provides listeners with a more immersive and authentic experience. Different musical styles incorporate varying instrumental components, each contributing to a piece's rhythmic and melodic analysis. For example, the rhythmic nature of percussion instruments makes it easier to capture the music's rhythm and flow, allowing for more precise vibration feedback.
[0003] In related technologies, methods that use the characteristics of music itself to generate vibrations often use rhythmic instruments such as drum beats to generate corresponding vibrations. However, this method is not suitable for music with slower rhythms. At the same time, existing technologies cannot generate vibrations of corresponding vibration levels by analyzing the strength of different rhythms in music, and the vibration feedback experience brought to users is relatively limited.
[0004] Therefore, it is necessary to provide a new tactile feedback method to obtain vibration output that more accurately matches the rhythm and beat of music. [Summary of the invention]
[0005] The technical problem to be solved by the present invention is to provide a method for producing a vibration output that more accurately matches the rhythm and beat of music.
[0006] To solve the above technical problems, in a first aspect, the present invention provides a tactile feedback method for music track matching vibration, the tactile feedback method is based on a deep learning model, and the tactile feedback method includes the following steps:
[0007] Get the original audio data;
[0008] Using a preset deep learning model to split the original audio data into tracks to obtain multiple tracks of audio data;
[0009] Calculate the energy proportion of each sub-track audio data in the original audio data;
[0010] Determine the weight of each corresponding track audio data according to the energy proportion;
[0011] Perform weighted calculation on all the sub-track audio data according to a preset weighting rule to obtain a time-frequency spectrum and output it;
[0012] generating a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum;
[0013] outputting the matching vibration signal as a driving signal of a driver to realize a haptic feedback effect.
[0014] Preferably, the step of calculating the energy proportion of each of the sub-track audio data in the original audio data specifically comprises:
[0015] performing short-time Fourier transform processing on each of the sub-track audio data to obtain corresponding transformed sub-track audio data;
[0016] calculating the energy proportion of the transformed sub-track audio data in the original audio data.
[0017] Preferably, the step of generating a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum specifically comprises:
[0018] performing normalization processing on the time-frequency spectrum to obtain a time-frequency curve;
[0019] setting vibration information corresponding to a part greater than a preset frequency threshold in the time-frequency curve;
[0020] outputting the time-frequency curve containing the vibration information as the matching vibration signal.
[0021] Preferably, the sub-track audio data at least comprises a first audio track, a second audio track, a third audio track and a fourth audio track, which are different in audio track characteristics.
[0022] Preferably, the preset weighting rule specifically comprises:
[0023] determining whether the energy proportion of the first audio track is the largest:
[0024] if yes:
[0025] determining whether the energy proportion of the second audio track is the second largest: if the energy proportion of the second audio track is the second largest, weighting and outputting the time-frequency spectrum of the first audio track and the second audio track; if the energy proportion of the second audio track is not the second largest, only outputting the time-frequency spectrum of the first audio track;
[0026] if no:
[0027] determining whether the energy proportion of the second audio track is the largest:
[0028] If the energy proportion of the second audio track is the maximum: determine whether the energy proportion of the first audio track is the second largest; if the energy proportion of the first audio track is the second largest, take the time-frequency spectrum weight of the first audio track and the second audio track as output; if the energy proportion of the first audio track is not the second largest, only take the time-frequency spectrum of the second audio track as output;
[0029] If the energy proportion of the second audio track is not the maximum: determine whether the energy proportion of the third audio track is the maximum; if the energy proportion of the third audio track is not the maximum, take the time-frequency spectrum of the fourth audio track as output; if the energy proportion of the third audio track is the maximum, take the time-frequency spectrum of the third audio track as output.
[0030] Preferably, the first audio track is percussion, the second audio track is other instrument audio track, the third audio track is human voice track, and the fourth audio track is bass track.
[0031] In a second aspect, the present application further provides a music track matching vibration tactile feedback system, comprising:
[0032] An original audio acquisition module is configured to acquire original audio data.
[0033] A track separation module is configured to separate the original audio data by using a preset deep learning model to obtain a plurality of separated audio data.
[0034] A proportion calculation module is configured to calculate an energy proportion corresponding to each of the separated audio data in the original audio data.
[0035] A weight calculation module is configured to determine a weight of each of the separated audio data according to the energy proportion.
[0036] A weighting calculation module is configured to perform weighting calculation on all the separated audio data according to a preset weighting rule to obtain a time-frequency spectrum and output.
[0037] A matching vibration module is configured to generate a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum.
[0038] A tactile feedback module is configured to output the matching vibration signal as a driving signal of a driver to realize a tactile feedback effect.
[0039] In a third aspect, the present application further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to realize the steps of the music track matching vibration tactile feedback method according to any one of the above aspects.
[0040] In a fourth aspect, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the music track matching vibration haptic feedback method according to any one of the above aspects.
[0041] Compared with the related art, in the haptic feedback method, the music is processed by a preset deep learning model to distinguish different audio tracks with large feature differences, the importance of different audio tracks in the original audio is determined according to the energy proportion of the different audio tracks, different weights are set according to the importance, the different audio tracks are flexibly weighted and combined, the audio data is matched with the vibration, and finally the vibration output that is more accurately matched with the rhythm and beat of the audio data is output, so that the user can obtain a better haptic feedback experience. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.
[0043] Figure 1 is a step flowchart of the music track matching vibration haptic feedback method provided by the embodiment of the present application;
[0044] Figure 2 is a structure diagram of the deep learning model provided by the embodiment of the present application;
[0045] Figure 3 is a diagram of the preset weighting rule provided by the embodiment of the present application;
[0046] Figure 4 is a diagram of the audio track processed by the deep learning model provided by the embodiment of the present application;
[0047] Figure 5 is a time-frequency spectrum comparison diagram of each audio track provided by the embodiment of the present application;
[0048] Figure 6 is a matching vibration signal diagram provided by the embodiment of the present application;
[0049] Figure 7 is a structure diagram of the haptic feedback effect generation system 200 provided by the embodiment of the present application;
[0050] Figure 8 is a structure diagram of the computer device provided by the embodiment of the present application.
DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0052] Please refer to Figure 1 , Figure 1 is a step flowchart of a music track matching vibration haptic feedback method provided by the embodiments of the present application. The haptic feedback method comprises the following steps:
[0053] S1, obtaining original audio data.
[0054] Specifically, the original audio data obtained by the embodiments of the present application is not specifically limited in the form of music it represents, such as pop music, rock music, symphony music, etc. The method for obtaining the original audio data includes but is not limited to: obtaining from existing audio data, or converting into a separate audio data file after real-time extraction by a recorder, a video camera, etc.
[0055] S2, using a preset deep learning model to separate the original audio data to obtain a plurality of separated audio data.
[0056] Specifically, the deep learning model is a neural network model for separating various different characteristic audios in the audio data. In the embodiments of the present application, the structure of the deep learning model used for separating the original audio data is as shown in Figure 2 The deep learning model comprises an encoding layer composed of a plurality of encoders, a neural network recurrent layer comprising an LSTM (Long Short-Term Memory) structure, and a decoding layer comprising a plurality of decoders. In the neural network recurrent layer, different LSTM modules can be set as needed to extract different characteristic audio tracks.
[0057] Preferably, the separated audio data comprises at least a first track, a second track, a third track and a fourth track, which are different in track characteristics.
[0058] S3, calculating the energy proportion of each of the separated audio data in the original audio data.
[0059] Preferably, the step of calculating the energy proportion of each of the separated audio data in the original audio data is specifically:
[0060] performing short-time Fourier transform on each of the split audio data to obtain corresponding transformed split audio data;
[0061] calculating the energy proportion of the transformed split audio data in the original audio data.
[0062] S4, determining the weight of each corresponding split audio data according to the energy proportion.
[0063] S5, performing weighted calculation on all the split audio data according to a preset weighting rule to obtain a time-frequency spectrum and output.
[0064] Preferably, the preset weighting rule is specifically as follows:
[0065] one of all the split audio data is used as the track used for generating the time-frequency spectrum.
[0066] Specifically, in one possible embodiment, four kinds of split audio data, the preset weighting rule is specifically as follows:
[0067] determining whether the energy proportion of the first track is the largest:
[0068] if yes:
[0069] determining whether the energy proportion of the second track is the second largest: if the energy proportion of the second track is the second largest, taking the time-frequency spectrum of the first track and the second track as the output; if the energy proportion of the second track is not the second largest, only taking the time-frequency spectrum of the first track as the output;
[0070] if no:
[0071] determining whether the energy proportion of the second track is the largest:
[0072] if the energy proportion of the second track is the largest: determining whether the energy proportion of the first track is the second largest: if the energy proportion of the first track is the second largest, taking the time-frequency spectrum of the first track and the second track as the output; if the energy proportion of the first track is not the second largest, only taking the time-frequency spectrum of the second track as the output;
[0073] if the energy proportion of the second track is not the largest: determining whether the energy proportion of the third track is the largest: if the energy proportion of the third track is not the largest, taking the time-frequency spectrum of the fourth track as the output; if the energy proportion of the third track is the largest, taking the time-frequency spectrum of the third track as the output.
[0074] Preferably, the first audio track is percussion, the second audio track is other instrument audio track, the third audio track is human voice audio track, and the fourth audio track is bass audio track. Please refer to Figure 3 , Figure 3 is a schematic diagram of the preset weighting rule provided by the embodiment of the present application. The bass audio track is the part with lower frequency in the audio. Correspondingly, the audio also includes middle tone and high tone. For the user, the listening experience brought by the bass change is more intense than that of the middle tone and high tone. The percussion and the instrument are the parts that emphasize the rhythm speed in the audio. Among them, the percussion is embodied as a regular frequency fluctuation, and the instrument other than percussion often embodies the type of music by combining with percussion. The human voice audio track is relatively special in the audio because the human voice does not have regularity, but the feedback of the human voice embodied in the music is also a great influence on the user experience. It should be noted that the number of audio tracks specifically divided in the embodiment of the present application can be flexibly changed.
[0075] According to the above-described preset weighting rule, the embodiment of the present application can take at least one of the audio data with the largest energy proportion in the audio data as the basis data of the time-frequency spectrum, so that the time-frequency spectrum focuses more on embodying the characteristics of the audio data that need to be matched to generate vibration feedback.
[0076] S6, generating a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum.
[0077] Preferably, the step of generating a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum specifically comprises:
[0078] normalizing the time-frequency spectrum to obtain a time-frequency curve;
[0079] corresponding setting vibration information for the part greater than the preset frequency threshold in the time-frequency curve;
[0080] outputting the time-frequency curve containing the vibration information as the matching vibration signal.
[0081] S7, outputting the matching vibration signal as a driving signal of a driver to realize a haptic feedback effect.
[0082] In the embodiment of the present application, the haptic feedback effect needs to be realized by a vibration feedback system with a motor-based driver.
[0083] For example, please refer to Figure 4 , Figure 4 is a schematic diagram of the audio track after being divided by the deep learning model in the embodiment of the present application, Figure 4 The audio tracks in the above-mentioned embodiment of the present application from top to bottom are: original audio data, bass audio track, percussion audio track, other instrument audio track, and human voice audio track. For comparison, please refer toFigure 5 As can be seen from the time-frequency spectrum contrast diagram of each audio track shown, the plurality of track audio data obtained by track separation from the original audio data has a large difference in the corresponding energy proportion due to the different basic track characteristics, and the matching vibration signal generated after weighting according to the preset weighting rule in the embodiment of the present application according to the different energy proportions is as shown in the figure. Figure 6 The first row is the general vibration signal without processing, and the third row is the matching vibration signal generated after weighting in the embodiment of the present application.
[0084] Compared with the related art, in the haptic feedback method of the present application, the music is processed by track separation through a preset deep learning model, different audio tracks with large feature differences are distinguished, the importance of different audio tracks in the original audio is determined according to the energy proportion, different weights are set for flexible weighting combination of different audio tracks, the audio data is matched with vibration, and finally the vibration output that is more accurately matched with the rhythm, beat, etc. of the audio data is output, so that the user can obtain a better haptic feedback experience.
[0085] The embodiment of the present application also provides a music track separation matching vibration haptic feedback system, please refer to Figure 7 , Figure 7 is a structural schematic diagram of the music track separation matching vibration haptic feedback system 200 provided by the embodiment of the present application, which comprises:
[0086] The original audio acquisition module 201 is configured to acquire original audio data.
[0087] The track separation module 202 is configured to separate the original audio data by using a preset deep learning model to obtain a plurality of track audio data.
[0088] The proportion calculation module 203 is configured to calculate the energy proportion of each track audio data in the original audio data.
[0089] The weight calculation module 204 is configured to determine the weight of each corresponding track audio data according to the energy proportion.
[0090] The weighting calculation module 205 is configured to calculate the weight of all track audio data according to a preset weighting rule to obtain a time-frequency spectrum and output.
[0091] The matching vibration module 206 is configured to generate a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum.
[0092] The haptic feedback module 207 is configured to output the matching vibration signal as a driving signal of a driver to realize a haptic feedback effect.
[0093] The music track matching vibration haptic feedback system 200 provided by the embodiment of the present application can realize the steps in the music track matching vibration haptic feedback method in the above embodiment, and can realize the same technical effects. For details, refer to the description in the above embodiment, which will not be repeated here.
[0094] The embodiment of the present application also provides a computer device, please refer to Figure 8 , Figure 8 is a structural schematic diagram of the computer device provided by the embodiment of the present application. The computer device 300 comprises a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301.
[0095] Please combine Figure 1 , the processor 301 calls the computer program stored in the memory 302, and realizes the steps in the music track matching vibration haptic feedback method in the above embodiment when executing the computer program, comprising:
[0096] obtain the original audio data;
[0097] track the original audio data using a preset deep learning model to obtain a plurality of track audio data;
[0098] calculate the energy proportion of each track audio data in the original audio data;
[0099] determine the weight of each corresponding track audio data according to the energy proportion;
[0100] According to the preset weighting rule, all the track audio data are weighted and calculated to obtain a time-frequency spectrum and output;
[0101] generate a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum;
[0102] output the matching vibration signal as a driver driving signal to realize the haptic feedback effect.
[0103] Preferably, the step of calculating the energy proportion of each track audio data in the original audio data is specifically:
[0104] performing short-time Fourier transform processing on each track audio data to obtain corresponding transformed track audio data;
[0105] calculate the energy proportion of the transformed track audio data in the original audio data.
[0106] Preferably, the step of generating a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum is specifically:
[0107] normalizing the time-frequency spectrum to obtain a time-frequency curve;
[0108] setting vibration information corresponding to a part greater than a preset frequency threshold in the time-frequency curve;
[0109] outputting the time-frequency curve containing the vibration information as the matching vibration signal.
[0110] Preferably, the split-track audio data at least includes a first audio track, a second audio track, a third audio track and a fourth audio track, each of which has different audio track features.
[0111] Preferably, the preset weighting rule is specifically:
[0112] determining whether the energy proportion of the first audio track is the largest:
[0113] if yes:
[0114] determining whether the energy proportion of the second audio track is the second largest: if the energy proportion of the second audio track is the second largest, weighting the time-frequency spectrum of the first audio track and the second audio track and taking the weighted time-frequency spectrum as output; if the energy proportion of the second audio track is not the second largest, only taking the time-frequency spectrum of the first audio track as output;
[0115] if no:
[0116] determining whether the energy proportion of the second audio track is the largest:
[0117] if the energy proportion of the second audio track is the largest: determining whether the energy proportion of the first audio track is the second largest: if the energy proportion of the first audio track is the second largest, weighting the time-frequency spectrum of the first audio track and the second audio track and taking the weighted time-frequency spectrum as output; if the energy proportion of the first audio track is not the second largest, only taking the time-frequency spectrum of the second audio track as output;
[0118] if the energy proportion of the second audio track is not the largest: determining whether the energy proportion of the third audio track is the largest: if the energy proportion of the third audio track is not the largest, taking the time-frequency spectrum of the fourth audio track as output; if the energy proportion of the third audio track is the largest, taking the time-frequency spectrum of the third audio track as output.
[0119] Preferably, the first audio track is percussion, the second audio track is other instrument audio track, the third audio track is human voice audio track, and the fourth audio track is bass audio track.
[0120] The computer device 300 provided by the embodiment of the present application can realize the steps in the music track matching vibration haptic feedback method in the above embodiment, and can realize the same technical effects. For details, refer to the description in the above embodiment, which will not be repeated here.
[0121] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to realize each process and step in the music track matching vibration haptic feedback method provided by the embodiment of the present application, and can realize the same technical effects. To avoid repetition, details will not be repeated here.
[0122] The above only describes the embodiments of the present application. It should be noted that, for those skilled in the art, improvements can be made without departing from the concept of the present application, and these improvements are within the protection scope of the present application.
Claims
1. A method of haptic feedback matching vibrations to music stem, characterized by, The haptic feedback method is based on a deep learning model, and the haptic feedback method comprises the following steps: Obtain original audio data; Split the original audio data using a preset deep learning model to obtain multiple split audio data; Calculate the energy proportion of each split audio data in the original audio data; Determine the weight of each corresponding split audio data according to the energy proportion; According to the preset weighting rule, all the split audio data are weighted and calculated to obtain a time-frequency spectrum and output; According to the time-frequency spectrum, a matching vibration signal corresponding to the original audio data is generated; The matching vibration signal is output as a driving signal of a vibrator to realize a haptic feedback effect.
2. The method of claim 1, wherein, The step of calculating the energy proportion of each split audio data in the original audio data is specifically: Performing short-time Fourier transform processing on each split audio data to obtain corresponding transformed split audio data; Calculate the energy proportion of the transformed split audio data in the original audio data.
3. The method of claim 1, wherein, The step of generating a matching vibration signal corresponding to the original audio data according to the time-frequency spectrum is specifically: Performing normalization processing on the time-frequency spectrum to obtain a time-frequency curve; Corresponding vibration information is set for the part greater than the preset frequency threshold in the time-frequency curve; The time-frequency curve containing the vibration information is output as the matching vibration signal.
4. The method of claim 1, wherein, The split audio data at least includes a first audio track, a second audio track, a third audio track and a fourth audio track, which are different in audio track characteristics.
5. The method of claim 4, wherein, The preset weighting rule is specifically: Determine whether the energy proportion of the first audio track is the largest: If yes: Determine whether the energy proportion of the second audio track is the second largest: if the energy proportion of the second audio track is the second largest, the time-frequency spectrum of the first audio track and the second audio track is weighted and taken as output; if the energy proportion of the second audio track is not the second largest, only the time-frequency spectrum of the first audio track is taken as output; If no: Determine whether the energy proportion of the second audio track is the largest: If the energy proportion of the second audio track is the largest: determine whether the energy proportion of the first audio track is the second largest: if the energy proportion of the first audio track is the second largest, the time-frequency spectrum of the first audio track and the second audio track is weighted and taken as output; if the energy proportion of the first audio track is not the second largest, only the time-frequency spectrum of the second audio track is taken as output; If the energy proportion of the second audio track is not the largest: determine whether the energy proportion of the third audio track is the largest: if the energy proportion of the third audio track is not the largest, the time-frequency spectrum of the fourth audio track is taken as output; if the energy proportion of the third audio track is the largest, the time-frequency spectrum of the third audio track is taken as output.
6. The method of claim 4, wherein, The first audio track is percussion, the second audio track is other instrument audio track, the third audio track is human voice audio track, and the fourth audio track is bass audio track.
7. A haptic feedback system that matches vibrations to music tracks, characterized by, It comprises: An original audio acquisition module for acquiring original audio data; A track splitting module is configured to split the original audio data by using a preset deep learning model to obtain a plurality of split audio data; A proportion calculation module is configured to calculate an energy proportion of each of the split audio data in the original audio data; A weight calculation module is configured to determine a weight of each of the split audio data according to the energy proportion; A weighted calculation module is configured to perform weighted calculation on all the split audio data according to a preset weighted rule to obtain a time-frequency spectrum and output the time-frequency spectrum; A matching vibration module is configured to generate a matching vibration signal of the original audio data according to the time-frequency spectrum; A haptic feedback module is configured to output the matching vibration signal as a driving signal of a driver to realize a haptic feedback effect.
8. A computer device, characterized in that: The application discloses a music track splitting matching vibration haptic feedback method and device. The application discloses a music track splitting matching vibration haptic feedback method and device.
9. A computer-readable storage medium, characterized in that, The application discloses a music track splitting matching vibration haptic feedback method and device.
Citation Information
Patent Citations
A method of extracting features from songs and transforming them into tactile sensations
CN109144257A
Tactile feedback method
CN109871120A