Sound effect adjustment method, device, headphone device and computer-readable storage medium

The user's footstep rhythm data is obtained through the headphone device, and the music playback speed is adjusted to make the music rhythm consistent with the movement rhythm, solving the problem of inconsistent rhythm during exercise and improving the user's exercise experience.

CN116347285BActive Publication Date: 2025-09-19GOERTEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310149750.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-09-19
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

When users exercise, the rhythm of the music played in the headphones is inconsistent with their own movement rhythm, which disrupts the movement rhythm and affects the exercise experience.

Method used

The user's footstep rhythm data is obtained through the headphone device, the next step is predicted, and the music playback speed is adjusted to make the music rhythm consistent with the action rhythm, specifically by adjusting the expected playback time and sampling rate of the target audio point.

Benefits of technology

The music rhythm is synchronized with the user's movement rhythm, which avoids the movement rhythm being affected by the music rhythm and improves the exercise experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116347285B_ABST
    Figure CN116347285B_ABST
Patent Text Reader

Abstract

The present invention discloses a sound effect adjustment method, device, headphone device, and computer-readable storage medium. The sound effect adjustment method is applied to the headphone device and includes: after playing a frame of historical music data in a target music audio, obtaining the user's footstep rhythm data and predicting the user's next footfall time based on the footstep rhythm data; using the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determining the expected playback time of the target audio point when playing the next frame of music data at the expected playback speed; and adjusting the expected playback speed based on the expected playback time and the next footfall time so that the target audio point is played at the next footfall time. The present invention achieves the goal of aligning the music rhythm played by the headphone with the user's movement rhythm, thereby improving the user's exercise experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of earphone technology, and in particular to a sound effect adjustment method, device, earphone equipment, and computer-readable storage medium. Background Art

[0002] As people's living standards improve, exercise and fitness have gradually become important needs. For example, simple and easy exercise methods such as running, walking, and skipping rope are becoming more and more popular. Currently, more and more people like to wear headphones and listen to music while exercising to relieve the boring exercise process. However, users usually have a fixed movement rhythm when exercising, and music also has its own musical rhythm. If users use headphones to listen to music while exercising, the music played by the headphones may not match the user's movement rhythm, resulting in the user's movement rhythm being affected by the music rhythm during exercise, disrupting the user's movement rhythm and affecting the user's exercise experience. Summary of the Invention

[0003] The main purpose of the present invention is to provide a sound effect adjustment method, device, headphone device and computer-readable storage medium, aiming to make the music rhythm played by the headphones consistent with the user's movement rhythm, thereby improving the user's exercise experience.

[0004] To achieve the above-mentioned object, the present invention provides a sound effect adjustment method, which is applied to a headphone device and includes:

[0005] The sound effect adjustment method is applied to a headphone device, and the sound effect adjustment method includes:

[0006] After playing a frame of historical music data in the target music audio, obtaining the user's footstep rhythm data and predicting the user's next step-down time based on the footstep rhythm data;

[0007] Taking the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determining the expected playback time of the target audio point when the next frame of music data is played at the expected playback speed;

[0008] The estimated playback speed is adjusted based on the estimated playback time and the next landing time so that the target audio point is played at the next landing time.

[0009] Optionally, the step of adjusting the estimated playback speed based on the estimated playback time and the next landing time includes:

[0010] If the estimated playback time is before the next landing time, reducing the estimated playback speed;

[0011] If the expected playback time is after the next landing time, the expected playback speed is increased.

[0012] Optionally, the step of adjusting the estimated playback speed based on the estimated playback time and the next landing time includes:

[0013] Determining a playback time interval between a playback end time of the historical music data and the estimated playback time;

[0014] Calculate the landing time interval between the playback end time and the next landing time based on the playback end time and the next landing time;

[0015] Calculate a target sampling rate based on the playback time interval, the landing time interval, and the basic sampling rate of the target music audio;

[0016] The target music audio is adjusted according to the target sampling rate to adjust the expected playback speed.

[0017] Optionally, the step of determining the playing time interval between the playing end time of the historical music data and the expected playing time includes:

[0018] Determining the position of the last sampling point in the historical music data in the target music audio as the first position in the target music audio corresponding to the playback end moment;

[0019] Determining that the estimated playing time corresponds to a second position of the target music audio;

[0020] The absolute value of the difference between the first position and the second position is divided by the basic sampling rate of the target music audio to obtain the playback time interval between the playback end time and the expected playback time.

[0021] Optionally, the step of obtaining the user's footstep rhythm data and predicting the user's next stepping moment according to the footstep rhythm data includes:

[0022] Acquiring an ambient audio signal of the user's environment through an external ear microphone of the headset;

[0023] Performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, wherein the footstep rhythm data is the time of footstep sampling points;

[0024] Determining the time corresponding to the last footstep sampling point in the footstep rhythm data as the user's last footstep time, and calculating the user's average footstep interval based on the time of each footstep sampling point;

[0025] The sum of the previous landing time and the average landing interval is determined as the next landing time of the user.

[0026] Optionally, the step of performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data containing a plurality of discrete points includes:

[0027] amplifying a frequency band where a preset footstep sound is located in the ambient audio signal to obtain filtered data;

[0028] Audio rhythm detection is performed on the filtered data to obtain footstep rhythm data.

[0029] Optionally, the step of taking the first accent in the next frame of music data of the historical music data in the target music audio as the target audio point includes:

[0030] Performing audio rhythm detection on the target music audio to obtain music rhythm data, wherein the music rhythm data is the position of each accent sampling point in the target music audio;

[0031] The first accent sampling point of the next frame of music data of the historical music data is determined according to the music rhythm data, and the first accent sampling point of the next frame of music data is used as the target audio point.

[0032] To achieve the above-mentioned object, the present invention further provides a sound effect adjustment device, which is deployed in a headphone device and includes:

[0033] an acquisition module, configured to acquire the user's footstep rhythm data after playing a frame of historical music data in the target music audio and predict the user's next step-down time based on the footstep rhythm data;

[0034] a determination module for determining a first accent sampling point in a next frame of music data of the historical music data in the target music audio as a target audio point, and determining an estimated playback time of the target audio point when the next frame of music data is played at an estimated playback speed;

[0035] An adjustment module is used to adjust the expected playback speed based on the expected playback time and the next landing time so that the target audio point is played at the next landing time.

[0036] To achieve the above-mentioned purpose, the present invention also provides a headphone device, which includes: a memory, a processor, and a sound effect adjustment program stored in the memory and executable on the processor. When the sound effect adjustment program is executed by the processor, the steps of the above-mentioned sound effect adjustment method are implemented.

[0037] In addition, to achieve the above-mentioned purpose, the present invention also proposes a computer-readable storage medium, on which a sound effect adjustment program is stored. When the sound effect adjustment program is executed by a processor, the steps of the above-mentioned sound effect adjustment method are implemented.

[0038] In the present invention, after playing a frame of historical music data in the target music audio, the headphone device obtains the user's footstep rhythm data and predicts the user's next landing time based on the footstep rhythm data, takes the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determines the expected playback time of the target audio point when playing the next frame of music data at the expected playback speed, and adjusts the expected playback speed based on the expected playback time and the next landing time so that the target audio point is played at the next landing time.

[0039] The present invention allows the rhythm of music played by headphones to be adjusted according to the user's movement rhythm, so that the music rhythm is consistent with the user's movement rhythm, avoiding the influence of the music rhythm on the user's movement rhythm during exercise, and improving the user's exercise experience. In addition, the present invention adjusts the music rhythm. Compared with the user adjusting their own movement rhythm according to the music rhythm, the present invention eliminates the need for the user to frequently change the movement rhythm according to the music rhythm, thereby improving the user's exercise experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 1. It is a flowchart of a first embodiment of a sound effect adjustment method of the present invention;

[0041] Figure 2 1 is a flow chart of an embodiment of a sound effect adjustment method according to the present invention;

[0042] Figure 3 Schematic diagram of the functional modules of a preferred embodiment of the sound effect adjustment device of the present invention.

[0043] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0044] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] Reference Figure 1 , Figure 1 FIG. 1 is a flow chart of a first embodiment of a sound effect adjustment method according to the present invention.

[0046] The present invention provides an embodiment of a method for adjusting sound effects. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order. The method for adjusting sound effects according to the present invention is applied to a headphone device. In this embodiment, the method for adjusting sound effects includes:

[0047] Step S10, after playing a frame of historical music data in the target music audio, obtaining the user's footstep rhythm data and predicting the user's next stepping moment based on the footstep rhythm data;

[0048] When exercising, users usually have a fixed movement rhythm, and music also has its own rhythm. If users listen to music while exercising, the user's movement rhythm during exercise may be affected by the rhythm of the music, causing the user's movement rhythm to be disrupted, thereby affecting the user's exercise experience.

[0049] Moreover, since the rhythm of the same music may change and the rhythm of different music may also be very different, if the user actively adjusts the rhythm of his or her movements according to the rhythm of the music, the user will need to frequently adjust the rhythm of his or her movements, which will also affect the user's exercise experience.

[0050] Therefore, in this embodiment, the music rhythm of the music played by the headphones is adjusted so that the music rhythm of the music listened to by the user during exercise is consistent with the user's movement rhythm, avoiding the user's movement rhythm during exercise being affected by the music rhythm, thereby improving the user's exercise experience.

[0051] Specifically, in this embodiment, the target music audio currently being played by the earphone is referred to as the target music audio. The target music audio comprises multiple frames of music data, with no limit on the number of sampling points in each frame. A frame of music data that has already been played is referred to as historical music data. After playing a frame of historical music data in the target music audio, the earphone plays the next frame of music data, at which point the tempo of the next frame of music data is adjusted.

[0052] In this embodiment, after playing a frame of historical music data in the target music audio, the earphones obtain the user's footstep rhythm data, where the footstep rhythm data can be the moment of each footstep of the user or the time interval between the footsteps of the user. Specifically, in one embodiment, the user's movement rhythm can be detected by a motion sensor provided on the earphones, and the moment of each footstep of the user is determined based on the movement rhythm, thereby obtaining the user's footstep rhythm data; in another embodiment, the user's environment audio signal can be obtained by an external ear microphone provided on the earphones, and the audio rhythm of the environment audio signal is detected to obtain the user's footstep rhythm data. The specific embodiment is not limited here and can be set according to actual needs.

[0053] After obtaining the user's footstep rhythm data, the user's next landing moment (hereinafter referred to as the next landing moment for distinction) is predicted based on the footstep rhythm data. Specifically, in one embodiment, when the footstep rhythm data is the moment of each user's landing, the time interval between each of the user's landings can be calculated based on the footstep rhythm data, and the last moment in the footstep rhythm data and the time interval between each of the user's landings can be added to obtain the next landing moment; in another embodiment, when the footstep rhythm data is the time interval between the user's landings, the last detected landing moment can be determined as the previous landing moment, and the next landing moment can be obtained by adding the previous landing moment and the time interval between the user's landings. The specific setting can be made according to different footstep rhythms and is not limited here.

[0054] Step S20, taking the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determining the expected playback time of the target audio point when the next frame of music data is played at the expected playback speed;

[0055] The musical rhythm of the target music audio can be reflected by the accent in the target music audio. Therefore, in this embodiment, the sampling point of the first accent in the next frame of music data of the historical music data in the target music audio (hereinafter referred to as the accent sampling point for distinction) is used as the first rhythm point that needs to be adjusted (hereinafter referred to as the target audio point for distinction).

[0056] Specifically, in one embodiment, it can be the first accent played after being determined according to the score of the target music audio. After determining the accent, the accent is matched to the position in the target music audio, and the target audio point is determined according to the position. Specifically, the accent is also called a drum beat, which is generally the first beat in a beat music. In this embodiment, the specific process of determining the accent can be: determining the score corresponding to a frame of music data received at the current moment, and determining the first accent in the score corresponding to the frame of music data as the first accent played after the current moment; in another embodiment, audio rhythm detection can also be performed on the target music audio to obtain music rhythm data, and the first accent sampling point of the next frame of data is determined according to the music rhythm data. It can be specifically set according to actual needs and is not limited here.

[0057] In this embodiment, the estimated playback time can be determined by calculation based on the historical music data, the target music audio and the music rhythm data of the target music audio. In one embodiment, the time interval between the time when the historical music data is played (hereinafter referred to as the playback end time for distinction) and the estimated playback time (hereinafter referred to as the playback time interval for distinction) can be calculated based on the music rhythm data and the target music audio, and the estimated playback time is calculated based on the playback time interval and the playback end time; in another embodiment, the playback time difference of the historical music data can be calculated based on the playback time of the historical music data without speed adjustment and the playback time of the historical music data after speed adjustment, the position of the target audio point in the target audio is determined based on the music rhythm data, the basic playback time of the target audio point when playing at the playback speed of the target music audio without speed adjustment is determined based on the position of the target audio point in the target audio and the time when the target music audio starts playing, and the estimated playback time is calculated using the playback time difference plus the basic playback time. The specific setting can be based on the actual needs of the user and is not limited here.

[0058] Step S30: adjusting the estimated playback speed based on the estimated playback time and the next landing time, so that the target audio point is played at the next landing time.

[0059] In this embodiment, the playback speed of the target music audio after playing a frame of historical music data is referred to as the expected playback speed. Before adjusting the target music audio based on the expected playback time and the next landing time, the expected playback speed can be the playback speed of the historical music data or the initial playback speed of the target music audio. After determining the target audio point, the expected playback speed can be adjusted based on the expected playback time and the next landing time so that the target audio point plays at the next landing time.

[0060] Specifically, in one embodiment, if the expected playback time is before the next landing time, the expected playback speed can be reduced; in another embodiment, if the expected playback time is after the next landing time, the expected playback speed can be increased; in another embodiment, if the expected playback time and the next landing time are consistent, the playback speed can be adjusted according to the initial playback speed of the target music audio. Specifically, if the expected playback speed is the playback speed of historical music data before the target music audio is adjusted based on the expected playback time and the next landing time, the expected playback speed is adjusted; if the expected playback speed is the initial playback speed of the target music audio before the target music audio is adjusted based on the expected playback time and the next landing time, the expected playback speed does not need to be adjusted.

[0061] In a specific embodiment, the adjustment of the expected playback speed can be achieved by adjusting the sampling rate of the target music audio. In this embodiment, the sampling rate of the target music audio when the speed is not adjusted is called the basic sampling rate, and the sampling rate after adjusting the expected playback speed is called the target sampling rate. In one embodiment, the target sampling rate can be calculated based on the basic sampling rate of the target music audio and the time interval between the expected playback moment and the next landing moment; in another embodiment, the target sampling rate can also be calculated based on the basic sampling rate of the target music audio, the playback time interval between the end of playback and the expected playback moment, and the time interval between the end of playback and the next landing moment (hereinafter referred to as the landing time interval for distinction). The specific setting can be based on actual needs and is not limited here.

[0062] Furthermore, in one embodiment, the estimated playback time can be determined by calculation based on the current time and the music rhythm data of the target music audio. The specific calculation process can be: based on the audio point corresponding to the current time, that is, the first audio point of the new frame data and the basic sampling rate of the target music audio, the time interval between the current time and the estimated playback time (hereinafter referred to as the playback time interval for distinction) is calculated, and the estimated playback time is calculated based on the playback time interval and the current time; in another embodiment, it can also be calculated based on the current time, the time when the target music audio starts to play, the duration of the target music audio and the music score of the target music audio. The specific process can be: based on the time when the target music audio starts to play and the current time, determine the time when the new frame of music data received at the current time starts to play, determine the playback duration of the new frame of music data based on the duration of the target music audio, and the estimated playback time can be calculated based on the time when the new frame of music data starts to play, the playback duration of the new frame of music data and the music score of the target music audio. The specific setting can be based on actual needs and is not limited here.

[0063] Furthermore, in some feasible implementations, in the above step S10, the step of obtaining the user's footstep rhythm data and predicting the user's next stepping moment based on the footstep rhythm data includes:

[0064] Step S101, obtaining an ambient audio signal of the user's environment through an external ear microphone of the headset;

[0065] In this embodiment, the audio signal of the user's environment (hereinafter referred to as the ambient audio signal for distinction) is obtained through the earphone's external ear microphone. In a specific embodiment, the external ear microphone can be a feedforward microphone or a call microphone, which is not limited here.

[0066] Step S102, performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, wherein the footstep rhythm data includes the time of footstep sampling points corresponding to the footsteps;

[0067] In this embodiment, after acquiring an ambient audio signal, audio rhythm detection is performed on the ambient audio signal to obtain data representing the rhythm of the user's footsteps during exercise, hereinafter referred to as footstep rhythm data for clarity. The footstep rhythm data includes the time instants of footstep sampling points corresponding to footsteps. Furthermore, in one embodiment, before performing audio rhythm detection on the ambient audio signal, filtering may be performed on the ambient audio signal to enhance the prominence of footsteps, thereby increasing the accuracy of the obtained footstep rhythm data.

[0068] In a specific embodiment, the ambient audio signal may be time domain data. In this embodiment, the process of performing audio rhythm detection on the ambient audio signal using 1024 sampling points in the ambient audio signal as a window may be:

[0069] Perform difference processing on the ambient audio signal, that is, cut the current window data from the previous window data to obtain the difference data. The specific difference processing calculation formula is:

[0070]

[0071] Where i is the sampling point, k is the current window, and k-1 is the previous window.

[0072] Performing Fourier transform on the difference data to obtain frequency domain data, and performing difference processing on the frequency domain data again to obtain frequency domain difference data;

[0073] The frequency domain difference data is quantized to obtain quantized data. In a specific implementation, the data quantization may be performed by a moving average method, which will not be described in detail here.

[0074] After converting the quantized data into time domain data, peak selection is performed on the time domain data to obtain time domain data with multiple discrete points. In the time domain data with multiple discrete points, each discrete point is a selected peak. In a specific implementation, the peak selection can be based on whether the peak amplitude is within a preset amplitude range, or it can be based on whether the peak amplitude is greater than a preset amplitude threshold. The specific setting can be based on actual needs and is not limited here.

[0075] Step S103, determining the time corresponding to the last footstep sampling point in the footstep rhythm data as the user's last footstep time, and calculating the user's average footstep interval based on the time of each footstep sampling point;

[0076] The time corresponding to the last footstep sampling point in the footstep rhythm data is determined as the user's last footstep time, and the user's average footstep interval is calculated based on the time of each footstep sampling point.

[0077] Step S104: Determine the sum of the previous landing time and the average landing interval as the next landing time of the user.

[0078] The sum of the last landing time and the average landing interval is determined as the user's next landing time. The specific calculation formula can be: t2 = t1 + T1, where t2 is the next landing time, t1 is the last landing time, and T1 is the average landing interval.

[0079] In this embodiment, audio detection is performed on the ambient audio signal to obtain footstep rhythm data, which is more accurate than obtaining the footstep rhythm data of the user during exercise through a motion sensor, so that the adjusted music rhythm in the headphones is more in line with the user's actual movement rhythm, thereby improving the user's exercise experience.

[0080] Furthermore, in a feasible implementation manner, the above step S102 includes:

[0081] Step S1021, amplifying the frequency band of the preset footsteps in the ambient audio signal to obtain filtered data;

[0082] In this embodiment, before performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, the ambient audio signal is filtered to amplify the frequency band where the footsteps are located. The data obtained after filtering is called filtered data. In a specific embodiment, the frequency band where the footsteps are located can be set to 30 to 50 Hz.

[0083] Step S1022: Perform audio rhythm detection on the filtered data to obtain footstep rhythm data.

[0084] Audio rhythm detection is performed on the filtered data to obtain footstep rhythm data. Compared to directly performing audio rhythm detection on the ambient audio signal, this embodiment uses filtering to make the frequency band of footsteps more prominent, preventing interference from sounds in other frequency bands on the audio rhythm detection. This improves the accuracy of the footstep rhythm data, allowing the adjustment of the target music audio to more closely match the user's movement rhythm, and enhancing the user's exercise experience.

[0085] Furthermore, in some feasible implementations, in the above step S20: taking the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point includes:

[0086] Step 201: Acquire the target music audio from the earphone, and perform audio rhythm detection on the target music audio to obtain music rhythm data;

[0087] In this embodiment, the target music audio may be subjected to audio rhythm detection to obtain rhythm data of the target music audio (hereinafter referred to as music rhythm data for distinction).

[0088] Specifically, in this embodiment, the target music audio of the earphone is obtained, and the audio rhythm detection is performed on the target music audio to obtain music rhythm data. In a specific embodiment, the specific process of performing audio rhythm detection on the target music audio to obtain music rhythm data can refer to step S102 and will not be repeated here.

[0089] Step S202: determining the first accent sampling point of the next frame of music data of the historical music data according to the music rhythm data, and taking the first accent sampling point of the next frame of music data as the target audio point.

[0090] In this embodiment, the first accent sampling point of the next frame of music data of the historical music data is determined according to the music rhythm data, and the first accent sampling point of the next frame of music data is used as the target audio point.

[0091] In this implementation, the target music audio played in the headphones is captured and audio rhythm detection is performed on the target music audio to obtain music rhythm data. Based on the music rhythm data, the first accent sampling point of the next frame of historical music data is determined, and the first accent sampling point of the next frame of music data is used as the target audio point. Compared to determining the target audio point based on the target music audio's score, the target music audio's playback duration, and the current moment, this implementation determines the target audio point with greater accuracy, allowing adjustments to the target music audio to better align with the user's movement rhythm, thereby improving the user's exercise experience.

[0092] In this embodiment, after playing a frame of historical music data in the target music audio, the headphone device obtains the user's footstep rhythm data and predicts the user's next landing time based on the footstep rhythm data, takes the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determines the expected playback time of the target audio point when playing the next frame of music data at the expected playback speed, and adjusts the expected playback speed based on the expected playback time and the next landing time so that the target audio point is played at the next landing time.

[0093] This embodiment allows the tempo of the music played by the headphones to be adjusted according to the user's movement rhythm, aligning the music tempo with the user's movement rhythm. This prevents the user's movement rhythm from being affected by the music tempo during exercise, thereby improving the user's exercise experience. Furthermore, this embodiment adjusts the music tempo. Compared to users adjusting their movement tempo according to the music tempo, this embodiment eliminates the need for users to frequently change their movement tempo according to the music tempo, thereby improving the user's exercise experience.

[0094] Furthermore, based on the first embodiment, a second embodiment of the sound effect adjustment method of the present invention is proposed. In this embodiment, the step S30 includes:

[0095] Step S301, if the estimated playback time is before the next landing time, reducing the estimated playback speed;

[0096] In this embodiment, whether the expected playback time of the target audio point is consistent with the next landing time is detected. In a specific implementation, the detection can be whether the expected playback time and the next landing time are the same; or the detection can be whether the playback time interval between the expected playback time and the current time and the time interval between the current time and the next landing time (hereinafter referred to as the landing time interval for distinction) are consistent. The specific configuration is not limited here and can be set according to actual needs.

[0097] In this embodiment, if the expected playback time is before the next landing time, that is, the playback time interval is less than the landing time interval, the expected playback speed is reduced so that the target audio point of the target audio point is played at the next landing time. In a specific implementation method, the expected playback speed can be reduced by interpolating the target music audio.

[0098] Step S303: If the expected playback time coincides with the next landing time and the expected playback time is after the next landing time, the target music audio is extracted and processed to increase the playback speed so that the target audio point is played at the next landing time.

[0099] In this embodiment, if the expected playback time is after the next landing time, that is, the playback time interval is greater than the landing time interval, the target music audio is extracted and processed to increase the playback speed so that the target audio point is played at the next landing time.

[0100] Furthermore, in one embodiment, if the expected playback time coincides with the next landing time, there is no need to adjust the target music audio.

[0101] Furthermore, in some feasible implementations, the above step S30 includes:

[0102] Step S304, determining a playback time interval between the current time and the expected playback time of the target audio point;

[0103] In this embodiment, the target sampling rate is calculated based on the basic sampling rate of the target music audio, the time interval between the current moment and the expected playback time, and the time interval between the current moment and the next landing time. Specifically, in this embodiment, the time interval between the current moment and the expected playback time of the target audio point (hereinafter referred to as the playback time interval for distinction) is determined, and the time interval between the current moment and the next landing time (hereinafter referred to as the landing time interval for distinction) is calculated based on the current moment and the next landing time.

[0104] In a specific implementation, the calculation formula for the landing time interval may be: T4=t2-t0, where T4 is the landing time interval, t2 is the next landing time, and t0 is the current time.

[0105] Step S305, calculating a landing time interval between the playback end time and the next landing time based on the playback end time and the next landing time;

[0106] The landing time interval between the playback end time and the next landing time is calculated based on the playback end time and the next landing time.

[0107] Specifically, in one embodiment, the expected playback time may have been determined, and the playback time interval may be calculated by combining the expected playback time and the current time; in another embodiment, the playback time interval may be calculated based on the audio point corresponding to the current time and the initial sampling of the target music audio; or the playback time interval may be calculated based on the current time, the time when the target music audio starts playing, the score of the target music audio, and the playback duration of the target music audio. The specific settings may be made according to actual needs and are not limited here.

[0108] Step S306, calculating a target sampling rate based on the playing time interval, the landing time interval, and the basic sampling rate of the target music audio;

[0109] In this embodiment, after the play time interval and the landing time interval are obtained, the target sampling rate can be calculated based on the play time interval, the landing time interval and the basic sampling rate of the target music audio.

[0110] The specific calculation formula may be: fs2=fs*(T4 / T3), where fs is the basic sampling rate, T4 is the landing time interval, and T3 is the playback time interval.

[0111] Furthermore, in another embodiment, the target sampling rate can also be calculated based on the time interval between the expected playback time and the next landing time (hereinafter referred to as the target time interval). Specifically, the calculation formula can be: fs2 = fs*((T4 / (T5+T4)), where fs is the basic sampling rate, T4 is the landing time interval, and T5 is the target time interval.

[0112] Step S307: Adjust the target music audio according to the target sampling rate to adjust the expected playback speed.

[0113] The target music audio is adjusted according to the target sampling rate to adjust the expected playback speed.

[0114] Specifically, according to the calculation formula, if T4 / T3 is greater than 1, that is, the playback time interval is less than the landing time interval, the target music audio can be interpolated according to the target sampling rate to reduce the expected playback speed so that the target audio point of the target audio point is played at the next landing time; if T4 / T3 is less than 1, that is, the playback time interval is greater than the landing time interval, the target music audio can be extracted according to the target sampling rate to increase the playback speed so that the target audio point is played at the next landing time; if T4 / T3 is equal to 1, the playback time interval is equal to the landing time interval, that is, the expected playback time is consistent with the next landing time, and the target music audio does not need to be adjusted. The specific process of processing according to the target sampling rate can be, in one embodiment, when the target sampling rate is 2, a sampling point is added after every two sampling points of the target music audio; in another embodiment, when the target sampling rate is 1 / 2, a sampling point is extracted after every two sampling points.

[0115] Furthermore, in a feasible implementation manner, in the above step S304, the step of determining the playing time interval between the playing end time of the historical music data and the expected playing time includes:

[0116] Step S3041, determining the position of the last sampling point in the historical music data in the target music audio as the first position in the target music audio corresponding to the playback end time;

[0117] In this embodiment, the playback time interval between the current moment and the expected playback time of the target audio point is determined based on the current moment, the target audio point, and the base sampling rate of the target music audio. Specifically, the position of the last sampling point in the historical music data in the target music audio is determined as the position in the target music audio corresponding to the playback end time (hereinafter referred to as the first position for clarity).

[0118] Step S3042, determining that the estimated playing time corresponds to a second position of the target music audio;

[0119] In this embodiment, the estimated playing time is determined to correspond to the second position of the target music audio (hereinafter referred to as the second position for distinction).

[0120] Step S3043: Use the absolute value of the difference between the first position and the second position to divide by the basic sampling rate of the target music audio to obtain the playback time interval between the playback end time and the expected playback time.

[0121] The playback time interval between the current moment and the expected playback time of the target audio point is calculated based on the current audio moment, the target audio moment and the basic sampling rate of the target music audio.

[0122] The specific calculation formula can be: T3 = (N1–N0) / fs, where T3 is the playback time interval, N1 is the target audio time, N0 is the current audio time, and fs is the basic sampling rate.

[0123] In this embodiment, the playback time interval between the current moment and the expected playback time of the target audio point is determined based on the current moment, the target audio point and the basic sampling rate of the target music audio. Compared with determining the playback time interval based on the current moment, the time when the target music audio starts playing, the playback duration of the target music audio and the music score of the target music audio, the playback time interval determined by this embodiment is more accurate, so that the adjustment of the target music audio is more in line with the user's movement rhythm, thereby improving the user's exercise experience.

[0124] In this embodiment, if the expected playback time is before the next landing time, the target music audio is interpolated to reduce the expected playback speed so that the target audio point is played at the next landing time. If the expected playback time is after the next landing time, the target music audio is extracted to increase the playback speed so that the target audio point is played at the next landing time. This embodiment plays the target audio point at the next landing time, thereby aligning the rhythm of the music listened to by the user with the rhythm of the user's movements, preventing the user's movement rhythm from being disrupted by the rhythm of the music, and improving the user's exercise experience.

[0125] Further, in one embodiment, referring to Figure 2 , Figure 2 This is a flow chart of an embodiment of the sound effect adjustment method of the present invention. In this embodiment, a memory for storing songs (i.e. Figure 2In a specific embodiment, the user can transfer the music file to the 001 module in the headset in advance through an app (application, a third-party application) compatible with the headset on the mobile phone or via Bluetooth transmission. After turning on the dynamic sound effect function, that is, the function of adjusting the music rhythm according to the rhythm of the user's movements, the headset will automatically detect the music files in the memory and play them in a loop.

[0126] The headset is provided with a soft decoder for decoding the music file (i.e. Figure 2 002 shown in ), the soft decoder decodes the music file and outputs the audio to the playback module (i.e. Figure 2 004 playback) and music rhythm detection module (ie Figure 2 005) In a specific embodiment, the soft decoder can adjust the decoding speed according to the consumption speed of the audio data.

[0127] The earphone is provided with a music rhythm detection module for detecting the audio rhythm of the audio data. The module can calculate the time corresponding to each music rhythm point in the audio data. In this embodiment, the external ear microphone provided on the earphone for collecting the ambient audio signal of the user's environment is a feedforward microphone (i.e. Figure 2 The headset is also provided with a motion rhythm detection module (ie, a motion rhythm detection module) for detecting the audio rhythm of the ambient audio signal. Figure 2 007 shown in ).

[0128] The headset is provided with a calculation module for calculating the playback speed (i.e. Figure 2 008 shown in the figure), this module calculates the data output by the music rhythm detection module and the action rhythm detection module to determine the expected playback speed so that the target audio point can be played at the next landing moment. The headset is also provided with a speed change module (i.e., Figure 2 003 shown in ).

[0129] In this embodiment, the specific process of the sound effect adjustment method is as follows:

[0130] The audio file is obtained from the memory by the soft decoder to decode the audio file, and the decoded audio data is transmitted to the music rhythm detection module for audio rhythm detection (i.e. Figure 2 ) to obtain music rhythm data;

[0131] The feedforward microphone obtains the ambient audio signal of the user's environment (i.e. Figure 2The FFmic (FeedForward microphone) shown in the figure collects the environmental sound signal, and performs audio rhythm detection on the environmental audio signal to obtain the footstep rhythm data (i.e. Figure 2 ) using the running cadence detection algorithm shown in ;

[0132] Determine the next landing time of the user after the current time according to the footstep rhythm data, determine the first accent played after the current time in the headphone target music audio as the target audio point, determine the playback time interval between the current time and the expected playback time of the target audio point, and calculate the landing time interval between the current time and the next landing time based on the current time and the next landing time, and calculate the target sampling rate (i.e., the target sampling rate) based on the playback time interval, the landing time interval and the basic sampling rate of the target music audio. Figure 2 Calculated playback speed as shown in );

[0133] Adjust the expected playback speed according to the target sampling rate so that the target audio point is played at the next landing moment (i.e. Figure 2 Shifting process shown in ).

[0134] After the sound effect adjustment is completed, use the speaker of the headset to play (i.e. Figure 2 (see playback in the ).

[0135] In addition, the embodiment of the present invention also provides a sound effect adjustment device, which is deployed in a headphone device. Figure 3 , the sound effect adjustment device includes:

[0136] An acquisition module 10 is configured to acquire the user's footstep rhythm data after playing a frame of historical music data in the target music audio and predict the user's next step-down time based on the footstep rhythm data;

[0137] a determination module 20 for determining a first accent sampling point in a next frame of music data of the historical music data in the target music audio as a target audio point, and determining an estimated playback time of the target audio point when the next frame of music data is played at an estimated playback speed;

[0138] The adjustment module 30 is configured to adjust the estimated playback speed based on the estimated playback time and the next landing time, so that the target audio point is played at the next landing time.

[0139] Furthermore, the adjustment module 30 is further configured to:

[0140] If the estimated playback time is before the next landing time, reducing the estimated playback speed;

[0141] If the expected playback time is after the next landing time, the expected playback speed is increased.

[0142] Furthermore, the adjustment module 30 is further configured to:

[0143] Determining a playback time interval between a playback end time of the historical music data and the estimated playback time;

[0144] Calculate the landing time interval between the playback end time and the next landing time based on the playback end time and the next landing time;

[0145] Calculate a target sampling rate based on the playback time interval, the landing time interval, and the basic sampling rate of the target music audio;

[0146] The target music audio is adjusted according to the target sampling rate to adjust the expected playback speed.

[0147] Furthermore, the adjustment module 30 is further configured to:

[0148] Determining the position of the last sampling point in the historical music data in the target music audio as the first position in the target music audio corresponding to the playback end moment;

[0149] Determining that the estimated playing time corresponds to a second position of the target music audio;

[0150] The absolute value of the difference between the first position and the second position is divided by the basic sampling rate of the target music audio to obtain the playback time interval between the playback end time and the expected playback time.

[0151] Furthermore, the acquisition module 10 is further configured to:

[0152] Acquiring an ambient audio signal of the user's environment through an external ear microphone of the headset;

[0153] Performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, wherein the footstep rhythm data is the time of footstep sampling points;

[0154] Determining the time corresponding to the last footstep sampling point in the footstep rhythm data as the user's last footstep time, and calculating the user's average footstep interval based on the time of each footstep sampling point;

[0155] The sum of the previous landing time and the average landing interval is determined as the next landing time of the user.

[0156] Furthermore, the sound effect adjustment device further includes a filtering module, which is used to:

[0157] amplifying a frequency band where a preset footstep sound is located in the ambient audio signal to obtain filtered data;

[0158] Audio rhythm detection is performed on the filtered data to obtain footstep rhythm data.

[0159] Furthermore, the determining module 20 is further configured to:

[0160] Performing audio rhythm detection on the target music audio to obtain music rhythm data, wherein the music rhythm data is the position of each accent sampling point in the target music audio;

[0161] The first accent sampling point of the next frame of music data of the historical music data is determined according to the music rhythm data, and the first accent sampling point of the next frame of music data is used as the target audio point.

[0162] The expanded content of the specific implementation of the sound effect adjustment device of the present invention is basically the same as the various embodiments of the above-mentioned sound effect adjustment method, and will not be repeated here.

[0163] The headphone device of the present invention comprises a structural housing, a communication module, a main control module (e.g., a microcontroller unit (MCU)), a speaker, and a memory. The main control module may include a microprocessor, an audio decoding unit, a power supply and power management unit, sensors required by the system, and other active or passive components (which may be replaced, deleted, or added based on actual functionality) to implement wireless audio reception and playback. The memory of the headphone may store a sound effect adjustment program, and the microprocessor may be used to call the sound effect adjustment program stored in the memory and perform the following operations:

[0164] The sound effect adjustment method is applied to a headphone device, and the sound effect adjustment method includes:

[0165] After playing a frame of historical music data in the target music audio, obtaining the user's footstep rhythm data and predicting the user's next step-down time based on the footstep rhythm data;

[0166] Taking the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determining the expected playback time of the target audio point when the next frame of music data is played at the expected playback speed;

[0167] The estimated playback speed is adjusted based on the estimated playback time and the next landing time so that the target audio point is played at the next landing time.

[0168] Furthermore, the operation of adjusting the estimated playback speed based on the estimated playback time and the next landing time includes:

[0169] If the estimated playback time is before the next landing time, reducing the estimated playback speed;

[0170] If the expected playback time is after the next landing time, the expected playback speed is increased.

[0171] Furthermore, the operation of adjusting the estimated playback speed based on the estimated playback time and the next landing time includes:

[0172] Determining a playback time interval between a playback end time of the historical music data and the estimated playback time;

[0173] Calculate the landing time interval between the playback end time and the next landing time based on the playback end time and the next landing time;

[0174] Calculate a target sampling rate based on the playback time interval, the landing time interval, and the basic sampling rate of the target music audio;

[0175] The target music audio is adjusted according to the target sampling rate to adjust the expected playback speed.

[0176] Furthermore, the operation of determining the playing time interval between the playing end time of the historical music data and the expected playing time includes:

[0177] Determining the position of the last sampling point in the historical music data in the target music audio as the first position in the target music audio corresponding to the playback end moment;

[0178] Determining that the estimated playing time corresponds to a second position of the target music audio;

[0179] The absolute value of the difference between the first position and the second position is divided by the basic sampling rate of the target music audio to obtain the playback time interval between the playback end time and the expected playback time.

[0180] Furthermore, the operation of obtaining the user's footstep rhythm data and predicting the user's next stepping moment according to the footstep rhythm data includes:

[0181] Acquiring an ambient audio signal of the user's environment through an external ear microphone of the headset;

[0182] Performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, wherein the footstep rhythm data is the time of footstep sampling points;

[0183] Determining the time corresponding to the last footstep sampling point in the footstep rhythm data as the user's last footstep time, and calculating the user's average footstep interval based on the time of each footstep sampling point;

[0184] The sum of the previous landing time and the average landing interval is determined as the next landing time of the user.

[0185] Furthermore, the operation of performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data containing a plurality of discrete points includes:

[0186] amplifying a frequency band where a preset footstep sound is located in the ambient audio signal to obtain filtered data;

[0187] Audio rhythm detection is performed on the filtered data to obtain footstep rhythm data.

[0188] Furthermore, the operation of taking the first accent in the next frame of music data of the historical music data in the target music audio as the target audio point includes:

[0189] Performing audio rhythm detection on the target music audio to obtain music rhythm data, wherein the music rhythm data is the position of each accent sampling point in the target music audio;

[0190] The first accent sampling point of the next frame of music data of the historical music data is determined according to the music rhythm data, and the first accent sampling point of the next frame of music data is used as the target audio point.

[0191] The various embodiments of the headphone device and the computer-readable storage medium of the present invention can all refer to the various embodiments of the sound effect adjustment method of the present invention, and will not be repeated here.

[0192] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0193] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0194] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0195] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A sound effect adjustment method, characterized in that: The sound effect adjustment method is applied to a headphone device, and the sound effect adjustment method includes: After playing a frame of historical music data in the target music audio, obtaining an ambient audio signal of the user's environment through the external ear microphone of the headset; Performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, wherein the footstep rhythm data is the time of footstep sampling points; Determining the time corresponding to the last footstep sampling point in the footstep rhythm data as the user's last footstep time, and calculating the user's average footstep interval based on the time of each footstep sampling point; Determine the sum of the previous landing time and the average landing interval as the next landing time of the user; Taking the first accent sampling point in the next frame of music data of the historical music data in the target music audio as the target audio point, and determining the expected playback time of the target audio point when the next frame of music data is played at the expected playback speed; The estimated playback speed is adjusted based on the estimated playback time and the next landing time so that the target audio point is played at the next landing time.

2. The sound effect adjustment method according to claim 1, wherein: The step of adjusting the estimated playback speed based on the estimated playback time and the next landing time includes: If the estimated playback time is before the next landing time, reducing the estimated playback speed; If the expected playback time is after the next landing time, the expected playback speed is increased.

3. The sound effect adjustment method according to claim 1, wherein: The step of adjusting the estimated playback speed based on the estimated playback time and the next landing time includes: Determining a playback time interval between a playback end time of the historical music data and the estimated playback time; Calculate the landing time interval between the playback end time and the next landing time based on the playback end time and the next landing time; Calculate a target sampling rate based on the playback time interval, the landing time interval, and the basic sampling rate of the target music audio; The target music audio is adjusted according to the target sampling rate to adjust the expected playback speed.

4. The sound effect adjustment method according to claim 3, wherein: The step of determining the playing time interval between the playing end time of the historical music data and the expected playing time comprises: Determining the position of the last sampling point in the historical music data in the target music audio as the first position in the target music audio corresponding to the playback end moment; Determining that the estimated playing time corresponds to a second position of the target music audio; The absolute value of the difference between the first position and the second position is divided by the basic sampling rate of the target music audio to obtain the playback time interval between the playback end time and the expected playback time.

5. The sound effect adjustment method according to claim 1, wherein: The step of performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data containing a plurality of discrete points comprises: amplifying a frequency band where a preset footstep sound is located in the ambient audio signal to obtain filtered data; Audio rhythm detection is performed on the filtered data to obtain footstep rhythm data.

6. The sound effect adjustment method according to any one of claims 1 to 5, wherein: The step of taking the first accent in the next frame of music data of the historical music data in the target music audio as the target audio point comprises: Performing audio rhythm detection on the target music audio to obtain music rhythm data, wherein the music rhythm data is the position of each accent sampling point in the target music audio; The first accent sampling point of the next frame of music data of the historical music data is determined according to the music rhythm data, and the first accent sampling point of the next frame of music data is used as the target audio point.

7. A sound effect adjustment device, characterized in that: The sound effect adjustment device is deployed on a headphone device, and the sound effect adjustment device includes: an acquisition module, configured to acquire the user's footstep rhythm data after playing a frame of historical music data in the target music audio and predict the user's next step-down time based on the footstep rhythm data; a determination module for determining a first accent sampling point in a next frame of music data of the historical music data in the target music audio as a target audio point, and determining an estimated playback time of the target audio point when the next frame of music data is played at an estimated playback speed; an adjusting module, configured to adjust the estimated playback speed based on the estimated playback time and the next landing time, so that the target audio point is played at the next landing time; The acquisition module is further used to: Acquiring an ambient audio signal of the user's environment through an external ear microphone of the headset; Performing audio rhythm detection on the ambient audio signal to obtain footstep rhythm data, wherein the footstep rhythm data is the time of footstep sampling points; Determining the time corresponding to the last footstep sampling point in the footstep rhythm data as the user's last footstep time, and calculating the user's average footstep interval based on the time of each footstep sampling point; The sum of the previous landing time and the average landing interval is determined as the next landing time of the user.

8. A headphone device, characterized in that: The headphone device includes: a memory, a processor, and a sound effect adjustment program stored in the memory and executable on the processor. When the sound effect adjustment program is executed by the processor, the steps of the sound effect adjustment method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a sound effect adjustment program, which, when executed by a processor, implements the steps of the sound effect adjustment method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Playing control method, device and system

    CN105959792A

  • Exercise apparatus

    JP2010005040A