An audio-based game hit mapping method and related apparatus

By combining the music beat information to correct the sound energy mutation points in the audio, the problem of traditional music game beat recognition methods being affected by musical instruments and human voices is solved, and the rhythm and user experience of the music game are improved.

CN114882902BActive Publication Date: 2025-10-17TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210468564.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-10-17
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

Traditional music game beat recognition methods are easily affected by factors such as instruments and vocals in the music, resulting in a more chaotic onset detected in the audio, affecting the rhythmicity of the interactive points.

Method used

The sound energy mutation points in the audio to be detected are corrected in combination with the beat information of the music, the beat information is identified through signal processing or deep learning neural network model, and the initial sound energy mutation points are corrected and adjusted according to the beat information, and points that do not conform to the rules are merged or deleted and mapped as interactive points in the music game.

Benefits of technology

The rhythm of the interactive points in the music game has been improved, making the interactive points more in line with the rules of the music and improving the user's gaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882902B_ABST
    Figure CN114882902B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a game beat mapping method based on audio and related devices, which are used to correct the Onset (sound energy mutation point) in the audio to be detected in combination with the beat information of a music piece, and map the corrected sound energy mutation point as an interactive point in a music game, so as to improve the rhythmicity of the interactive point in the music game. The method of the embodiments of the present application comprises: acquiring audio to be detected; identifying initial sound energy mutation points in the audio to be detected; acquiring beat information of the audio to be detected; correcting the initial sound energy mutation points in the audio to be detected in combination with the beat information of the audio to be detected, to obtain corrected sound energy mutation points; and mapping the corrected sound energy mutation points as interactive points in a music game.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of music data processing, and in particular to a game beat mapping method based on audio and related devices. BACKGROUND

[0002] Music game beat: refers to that a user completes corresponding tapping according to given tapping points according to song rhythm information. Scoring is performed according to the position of the tapping points hit by the user, and the higher the completion degree is, the higher the score is. Common music game beat design is to set tapping points manually according to song rhythm information.

[0003] The traditional game beat is only set based on the Onset (sound energy mutation point) point in the audio, but the traditional Onset recognition method is easily affected by factors such as musical instruments and vocals in the music, so that the Onset detected in the audio is relatively messy. SUMMARY

[0004] The embodiments of the present application provide a game beat mapping method based on audio and related devices, which is used for correcting the Onset (sound energy mutation point) in the audio to be detected in combination with the beat information of the music, and mapping the corrected sound energy mutation point as an interactive point in a music game, so as to improve the rhythm of the interactive point in the music game.

[0005] The first aspect of the embodiments of the present application provides a game beat mapping method based on audio, and the method comprises:

[0006] obtaining audio to be detected;

[0007] identifying an initial sound energy mutation point in the audio to be detected;

[0008] obtaining beat information of the audio to be detected;

[0009] combining the beat information of the audio to be detected, correcting the initial sound energy mutation point in the audio to be detected to obtain a corrected sound energy mutation point;

[0010] mapping the corrected sound energy mutation point as an interactive point in a music game.

[0011] Optionally, the beat information of the audio to be detected is obtained, comprising:

[0012] detecting beat point information and heavy beat information in the audio to be detected based on a signal processing or a deep learning neural network model;

[0013] obtaining distribution rules of the heavy beat information and the beat point information in the audio to be detected;

[0014] According to the distribution rule, beat type information of the audio to be detected is obtained.

[0015] Optionally, the beat information comprises the beat type information and the number of beats per unit time.

[0016] The beat information of the audio to be detected is combined to correct the sound energy mutation points in the audio to be detected to obtain corrected sound energy mutation points, comprising:

[0017] According to the beat type information of the audio to be detected and the number of beats per unit time, a minimum time interval of the initial sound energy mutation points is calculated.

[0018] If multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, a deletion operation and / or an adjustment operation are performed on the multiple initial sound energy mutation points to obtain corrected sound energy mutation points.

[0019] Optionally, if multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, a deletion operation is performed on the multiple initial sound energy mutation points, comprising:

[0020] If multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, at least two initial sound energy mutation points in the multiple initial sound energy mutation points are combined into one initial sound energy mutation point.

[0021] Optionally, after the at least two initial sound energy mutation points in the multiple initial sound energy mutation points are combined into one initial sound energy mutation point, the method further comprises:

[0022] According to the beat type information of the audio to be detected, a maximum number of sound energy mutation points allowed to appear in adjacent beat points of the audio to be detected is obtained.

[0023] If the number of initial sound energy mutation points appearing in the adjacent beat points of the audio to be detected is greater than the maximum number, other initial sound energy mutation points except the maximum number are deleted.

[0024] Optionally, if multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, an adjustment operation is performed on the multiple initial sound energy mutation points, comprising:

[0025] A time period between adjacent beat points in which the maximum number of initial sound energy mutation points is obtained.

[0026] The maximum number of initial sound energy mutation points is divided into each time point in the time period.

[0027] Optionally, the method further comprises:

[0028] identifying at least one of beat information, heavy beat information and long beat information in the audio to be detected;

[0029] mapping the at least one of beat information, heavy beat information and long beat information in the audio to be detected as an interaction point in the music game.

[0030] Optionally, the identifying the long beat information in the audio to be detected comprises:

[0031] obtaining a starting time and / or a duration of each note and / or each lyric in the audio to be detected;

[0032] determining a note and / or a lyric with a duration greater than a first preset duration as the long beat information in the audio to be detected.

[0033] Optionally, the identifying the beat information and / or the heavy beat information in the audio to be detected comprises:

[0034] detecting the beat information and / or the heavy beat information in the audio to be detected based on signal processing or a deep learning neural network model.

[0035] Optionally, before mapping the corrected sound energy mutation point in the audio to be detected and the at least one of beat information, heavy beat information and long beat information as an interaction point in the music game, the method further comprises:

[0036] sequentially marking the long beat information, the heavy beat information, the beat information and the corrected sound energy mutation point in the audio to be detected;

[0037] if the corrected sound energy mutation point overlaps with the heavy beat information and / or the beat information, marking the corrected sound energy mutation point as the heavy beat information and / or the beat information.

[0038] Optionally, the mapping the corrected sound energy mutation point in the audio to be detected and the at least one of beat information, heavy beat information and long beat information as an interaction point in the music game comprises:

[0039] mapping the corrected sound energy mutation point as a single click in the music game, and / or;

[0040] mapping the beat information as a single click in the music game; and / or;

[0041] mapping the heavy beat information as a double click in the music game; and / or;

[0042] Map the long tap information to a preset time length of continuous pressing in a music game.

[0043] Optionally, before identifying the initial sound energy mutation point in the audio to be detected, the method further comprises:

[0044] performing preprocessing on the audio to be detected to reduce the data processing amount of the audio to be detected, wherein the preprocessing comprises at least one of converting the audio to be detected into a single-channel speech signal and resampling a sampling rate of the audio to be detected to a standard sampling rate.

[0045] The second aspect of the embodiments of the present application provides a game beat mapping device based on audio, the device comprising:

[0046] an acquisition unit configured to acquire audio to be detected;

[0047] an identification unit configured to identify an initial sound energy mutation point in the audio to be detected;

[0048] The acquisition unit is further configured to acquire beat information of the audio to be detected.

[0049] a correction unit configured to correct the initial sound energy mutation point in the audio to be detected in combination with the beat information of the audio to be detected, to obtain a corrected sound energy mutation point.

[0050] a mapping unit configured to map the corrected sound energy mutation point to an interaction point in a music game.

[0051] Optionally, the acquisition unit is specifically configured to:

[0052] detect beat information and heavy tap information in the audio to be detected based on a signal processing or a deep learning neural network model;

[0053] obtain a distribution rule of the heavy tap information and the beat information in the audio to be detected;

[0054] obtain beat type information of the audio to be detected according to the distribution rule.

[0055] Optionally, the beat information comprises the beat type information and the number of beats per unit time.

[0056] The correction unit is specifically configured to:

[0057] calculate a minimum time interval of the initial sound energy mutation point according to the beat type information of the audio to be detected and the number of beats per unit time.

[0058] If multiple initial sound energy mutation points occur within a time interval not greater than half of the minimum time interval, the multiple initial sound energy mutation points are subjected to a pruning operation and / or an adjustment operation to obtain corrected sound energy mutation points.

[0059] Optionally, the correction unit is specifically configured to:

[0060] If multiple initial sound energy mutation points occur within a time interval not greater than half of the minimum time interval, at least two initial sound energy mutation points among the multiple initial sound energy mutation points are combined into one initial sound energy mutation point.

[0061] Optionally, the correction unit is further configured to:

[0062] According to the beat type information of the audio to be detected, a maximum number of sound energy mutation points allowed to occur within adjacent beat points of the audio to be detected is obtained.

[0063] If the number of initial sound energy mutation points occurring within the adjacent beat points of the audio to be detected is greater than the maximum number, other initial sound energy mutation points other than the maximum number are deleted.

[0064] Optionally, the correction unit is specifically configured to:

[0065] A time interval between adjacent beat points in which the maximum number of initial sound energy mutation points is obtained.

[0066] The maximum number of initial sound energy mutation points is evenly divided into each time point within the time interval.

[0067] Optionally, the device further comprises:

[0068] The identification unit is configured to identify at least one of beat information, double beat information and long beat information in the audio to be detected.

[0069] The mapping unit is further configured to:

[0070] Map the at least one of the beat information, the double beat information and the long beat information in the audio to be detected as an interaction point in the music game.

[0071] Optionally, the identification unit is specifically configured to:

[0072] Obtain a starting time and / or a duration of each note and / or each lyric in the audio to be detected.

[0073] Determine a note and / or a lyric with a duration greater than a first preset duration as long beat information in the audio to be detected.

[0074] Optionally, the identification unit is specifically configured to:

[0075] detecting the beat information and / or the double beat information in the to-be-detected audio based on a signal processing or a deep learning neural network model.

[0076] Optionally, the apparatus further comprises:

[0077] a marking unit, configured to sequentially mark the long beat information, the double beat information, the beat information and the corrected sound energy mutation point in the to-be-detected audio before mapping the corrected sound energy mutation point and at least one of the beat information, the double beat information and the long beat information in the to-be-detected audio into an interactive point in the music game;

[0078] The marking unit is further configured to mark the corrected sound energy mutation point as the double beat information and / or the beat information if the corrected sound energy mutation point overlaps with the double beat information and / or the beat information.

[0079] Optionally, the mapping unit is specifically configured to:

[0080] map the corrected sound energy mutation point into a single click in the music game, and / or;

[0081] map the beat information into a single click in the music game; and / or;

[0082] map the double beat information into a double click in the music game; and / or;

[0083] map the long beat information into a continuous press of a preset time length in the music game.

[0084] Optionally, the apparatus further comprises:

[0085] a preprocessing unit, configured to perform preprocessing on the to-be-detected audio before identifying the initial sound energy mutation point in the to-be-detected audio, so as to reduce the data processing amount of the to-be-detected audio, wherein the preprocessing comprises at least one of converting the to-be-detected audio into a monophonic speech signal and resampling a sampling rate of the to-be-detected audio to a standard sampling rate.

[0086] The embodiment of the present application further provides a computer apparatus comprising a processor and a memory, wherein the processor is configured to implement the beat mapping method based on audio game in the first aspect of the embodiment of the present application when executing a computer program stored in the memory.

[0087] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is configured to implement the beat mapping method based on audio game in the first aspect of the embodiment of the present application when executed by a processor.

[0088] From the above technical solutions, the embodiments of the present application have the following advantages:

[0089] In the embodiments of the present application, the audio to be detected is acquired; an initial sound energy mutation point in the audio to be detected is identified; beat information of the audio to be detected is acquired; the initial sound energy mutation point in the audio to be detected is corrected in combination with the beat information of the audio to be detected to obtain a corrected sound energy mutation point; and the corrected sound energy mutation point is mapped as an interactive point in a music game.

[0090] Because the embodiments of the present application can correct the initial sound energy mutation point in the audio to be detected based on the beat information of the audio to be detected, the corrected sound energy mutation point is more in line with the rules of the music, and the corrected sound energy mutation point is mapped as the interactive point in the music game, thereby improving the rhythmicity of the interactive point in the music game. BRIEF DESCRIPTION OF DRAWINGS

[0091] Figure 1 FIG. 1 is a schematic diagram of an embodiment of the audio-based game beat mapping method in the embodiments of the present application;

[0092] Figure 2 FIG. 2 is a schematic diagram of the audio-based game beat mapping method in the embodiments of the present application; Figure 1 FIG. 3 is a detailed step of the embodiment step 103;

[0093] Figure 3 FIG. 4 is a schematic diagram of the structure of a measure in a 4 / 4 measure and a 3 / 4 measure of music;

[0094] Figure 4 FIG. 5 is a schematic diagram of the audio-based game beat mapping method in the embodiments of the present application; Figure 1 FIG. 6 is a detailed step of the embodiment step 104;

[0095] Figure 5 FIG. 7 is a schematic diagram of a song energy envelope line before deleting the redundant initial sound energy mutation point in the embodiments of the present application;

[0096] Figure 6 FIG. 8 is a schematic diagram of a song energy envelope line after deleting the redundant initial sound energy mutation point in the embodiments of the present application;

[0097] Figure 7 FIG. 9 is a schematic diagram of another embodiment of the audio-based game beat mapping method in the embodiments of the present application;

[0098] Figure 8 FIG. 10 is a schematic diagram of an embodiment of the audio-based game beat mapping device in the embodiments of the present application. DETAILED DESCRIPTION

[0099] The embodiment of the present application provides a game beat mapping method based on audio and a related device, which is used for correcting an initial sound energy mutation point of to-be-detected audio, so that the corrected sound energy mutation point is more in line with the law of a music piece, and the corrected sound energy mutation point is mapped as an interactive point in a music game, so that the rhythm of the interactive point in the music game is improved.

[0100] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative efforts should belong to the protection scope of the present application.

[0101] The terms "first", "second", "third", "fourth" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0102] Based on the prior art, the detected Onset (sound energy mutation point) is easily affected by musical instruments, vocals and other factors in a music piece, so that the detected Onset in the audio is relatively chaotic. The present application provides a game beat mapping method based on audio and a related device, which is used for correcting a sound energy mutation point in to-be-detected audio in combination with beat information of the to-be-detected audio, and then mapping the corrected sound energy mutation point as an interactive point in a music game, so as to improve the rhythm of the interactive point in the music game.

[0103] For the convenience of understanding, the game beat mapping method based on audio in the embodiments of the present application will be described in detail below. Please refer to Figure 1 One embodiment of the game beat mapping method based on audio in the embodiments of the present application comprises the following steps.

[0104] 101、acquiring to-be-detected audio;

[0105] It is easy to understand that the present application needs to obtain the audio to be detected before mapping the onset (sound energy mutation point) in the audio to the interaction point in the music game, wherein the "obtaining" in the embodiments of the present application can be active obtaining or passive receiving.

[0106] Specifically, the execution subject of the method of the present application can be various computer devices (such as desktop computers, notebooks, tablets or wearable devices), and APPs or small programs installed in various computer devices.

[0107] Further, the onset of the audio, also known as the sound energy mutation point in the audio, refers to the starting point of the sudden appearance of other sounds in the process of playing the audio, and the other sounds can be any kind of sound other than the currently played sound, such as the sound of other musical instruments (violin, drum, piano) or the sound of other choristers.

[0108] 102, identifying the initial sound energy mutation point in the audio to be detected;

[0109] Specifically, the initial sound energy mutation point in the embodiments of the present application refers to the sound energy mutation point in the audio identified by a traditional method without modification, and there are many traditional methods for identifying the sound energy mutation in the audio, which can be based on existing audio processing tools such as Librosa tool kit for processing or deep neural network algorithm for identification, which is not limited here.

[0110] 103, obtaining the beat information of the audio to be detected;

[0111] In order to modify the initial sound energy mutation point in the audio to be detected, the beat information of the audio to be detected also needs to be obtained in the embodiments of the present application, wherein the beat information here includes at least one of beat point information, downbeat information, beat type information and the number of beats per unit time, which is not limited here.

[0112] Specifically, the beat point information of the audio includes the time when the beat appears in the audio, the downbeat information includes the time when the downbeat appears, and the beat type information is the number of beat points (beats) in a measure, such as specific beat type information which can include 4 / 4 beat, 3 / 4 beat, etc., i.e. there can be 4 beats or 3 beats in a measure, etc., and the number of beats per unit time is also called BPM, i.e. the number of beats per unit time.

[0113] 104, modifying the initial sound energy mutation point in the audio to be detected in combination with the beat information of the audio to be detected to obtain the modified sound energy mutation point;

[0114] In the embodiments of the present application, in order to make the sound energy mutation point in the to-be-detected audio more consistent with the rules of the music, the beat information of the to-be-detected audio is combined to correct the initial sound energy mutation point in the to-be-detected audio, so as to obtain a corrected sound energy mutation point.

[0115] The specific process of how to combine the beat information of the to-be-detected audio to correct the initial sound energy mutation point in the to-be-detected audio to obtain a corrected sound energy mutation point will be described in the following embodiments, and will not be described here.

[0116] 105. mapping the corrected sound energy mutation point as an interactive point in the music game.

[0117] In order to make the interactive point in the music game more consistent with the rules of the music, that is, more rhythmic, the embodiments of the present application map the corrected sound energy mutation point as an interactive point in the music game, so as to improve the rhythmicity of the interactive point in the music game.

[0118] In the embodiments of the present application, the to-be-detected audio is obtained; the initial sound energy mutation point in the to-be-detected audio is identified; the beat information of the to-be-detected audio is obtained; the beat information of the to-be-detected audio is combined to correct the initial sound energy mutation point in the to-be-detected audio, so as to obtain a corrected sound energy mutation point; and the corrected sound energy mutation point is mapped as an interactive point in the music game.

[0119] Because the embodiments of the present application can correct the initial sound energy mutation point of the to-be-detected audio based on the beat information of the to-be-detected audio, so that the corrected sound energy mutation point is more consistent with the rules of the music, and the corrected sound energy mutation point is mapped as an interactive point in the music game, thereby improving the rhythmicity of the interactive point in the music game.

[0120] Based on Figure 1 In the embodiments, when the beat information in step 102 is the beat type information of the to-be-detected audio, the process of obtaining the beat type information of the to-be-detected audio is described below, please refer to Figure 2 , Figure 2 For Figure 1 The detailed steps of step 103 in the embodiments are as follows:

[0121] 201. detecting the beat point information and the repeat information in the to-be-detected audio based on signal processing or a deep learning neural network model;

[0122] Specifically, there are many methods for identifying beat information (beat) and downbeat information in the audio to be detected, such as traditional signal processing or deep learning neural network model, and the specific deep learning neural network model can be madmom, an open source library based on python, for detecting beat information (beat) and downbeat information in the audio to be detected.

[0123] 202. Obtain the distribution rule of the downbeat information and the beat information in the audio to be detected.

[0124] In a song, the beat type of the song refers to the total length of the notes in each measure, and common beat types of songs include 1 / 4, 2 / 4, 3 / 4, 4 / 4, 3 / 8, 6 / 8, etc. According to different beat types of songs, the downbeat information and the beat information often follow certain distribution rules. For example, in a 4 / 4 measure, a complete measure is generally composed of a strong beat (downbeat), a weak beat, a secondary strong beat, and a weak beat, while in a 3 / 4 measure, a complete measure is generally composed of a strong beat (downbeat), a weak beat, and a secondary strong beat.

[0125] For the convenience of understanding, Figure 3 The structure diagrams of measures in 4 / 4 and 3 / 4 are given in FIGS. 1 and 2, and the distribution rules of the downbeat information and the beat information are also marked in the figures.

[0126] 203. Obtain the beat type information of the audio to be detected according to the distribution rule.

[0127] Specifically, as described in step 202, because in a 4 / 4 measure, a complete measure is generally composed of a strong beat (downbeat), a weak beat, a secondary strong beat, and a weak beat, while in a 3 / 4 measure, a complete measure is generally composed of a strong beat (downbeat), a weak beat, and a secondary strong beat, after identifying the distribution rule of the downbeat information and the beat information in the audio to be detected, the beat type information of the audio to be detected can be obtained according to the distribution rule of the downbeat information and the beat information in the audio to be detected.

[0128] The specific process for obtaining the beat type information of the audio to be detected is given in the embodiments of the present application, which improves the reliability of the process of identifying the beat type information of the audio to be detected.

[0129] Based on Figure 1 and Figure 2 the embodiments described above, the step 104 in the embodiments will be described in detail below. Please refer to Figure 1 , Figure 4 , Figure 4 for Figure 1 the detailed steps of step 104 in the embodiments:

[0130] 401、calculate the minimum time interval of the initial sound energy mutation point in the audio to be detected according to the beat type information of the audio to be detected and the number of beats per unit time;

[0131] Specifically, the beat information of the audio to be detected in the embodiment of the application includes beat type information of the audio to be detected and the number of beats per unit time. The number of beats per unit time of the audio to be detected can be detected by using the open source library madmom of python.

[0132] And in Figure 2 After obtaining the beat type information of the audio to be detected in the embodiment, the minimum time interval of the initial sound energy mutation point in the audio to be detected can be calculated according to the beat type information of the audio to be detected and the number of beats per unit time.

[0133] For the convenience of understanding, the following examples are given:

[0134] Suppose that the BPM (the number of beats per unit time 1s) of the audio to be detected is 87, and the beat type of the audio to be detected is 4 / 4 beat, then the minimum time interval of the initial sound energy mutation point is

[0135] It should be noted that the embodiment of the application is exemplified by 4 / 4 beat, and when the audio to be detected is 3 / 4 beat, the minimum time interval of the initial sound energy mutation point is That is, the minimum time interval of the initial sound energy mutation point in the embodiment of the application is related to the beat type information of the audio to be detected.

[0136] 402, determine whether multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, if yes, execute step 403, if no, execute step 404;

[0137] After obtaining the minimum time interval of the initial sound energy mutation point in the audio to be detected in step 401, it is further determined whether multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, wherein the multiple in the embodiment of the application is at least 2, and if yes, step 403 is executed, and if no, step 404 is executed.

[0138] 403, if multiple initial sound energy mutation points appear within a time not greater than half of the minimum time interval, a deletion operation and / or an adjustment operation are performed on the multiple initial sound energy mutation points to obtain a corrected sound energy mutation point;

[0139] If multiple initial sound energy mutation points occur within a time period not greater than half of the minimum time interval, a pruning operation and / or an adjustment operation is performed on the multiple initial sound energy mutation points to obtain corrected sound energy mutation points.

[0140] Specifically, performing a pruning operation on the multiple initial sound energy mutation points includes:

[0141] 1. If multiple initial sound energy mutation points occur within a time period not greater than half of the minimum time interval, at least two initial sound energy mutation points among the multiple initial sound energy mutation points are combined into one initial sound energy mutation point.

[0142] To maintain consistency of the description, the example of 4 / 4 beats in step 401 is continued: it is assumed that multiple (greater than or equal to 2) initial sound energy mutation points occur within the minimum time interval 0.1724s, and at least two initial sound energy mutation points among the multiple initial sound energy mutation points are combined into one initial sound energy mutation point.

[0143] That is, it is assumed that three initial sound energy mutation points occur within the minimum time interval 0.1724s, and at least two of the three initial sound energy mutation points are combined into one initial sound energy mutation point, so that at least one initial sound energy mutation point is maintained within the minimum time interval 0.1724s.

[0144] 2. According to the beat type information of the audio to be detected, the maximum number of initial sound energy mutation points allowed to occur within adjacent beat points of the audio to be detected is obtained; it is determined whether the number of initial sound energy mutation points occurring within the adjacent beat points of the audio to be detected is greater than the maximum number, if yes, the initial sound energy mutation points other than the maximum number are deleted, and if no, the number of initial sound energy mutation points occurring within the adjacent beat points is maintained, and the time period between the adjacent beat points in which the number of initial sound energy mutation points occurs is obtained, and the number of initial sound energy mutation points occurring within the adjacent beat points is divided into each time within the time period between the adjacent beat points.

[0145] After at least two initial sound energy mutation points among the multiple initial sound energy mutation points are combined into one initial sound energy mutation point, to further make the initial sound energy mutation points more consistent with the rules of music theory, the embodiment of the present application further obtains, according to the beat type information of the audio to be detected, the maximum number of initial sound energy mutation points allowed to occur within adjacent beat points of the audio to be detected, wherein the maximum number of initial sound energy mutation points allowed to occur within adjacent beat points of the audio to be detected is as follows:

[0146] If the to-be-detected audio is 4 / 4 beat, the maximum number of sound energy mutation points allowed to appear in the adjacent beat points of the to-be-detected audio is 4-1=3, and if the to-be-detected audio is 3 / 4 beat, the maximum number of sound energy mutation points allowed to appear in the adjacent beat points of the to-be-detected audio is 3-1=2.

[0147] If the number of initial sound energy mutation points appearing in the adjacent beat points of the to-be-detected audio is greater than the maximum number after at least two initial sound energy mutation points in the plurality of initial sound energy mutation points are combined into one initial sound energy mutation point, the initial sound energy mutation points other than the maximum number are deleted.

[0148] The following examples are given:

[0149] Suppose that 3 initial sound energy mutation points appear in the minimum time interval 0.2299s after at least two initial sound energy mutation points in the plurality of initial sound energy mutation points are combined into one initial sound energy mutation point, and the maximum number of sound energy mutation points allowed to appear in the adjacent beat points of the to-be-detected audio is 2, then the initial sound energy mutation points other than the maximum number (2) are deleted.

[0150] Specifically, when deleting the initial sound energy mutation points other than the maximum number (2), one initial sound energy mutation point can be randomly deleted, or the weakest one of the initial sound energy mutation points can be deleted, and the way of deletion is not limited here.

[0151] For the convenience of understanding, Figure 5 and Figure 6 The song energy envelope line diagrams before and after deleting the redundant initial sound energy mutation points are given respectively.

[0152] If the number of initial sound energy mutation points appearing in the adjacent beat points of the to-be-detected audio is not greater than the maximum number after at least two initial sound energy mutation points in the plurality of initial sound energy mutation points are combined into one initial sound energy mutation point, the number of initial sound energy mutation points appearing in the adjacent beat points is maintained, and the time period between the adjacent beat points in which the maximum number of initial sound energy mutation points are located is obtained. The number of initial sound energy mutation points in the adjacent beat points is divided into each time point in the time period between the adjacent beat points to improve the regularity of the distribution of the initial sound energy mutation points.

[0153] 3. Obtain the time period between the adjacent beat points in which the maximum number of initial sound energy mutation points are located; divide the maximum number of initial sound energy mutation points into each time point in the time period.

[0154] After deleting the initial sound energy mutation points other than the maximum number of initial sound energy mutation points, in order to further improve the regularity of the distribution of the sound energy mutation points, the embodiment of the present application further obtains the time period between the adjacent beat points where the maximum number of initial sound energy mutation points are located, and divides the maximum number of initial sound energy mutation points into each time in the time period.

[0155] The following will be described in combination with Figure 6 It is assumed that two (maximum number) sound energy mutation points appear in the beat points located at x1 moment and x2 moment, and the two sound energy mutation points (maximum number of sound energy mutation points) between x1 moment and x2 moment are divided into x1 moment and x2 moment, such as setting the first sound energy mutation point at x1+(x2-x1) / 3 moment and the second sound energy mutation point at x1+2(x2-x1) / 3 moment.

[0156] 404、If an initial sound energy mutation point appears within a time not greater than half of the minimum time interval, the single initial sound energy mutation point is kept, and the time period between the adjacent beat points where the single initial sound energy mutation point is located is obtained, and the single initial sound energy mutation point is divided into the time period between the adjacent beat points where the single initial sound energy mutation point is located to obtain the corrected sound energy mutation point.

[0157] If only one initial sound energy mutation point appears within a time not greater than half of the minimum time interval, the initial sound energy mutation point is not processed, that is, the initial sound energy mutation point is kept, and the single initial sound energy mutation point is divided into the time period between the adjacent beat points where the single initial sound energy mutation point is located.

[0158] The following is an example:

[0159] It is assumed that only one sound energy mutation point appears in the beat points located at x1 moment and x2 moment, and the single sound energy mutation point is directly set at (x2-x1) / 2 moment.

[0160] In the embodiment of the present application, the specific process of correcting the sound energy mutation points in the to-be-detected audio based on the beat information of the to-be-detected audio to obtain the corrected sound energy mutation points is described in detail, thereby improving the reliability of the process of obtaining and correcting the sound energy mutation points in the present application, and making the corrected sound energy mutation points more consistent with the rules of music theory.

[0161] The beat point mapping method based on audio in the embodiment of the present application will be described in detail below in combination with the above embodiment, please refer to Figure 7 , Figure 7For another embodiment of the audio-based game dot mapping method in the embodiments of the present application:

[0162] 701. Obtain audio to be detected;

[0163] It should be noted that steps 701 and 702 in the embodiments of the present application are similar to the descriptions of steps 101 and 102 in the embodiments of the present application, and will not be described here. Figure 1 The description of step 101 in the embodiments is similar, which will not be described here.

[0164] 702. Perform preprocessing on the audio to be detected to reduce the data processing amount of the audio to be detected, wherein the preprocessing includes at least one of converting the audio to be detected into a single-channel speech signal and resampling the sampling rate of the audio to be detected to a standard sampling rate;

[0165] After obtaining the audio to be detected, in order to reduce the data amount of audio operation and thus reduce the calculation pressure to improve the operation efficiency, the embodiments of the present application can also perform preprocessing on the audio to be detected to reduce the data processing amount of the audio to be detected, wherein the preprocessing includes at least one of converting the audio to be detected into a single-channel speech signal and resampling the sampling rate of the audio to be detected to a standard sampling rate.

[0166] Specifically, after obtaining the audio to be detected, if the audio to be detected is a non-single-channel signal, the audio to be detected can be converted into a single-channel speech signal. For example, if the audio to be detected is a double-channel signal, because the double-channel signal outputs the same waveform signal in the double channels, a single-channel signal can be obtained by averaging the signal energy of the two channels.

[0167] If the sampling rate of the audio to be detected is higher than the preset standard sampling rate (such as higher than 8 kHz), the audio signal can be resampled to the standard sampling rate (such as 8 kHz). Specifically, when resampling the audio signal, open source tools (libresample) or direct sequence extraction operations can be used to realize the resampling of the audio signal.

[0168] It should be noted that 8 kHz here is only an example of the preset standard sampling rate, and is not a specific limitation. The preset standard sampling rate is not specifically limited here, for example, the preset standard sampling rate can be 44100 Hz, or 48000 Hz, or 96000 Hz, etc.

[0169] 703. Identify an initial sound energy mutation point in the audio to be detected;

[0170] 704. Obtain beat information of the audio to be detected;

[0171] It should be noted that steps 703 and 704 in the embodiments of the present application are similar to the descriptions of steps 103 and 104 in the embodiments of the present application, and will not be described here. Figure 1The descriptions of 102 to 103 in the embodiment are similar and will not be repeated here.

[0172] 705. The beat information of the audio to be detected includes beat type information of the audio to be detected and the number of beats per unit time. The minimum time interval between occurrences of the initial sound energy mutation point is calculated based on the beat type information of the audio to be detected and the number of beats per unit time.

[0173] 706. Determine whether multiple initial sound energy mutation points appear within a time period not greater than half of the minimum time interval. If so, execute step 707; if not, execute step 708.

[0174] 707. If multiple initial sound energy mutation points appear within a time period no greater than half of the minimum time interval, performing a deletion operation and / or an adjustment operation on the multiple initial sound energy mutation points to obtain a revised sound energy mutation point;

[0175] 708. If an initial sound energy mutation point occurs within a time period no greater than half of the minimum time interval, retain the initial sound energy mutation point, obtain the time period between adjacent beats where the single initial sound energy mutation point is located, and evenly divide the single initial sound energy mutation point into the time period between adjacent beats where the single initial sound energy mutation point is located to obtain a corrected sound energy mutation point.

[0176] 709. Identify at least one of beat information, rebeat information, and long beat information in the audio to be detected;

[0177] Specifically, when identifying the beat information and rebeat information in the audio to be detected, the identification can be based on traditional signal processing methods or based on a deep learning neural network model, such as using the DBNDownBeatTrackingProcessor algorithm in the open source library madmom for identification.

[0178] Furthermore, when identifying long beat information in the audio to be detected, the identification can be performed based on the following methods:

[0179] For the input lyrics file or midi file (wherein the lyrics file or midi file will display the starting time and duration of each lyric or each note), the starting time and duration of each note and / or each lyric in the audio to be detected are obtained; the notes and / or lyrics whose duration is greater than the first preset duration are determined as long beat information in the audio to be detected.

[0180] For example, the last line of the song "lemom" is:

[0181] Now (238391, 284) also (239005, 303) a (239586, 294) ta (240295, 313) wa (240608, 318) ta (241630, 395) light (242025, 2663).

[0182] Wherein, the first time point after the lyrics is the start time of the lyrics (unit: ms), and the second time is the duration of the lyrics (unit: ms). For example, in now (238391, 284), 238391 is the start time point of "now", and 284 is the duration of "now".

[0183] Suppose the first preset duration in this song is 0.51724s, then in the last sentence of "lemom", only "light" is a long beat.

[0184] 710, mark the long beat information, the heavy beat information, the beat point information and the corrected sound energy mutation point in the to-be-detected audio in sequence;

[0185] After identifying the long beat information, the heavy beat information, the beat point information and the corrected sound energy mutation point in the to-be-detected audio, then mark the long beat information, the heavy beat information, the beat point information and the corrected sound energy mutation point in the to-be-detected audio in sequence, and execute step 712 in the marking process.

[0186] 711, judge whether the corrected sound energy mutation point overlaps with the heavy beat information and / or the beat point information, if yes, execute step 712, if not, execute step 713.

[0187] In order to enhance the rhythmicity of the to-be-detected audio, in the process of marking the long beat information, the heavy beat information, the beat point information and the corrected sound energy mutation point in the to-be-detected audio in sequence, judge whether the corrected sound energy mutation point overlaps with the heavy beat information and / or the beat point information, if yes, execute step 712, if not, execute step 713.

[0188] 712, if the corrected sound energy mutation point overlaps with the heavy beat information and / or the beat point information, mark the corrected sound energy mutation point as the heavy beat information and / or the beat point information.

[0189] If the corrected sound energy mutation point overlaps with the heavy beat information and / or the beat point information, in order to enhance the rhythmicity of the to-be-detected audio, mark the corrected sound energy mutation point as the heavy beat information and / or the beat point information.

[0190] 713、map the corrected sound energy mutation point in the audio to be detected and at least one of the beat information, the re-beat information, and the long-beat information to an interaction point in the music game.

[0191] As a possible implementation, in the mapping process, the beat information in the audio to be detected and the sound energy mutation point can be mapped to a single click in the music game, the re-beat information in the audio to be detected can be mapped to a double click in the music game, and the long-beat in the audio to be detected can be mapped to a continuous press of a preset length in the music game.

[0192] In addition, according to actual application scenarios, the beat information, the re-beat information, the long-beat information, and the sound energy mutation point in the audio to be detected can be mapped to an interaction point in the music game in other manners, such as mapping the beat information in the audio to be detected to a double click in the music game, mapping the re-beat information in the audio to be detected to a heavy click in the music game, mapping the long-beat information in the audio to be detected to N consecutive clicks in the music game, and mapping the sound energy mutation point in the audio to be detected to a single click in the music game. The specific manner of mapping the beat information, the re-beat information, the long-beat information, and the sound energy mutation point in the audio to be detected to an interaction point in the music game is not limited herein.

[0193] The preprocessing process of the audio to be detected is described in detail in the embodiments of the present application, which improves the operation efficiency of the audio to be processed. The process of obtaining the long-beat information, the re-beat information, and the beat information in the audio to be detected, and the process of mapping the long-beat information, the re-beat information, and the beat information in the audio to be detected and the corrected sound energy mutation point to an interaction point in the music game are described in detail, which respectively improves the reliability of each process.

[0194] The game beat mapping method based on audio in the present application is described in detail above. Next, the game beat mapping device based on audio in the embodiments of the present application is described. Please refer to Figure 8 One embodiment of the game beat mapping device based on audio in the embodiments of the present application includes:

[0195] The obtaining unit 801 is configured to obtain an audio to be detected.

[0196] The identifying unit 802 is configured to identify an initial sound energy mutation point in the audio to be detected.

[0197] The obtaining unit 801 is further configured to obtain beat information of the audio to be detected.

[0198] The correcting unit 803 is configured to correct the initial sound energy mutation point in the audio to be detected in combination with the beat information of the audio to be detected, to obtain a corrected sound energy mutation point.

[0199] map the corrected sound energy mutation point to an interaction point in the music game.

[0200] Optionally, the acquisition unit 801 is specifically configured to:

[0201] detect beat point information and heavy beat information in the to-be-detected audio based on signal processing or a deep learning neural network model;

[0202] acquire distribution rules of the heavy beat information and the beat point information in the to-be-detected audio;

[0203] acquire beat type information of the to-be-detected audio according to the distribution rules.

[0204] Optionally, the beat information includes the beat type information and the number of beats per unit time.

[0205] The correction unit 803 is specifically configured to:

[0206] calculate a minimum time interval of occurrence of the initial sound energy mutation point according to the beat type information of the to-be-detected audio and the number of beats per unit time.

[0207] If a plurality of initial sound energy mutation points occur within a time not greater than half of the minimum time interval, the plurality of initial sound energy mutation points are subjected to a deletion operation and / or an adjustment operation to obtain a corrected sound energy mutation point.

[0208] Optionally, the correction unit 803 is specifically configured to:

[0209] If a plurality of initial sound energy mutation points occur within a time not greater than half of the minimum time interval, at least two initial sound energy mutation points in the plurality of initial sound energy mutation points are combined into one initial sound energy mutation point.

[0210] Optionally, the correction unit 803 is further configured to:

[0211] acquire a maximum number of sound energy mutation points allowed to occur within adjacent beat points of the to-be-detected audio according to the beat type information of the to-be-detected audio.

[0212] If the number of initial sound energy mutation points occurring within the adjacent beat points of the to-be-detected audio is greater than the maximum number, other initial sound energy mutation points other than the maximum number are deleted.

[0213] Optionally, the correction unit 803 is specifically configured to:

[0214] acquire a time period between adjacent beat points in which the maximum number of initial sound energy mutation points are located.

[0215] evenly divide the maximum number of initial sound energy mutation points to each time point in the time period.

[0216] Optionally, the identification unit 802 is further configured to identify at least one of beat information, heavy beat information, and long beat information in the to-be-detected audio.

[0217] The mapping unit 804 is further configured to:

[0218] map the at least one of beat information, heavy beat information, and long beat information in the to-be-detected audio to an interaction point in the music game.

[0219] Optionally, the identification unit 802 is specifically configured to:

[0220] obtain a starting time and / or a duration of each note and / or each lyric in the to-be-detected audio;

[0221] determine a note and / or a lyric with a duration greater than a first preset duration as long beat information in the to-be-detected audio.

[0222] Optionally, the identification unit 802 is specifically configured to:

[0223] detect beat information and / or heavy beat information in the to-be-detected audio based on signal processing or a deep learning neural network model.

[0224] Optionally, the apparatus further includes:

[0225] a marking unit 805, configured to sequentially mark long beat information, heavy beat information, beat information, and the corrected sound energy mutation point in the to-be-detected audio before mapping the corrected sound energy mutation point in the to-be-detected audio and the at least one of beat information, heavy beat information, and long beat information to an interaction point in the music game.

[0226] The marking unit 805 is further configured to, if the corrected sound energy mutation point overlaps with the heavy beat information and / or the beat information, mark the corrected sound energy mutation point as the heavy beat information and / or the beat information.

[0227] Optionally, the mapping unit 804 is specifically configured to:

[0228] map the modified sound energy mutation point to a single click in the music game, and / or;

[0229] map the beat information to a single click in the music game; and / or;

[0230] map the heavy beat information to a double click in the music game; and / or;

[0231] mapping the long tap information as a continuous press of a preset time length in the music game.

[0232] Optionally, the apparatus further comprises:

[0233] a preprocessing unit 806, configured to perform preprocessing on the to-be-detected audio to reduce a data processing amount of the to-be-detected audio before identifying the initial sound energy mutation point in the to-be-detected audio, wherein the preprocessing comprises at least one of converting the to-be-detected audio into a single-channel speech signal and resampling a sampling rate of the to-be-detected audio to a standard sampling rate.

[0234] In the embodiments of the present application, the to-be-detected audio is acquired by the acquisition unit 801; the initial sound energy mutation point in the to-be-detected audio is identified by the identification unit 802; the beat information of the to-be-detected audio is acquired; the initial sound energy mutation point in the to-be-detected audio is corrected by the correction unit 803 in combination with the beat information of the to-be-detected audio to obtain a corrected sound energy mutation point; and the corrected sound energy mutation point is mapped as an interactive point in the music game by the mapping unit 804.

[0235] Because the embodiments of the present application can correct the initial sound energy mutation point of the to-be-detected audio based on the beat information of the to-be-detected audio, the corrected sound energy mutation point is more in line with the rules of the music, and the corrected sound energy mutation point is mapped as an interactive point in the music game, thereby improving the rhythmicity of the interactive point in the music game.

[0236] The above describes the audio-based game beat point calculation apparatus in the embodiments of the present application from the perspective of modularized functional entities, and the following describes the computer apparatus in the embodiments of the present application from the perspective of hardware processing:

[0237] The computer apparatus is used to implement the functions on the gateway device side, and one embodiment of the computer apparatus in the embodiments of the present application comprises:

[0238] a processor and a memory;

[0239] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the following steps can be implemented:

[0240] acquiring a to-be-detected audio;

[0241] identifying an initial sound energy mutation point in the to-be-detected audio;

[0242] acquiring beat information of the to-be-detected audio;

[0243] The initial sound energy mutation point in the audio to be detected is corrected based on the beat information of the audio to be detected, to obtain a corrected sound energy mutation point.

[0244] The corrected sound energy mutation point is mapped as an interaction point in the music game.

[0245] In some embodiments of the present application, the processor can be further configured to implement the following steps:

[0246] The beat point information and the heavy beat information in the audio to be detected are detected based on a signal processing or a deep learning neural network model.

[0247] The distribution rule of the heavy beat information and the beat point information in the audio to be detected is obtained.

[0248] The beat type information of the audio to be detected is obtained according to the distribution rule.

[0249] In some embodiments of the present application, the beat information includes the beat type information and the number of beats per unit time, and the processor can be further configured to implement the following steps:

[0250] The minimum time interval of the initial sound energy mutation point is calculated according to the beat type information of the audio to be detected and the number of beats per unit time.

[0251] If multiple initial sound energy mutation points appear within a time interval not greater than half of the minimum time interval, a deletion operation and / or an adjustment operation are performed on the multiple initial sound energy mutation points, to obtain a corrected sound energy mutation point.

[0252] In some embodiments of the present application, the processor can be further configured to implement the following steps:

[0253] If multiple initial sound energy mutation points appear within a time interval not greater than half of the minimum time interval, at least two initial sound energy mutation points in the multiple initial sound energy mutation points are combined into one initial sound energy mutation point.

[0254] In some embodiments of the present application, after the at least two sound energy mutation points in the multiple sound energy mutation points are combined into one sound energy mutation point, the processor can be further configured to implement the following steps:

[0255] The maximum number of sound energy mutation points allowed to appear in adjacent beat points of the audio to be detected is obtained according to the beat type information of the audio to be detected.

[0256] If the number of initial sound energy mutation points appearing in the adjacent beat points of the audio to be detected is greater than the maximum number, the initial sound energy mutation points other than the maximum number are deleted.

[0257] In some embodiments of the present application, the processor can be further specifically configured to implement the following steps:

[0258] acquire the time period between the adjacent beat points where the maximum number of initial sound energy mutation points are located;

[0259] divide the maximum number of initial sound energy mutation points into each time point in the time period.

[0260] In some embodiments of the present application, the processor can be further configured to implement the following steps:

[0261] identify at least one of the beat point information, the double beat information and the long beat information in the audio to be detected;

[0262] map the at least one of the beat point information, the double beat information and the long beat information in the audio to be detected into the interactive point in the music game.

[0263] In some embodiments of the present application, the processor can be further specifically configured to implement the following steps:

[0264] acquire the starting time and / or the duration of each note and / or each lyric in the audio to be detected;

[0265] determine the note and / or the lyric with a duration greater than the first preset duration as the long beat information in the audio to be detected.

[0266] In some embodiments of the present application, the processor can be further configured to implement the following steps:

[0267] detect the beat point information and / or the double beat information in the audio to be detected based on signal processing or a deep learning neural network model.

[0268] In some embodiments of the present application, before mapping the corrected sound energy mutation point in the audio to be detected and the at least one of the beat point information, the double beat information and the long beat information into the interactive point in the music game, the processor can be further configured to implement the following steps:

[0269] mark the long beat information, the double beat information, the beat point information and the corrected sound energy mutation point in the audio to be detected in sequence;

[0270] if the corrected sound energy mutation point overlaps with the double beat information and / or the beat point information, mark the corrected sound energy mutation point as the double beat information and / or the beat point information.

[0271] In some embodiments of the present application, the processor can be further configured to implement the following steps:

[0272] map the modified sound energy mutation point as a single click in the music game, and / or;

[0273] map the beat point information as a single click in the music game; and / or;

[0274] map the heavy beat information as a double click in the music game; and / or;

[0275] map the long beat information as a continuous press of a preset duration in the music game.

[0276] In some embodiments of the present application, before identifying the initial sound energy mutation point in the audio to be detected, the processor can be further configured to implement the following steps:

[0277] perform pre-processing on the audio to be detected to reduce the data processing amount of the audio to be detected, wherein the pre-processing includes at least one of converting the audio to be detected into a mono voice signal and resampling the sampling rate of the audio to be detected to a standard sampling rate.

[0278] It can be understood that the processor in the computer device described above can also implement the functions of each unit in the corresponding device embodiments described above when executing the computer program, and thus will not be described here. For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the audio-based game beat point calculation device. For example, the computer program can be divided into the units in the audio-based game beat point calculation device described above, and each unit can implement the specific functions as described above in the corresponding audio-based game beat point calculation device.

[0279] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device can include but is not limited to a processor and a memory. Those skilled in the art can understand that the processor and the memory are only examples of the computer device, and do not constitute a limitation on the computer device, and can include more or fewer components, or combine certain components, or different components, for example, the computer device can also include an input / output device, a network access device, a bus, and the like.

[0280] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor is a control center of the computer device, and connects various parts of the computer device through various interfaces and lines.

[0281] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required by a function, etc.; and the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.

[0282] The application further provides a computer readable storage medium for realizing the function of the audio-based game beat calculation device, and the computer readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the processor can be used to execute the following steps:

[0283] Obtaining to-be-detected audio;

[0284] Identifying an initial sound energy mutation point in the to-be-detected audio;

[0285] Obtaining beat information of the to-be-detected audio;

[0286] Combining the beat information of the to-be-detected audio, correcting the initial sound energy mutation point in the to-be-detected audio to obtain a corrected sound energy mutation point;

[0287] Mapping the corrected sound energy mutation point to an interactive point in a music game.

[0288] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0289] detecting beat point information and re-beat information in the to-be-detected audio based on signal processing or a deep learning neural network model;

[0290] obtaining distribution rules of the re-beat information and the beat point information in the to-be-detected audio;

[0291] obtaining beat type information of the to-be-detected audio according to the distribution rules.

[0292] In some embodiments of the present application, the beat information includes the beat type information and the number of beats per unit time, and when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be specifically used to implement the following steps:

[0293] calculating a minimum time interval of occurrence of the initial sound energy mutation point according to the beat type information of the to-be-detected audio and the number of beats per unit time;

[0294] If multiple initial sound energy mutation points occur within a time not greater than half of the minimum time interval, performing a deletion operation and / or an adjustment operation on the multiple initial sound energy mutation points to obtain a corrected sound energy mutation point.

[0295] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be specifically used to implement the following steps:

[0296] If multiple initial sound energy mutation points occur within a time not greater than half of the minimum time interval, at least two initial sound energy mutation points in the multiple initial sound energy mutation points are combined into one initial sound energy mutation point.

[0297] In some embodiments of the present application, after at least two sound energy mutation points in the multiple sound energy mutation points are combined into one sound energy mutation point, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0298] obtaining a maximum number of sound energy mutation points allowed to occur within adjacent beat points of the to-be-detected audio according to the beat type information of the to-be-detected audio;

[0299] If the number of initial sound energy mutation points occurring within the adjacent beat points of the to-be-detected audio is greater than the maximum number, deleting other initial sound energy mutation points other than the maximum number.

[0300] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be specifically used to implement the following steps:

[0301] obtaining a time period between adjacent beat points where the maximum number of initial sound energy mutation points are located;

[0302] dividing the maximum number of initial sound energy mutation points into each time point in the time period.

[0303] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0304] identifying at least one of beat point information, heavy beat information and long beat information in the audio to be detected;

[0305] mapping the at least one of beat point information, heavy beat information and long beat information in the audio to be detected into an interactive point in the music game.

[0306] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be specifically used to implement the following steps:

[0307] obtaining a starting time and / or a duration of each note and / or each lyric in the audio to be detected;

[0308] determining a note and / or a lyric with a duration greater than a first preset duration as long beat information in the audio to be detected.

[0309] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0310] detecting beat point information and / or heavy beat information in the audio to be detected based on signal processing or a deep learning neural network model.

[0311] In some embodiments of the present application, before mapping the corrected sound energy mutation points in the audio to be detected and at least one of the beat point information, the heavy beat information and the long beat information into the interactive point in the music game, the computer program stored in the computer readable storage medium is executed by the processor, and the processor can also be used to implement the following steps:

[0312] sequentially marking long beat information, heavy beat information, beat point information and corrected sound energy mutation points in the audio to be detected;

[0313] If the modified sound energy mutation point overlaps with the re-hit information and / or the beat point information, the modified sound energy mutation point is marked as the re-hit information and / or the beat point information.

[0314] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0315] mapping the modified sound energy mutation point as a single click in the music game, and / or;

[0316] mapping the beat point information as a single click in the music game; and / or;

[0317] mapping the re-hit information as a double click in the music game; and / or;

[0318] mapping the long beat information as a continuous press of a preset length in the music game.

[0319] In some embodiments of the present application, before identifying the initial sound energy mutation point in the audio to be detected, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0320] performing preprocessing on the audio to be detected to reduce the data processing amount of the audio to be detected, wherein the preprocessing includes at least one of converting the audio to be detected into a single-channel speech signal and resampling the sampling rate of the audio to be detected to a standard sampling rate.

[0321] It can be understood that the integrated units, if implemented in the form of software function units and sold or used as independent products, can be stored in a corresponding computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned corresponding embodiment methods can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can include any entity or device, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code. It should be noted that the contents included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0322] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0323] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0324] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0325] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0326] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for mapping game beats based on audio, characterized in that: The method comprises: Get the audio to be detected; Identifying an initial sound energy mutation point in the audio to be detected; Obtaining beat information of the audio to be detected, where the beat information includes beat type information and the number of beats per unit time; Calculate the minimum time interval of the initial sound energy mutation point according to the beat type information of the audio to be detected and the number of beats per unit time; Based on the minimum time interval and the number of the initial sound mutation points occurring within the minimum time interval, the initial sound energy mutation point in the audio to be detected is corrected to obtain a corrected sound energy mutation point; The modified sound energy mutation points are mapped as interaction points in the music game.

2. The method according to claim 1, characterized in that The obtaining of the beat information of the audio to be detected includes: Detecting beat information and rebeat information in the audio to be detected based on signal processing or a deep learning neural network model; Obtaining a distribution pattern of the rebeat information and the beat point information in the audio to be detected; According to the distribution rule, the beat type information of the audio to be detected is obtained.

3. The method according to claim 2, characterized in that The step of correcting the initial sound energy mutation point in the audio to be detected based on the minimum time interval and the number of the initial sound mutation points occurring within the minimum time interval to obtain a corrected sound energy mutation point includes: If multiple initial sound energy mutation points appear within a time period not greater than half of the minimum time interval, deletion operations and / or adjustment operations are performed on the multiple initial sound energy mutation points to obtain modified sound energy mutation points.

4. The method according to claim 3, characterized in that If multiple initial sound energy mutation points appear within a time period not greater than half of the minimum time interval, performing a deletion operation on the multiple initial sound energy mutation points, including: If multiple initial sound energy mutation points appear within a time period not greater than half of the minimum time interval, at least two initial sound energy mutation points among the multiple initial sound energy mutation points are merged into one initial sound energy mutation point.

5. The method according to claim 4, characterized in that After merging at least two of the multiple initial sound energy mutation points into one initial sound energy mutation point, the method further includes: According to the beat type information of the audio to be detected, obtaining the maximum number of sound energy mutation points allowed to appear between adjacent beat points of the audio to be detected; If the number of initial sound energy mutation points appearing within the adjacent beat points of the audio to be detected is greater than the maximum number, the other initial sound energy mutation points beyond the maximum number are deleted.

6. The method according to claim 5, characterized in that If multiple initial sound energy mutation points appear within a time period not greater than half of the minimum time interval, performing an adjustment operation on the multiple initial sound energy mutation points includes: Obtaining the time period between adjacent beat points where the maximum number of initial sound energy mutation points are located; The maximum number of initial sound energy mutation points is evenly distributed to each moment in the time period.

7. The method according to claim 1, characterized in that The method further comprises: Identifying at least one of beat information, rebeat information, and long beat information in the audio to be detected; At least one of the beat information, the rebeat information, and the long beat information in the audio to be detected is mapped to an interaction point in a music game.

8. The method according to claim 7, characterized in that The identifying long beat information in the audio to be detected includes: Obtaining the starting time and / or duration of each note and / or each lyric in the audio to be detected; The notes and / or lyrics that last longer than the first preset duration are determined as long beat information in the audio to be detected.

9. The method according to claim 7, characterized in that The identifying the beat information and / or rebeat information in the audio to be detected includes: The beat information and / or rebeat information in the audio to be detected is detected based on signal processing or a deep learning neural network model.

10. The method according to claim 7, characterized in that Before mapping the corrected sound energy mutation point in the audio to be detected and at least one of the beat information, the rebeat information, and the long beat information as an interaction point in the music game, the method further includes: Sequentially mark the long beat information, re-beat information, beat point information and corrected sound energy mutation points in the audio to be detected; If the corrected sound energy mutation point overlaps with the rebeat information and / or the beat point information, the corrected sound energy mutation point is marked as the rebeat information and / or the beat point information.

11. The method according to claim 7, characterized in that Mapping the corrected sound energy mutation point in the audio to be detected and at least one of the beat information, the rebeat information, and the long beat information into an interaction point in the music game includes: Mapping the modified sound energy mutation point to a single click in a music game, and / or; Mapping the beat information into single clicks in a music game; and / or; mapping the re-beat information into a double-beat in a music game; and / or; The long beat information is mapped to continuous pressing of a preset duration in the music game.

12. The method according to claims 1 to 11, characterized in that Before identifying the initial sound energy mutation point in the audio to be detected, the method further includes: Preprocessing is performed on the audio to be detected to reduce the amount of data processing for the audio to be detected, wherein the preprocessing includes at least one of converting the audio to be detected into a monophonic speech signal and resampling the sampling rate of the audio to be detected to a standard sampling rate.

13. A computer device comprising a processor and a memory, characterized in that: When executing the computer program stored in the memory, the processor is configured to implement the audio-based game beat mapping method according to any one of claims 1 to 12.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the audio-based game beat mapping method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Audio processing method and device and storage medium

    CN109817241A

  • Audio popping detection method and device and storage medium

    CN110265064A