Audio synthesis method, device, equipment and computer-readable storage medium

By obtaining and matching sub-audio with hearing impaired hearing tone, the problem that hearing impaired patients cannot hear high-frequency components of music is solved, high-quality music synthesized audio is achieved, and the music listening experience of hearing impaired patients is improved.

CN113936628BActive Publication Date: 2025-06-13TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111189249.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2025-06-13
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

When listening to music, hearing impaired patients are not sensitive to the high-frequency components of the sound, and cannot hear the high-frequency components of the music, resulting in the intermittent music and poor sound quality.

Method used

By obtaining the music score data of the target music, matching the sub-audio with hearing impaired hearing tone, and fusion processing is performed based on the performance time information of these sub-audios, synthetic audio that can be heard by hearing impaired patients is generated.

Benefits of technology

It solves the problems of intermittent and poor sound quality when listening to music by hearing impaired patients, providing a smooth and good sound quality music experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113936628B_ABST
    Figure CN113936628B_ABST
Patent Text Reader

Abstract

The present application discloses an audio synthesis method, apparatus, device and computer-readable storage medium, belonging to the field of computer technology. The method includes: obtaining score data of target music, wherein the score data includes audio data identifiers corresponding to a plurality of sub-audios and performance time information, and the musical instrument timbre corresponding to each sub-audio matches the hearing-impaired hearing timbre; obtaining the corresponding sub-audio based on each audio data identifier; and performing fusion processing on each sub-audio based on the performance time information corresponding to each sub-audio to generate a synthesized audio of the target music. The synthesized audio obtained based on this method can be completely heard by hearing-impaired patients without distortion, enabling hearing-impaired patients to hear smooth music, having a better listening experience for hearing-impaired patients, being able to improve the sound quality when hearing-impaired patients listen to music, and improving the listening effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and particularly to an audio synthesis method, apparatus, device, and computer-readable storage medium. Background Art

[0002] With the continuous enrichment of audio resources (such as music), people can listen to the music they want at any time and place. However, due to the insufficient sensitivity to the high-frequency components of sound, hearing-impaired patients are prone to the problem of not being able to hear when listening to audio. Therefore, there is an urgent need for an audio synthesis method to synthesize audio that can be heard by hearing-impaired patients.

[0003] In the related art, taking an audio resource as music as an example. When a hearing-impaired patient listens to music, if not wearing a hearing aid, they can only hear the sound of the low-frequency components in the music and cannot hear the sound of the high-frequency components in the music, making the music heard by the hearing-impaired patient intermittent and not smooth enough. Furthermore, it leads to relatively distorted music and poor sound quality heard by the hearing-impaired patient, resulting in a poor music listening effect for the hearing-impaired patient. Summary of the Invention

[0004] Embodiments of the present application provide an audio synthesis method, apparatus, device, and computer-readable storage medium, which can be used to solve the problems in the related art. The technical solutions are as follows:

[0005] On the one hand, embodiments of the present application provide an audio synthesis method, and the method includes:

[0006] Obtain the score data of the target music, where the score data includes audio data identifiers corresponding to multiple sub-audios and performance time information, and the instrument timbre corresponding to each sub-audio matches the hearing-impaired hearing timbre;

[0007] Obtain the corresponding sub-audio based on each audio data identifier;

[0008] Based on the performance time information corresponding to each sub-audio, perform a fusion process on each sub-audio to generate the synthesized audio of the target music.

[0009] Optionally, in the spectrum of the instrument corresponding to each sub-audio, the ratio of the energy of the low-frequency band to the energy of the high-frequency band is greater than a ratio threshold, the low-frequency band is the band below the frequency threshold, and the high-frequency band is the band above the frequency threshold, where the ratio threshold is used to indicate the condition that the ratio of the energy of the low-frequency band to the energy of the high-frequency band in the spectrum of the audio that can be heard by hearing-impaired patients needs to meet.

[0010] Optionally, the obtaining the score data of the target music includes:

[0011] Determine the audio data identifier and performance time information corresponding to the multiple sub-audios based on the tempo, time signature, and chord list of the target music.

[0012] Optionally, the multiple sub-audios include a drumbeat sub-audio and a chord sub-audio;

[0013] The determining the audio data identifier and performance time information corresponding to the multiple sub-audios based on the tempo, time signature, and chord list of the target music includes:

[0014] Determine the audio data identifier and performance time information corresponding to the drumbeat sub-audio based on the tempo and time signature of the target music;

[0015] Determine the audio data identifier and performance time information corresponding to the chord sub-audio based on the tempo, time signature, and chord list of the target music;

[0016] The audio data identifier and performance time information corresponding to the drumbeat sub-audio, and the audio data identifier and performance time information corresponding to the chord sub-audio, constitute the audio data identifier and performance time information corresponding to the multiple sub-audios.

[0017] Optionally, the determining the audio data identifier and performance time information corresponding to the drumbeat sub-audio based on the tempo and time signature of the target music includes:

[0018] Determine the audio data identifier corresponding to the time signature and tempo of the target music, and use the audio data identifier corresponding to the time signature and tempo of the target music as the audio data identifier corresponding to the drumbeat sub-audio;

[0019] Based on the time signature and tempo of the target music, determine the performance time information corresponding to the drumbeat sub-audio.

[0020] Optionally, the chord list includes a chord identifier and the performance time information corresponding to the chord identifier;

[0021] The determining the audio data identifier and performance time information corresponding to the chord sub-audio based on the tempo, time signature, and chord list of the target music includes:

[0022] Based on the tempo and time signature of the target music, determine the audio data identifier corresponding to the chord identifier;

[0023] Determine the performance time information and audio data identifier corresponding to the chord sub-audio as the performance time information and audio data identifier corresponding to the chord identifier.

[0024] Optionally, the performing a fusion process on each sub-audio based on the performance time information corresponding to each sub-audio to generate a synthesized audio of the target music includes:

[0025] Based on the performance time information corresponding to each sub-audio, perform fusion processing on each sub-audio to obtain an intermediate audio of the target music;

[0026] Perform frequency domain compression processing on the intermediate audio of the target music to obtain a synthesized audio of the target music.

[0027] Optionally, the performing frequency domain compression processing on the intermediate audio of the target music to obtain the synthesized audio of the target music includes:

[0028] Obtain a first sub-audio in a first frequency range and a second sub-audio in a second frequency range corresponding to the intermediate audio, where the frequency of the first frequency range is less than the frequency of the second frequency range;

[0029] Based on a first gain coefficient, perform gain compensation on the first sub-audio to obtain a third sub-audio, and based on a second gain coefficient, perform gain compensation on the second sub-audio to obtain a fourth sub-audio;

[0030] Perform compression frequency shift processing on the fourth sub-audio to obtain a fifth sub-audio, where the lower limit of a third frequency range corresponding to the fifth sub-audio is equal to the lower limit of the second frequency range;

[0031] Perform fusion processing on the third sub-audio and the fifth sub-audio to obtain the synthesized audio of the target music.

[0032] Optionally, the performing compression frequency shift processing on the fourth sub-audio to obtain the fifth sub-audio includes:

[0033] Perform frequency compression of a target ratio on the fourth sub-audio to obtain a sixth sub-audio;

[0034] Perform frequency upward shift of a target value on the sixth sub-audio to obtain the fifth sub-audio, where the target value is equal to the difference between the lower limit of the second frequency range and the lower limit of a fourth frequency range corresponding to the sixth sub-audio.

[0035] On the other hand, an embodiment of the present application provides an audio synthesis device, and the device includes:

[0036] An acquisition module, configured to acquire score data of a target music, where the score data includes audio data identifiers and performance time information corresponding to a plurality of sub-audios, and the musical instrument timbre corresponding to each sub-audio matches the hearing-impaired hearing timbre;

[0037] The acquisition module is configured to acquire corresponding sub-audios based on each audio data identifier;

[0038] A generation module, configured to perform a fusion process on each of the sub-audios based on the performance time information corresponding to each of the sub-audios, so as to generate a synthesized audio of the target music.

[0039] Optionally, in the spectrum of the instrument corresponding to each of the sub-audios, the ratio of the energy of the low-frequency band to the energy of the high-frequency band is greater than a ratio threshold, where the low-frequency band is a band lower than a frequency threshold, and the high-frequency band is a band higher than the frequency threshold, and the ratio threshold is used to indicate a condition that needs to be satisfied by the ratio of the energy of the low-frequency band to the energy of the high-frequency band in the spectrum of an audio that can be heard by a hearing-impaired patient.

[0040] Optionally, the acquisition module is configured to determine the audio data identifier and the performance time information corresponding to the plurality of sub-audios based on the tempo, time signature, and chord list of the target music.

[0041] Optionally, the plurality of sub-audios include a drumbeat sub-audio and a chord sub-audio;

[0042] The acquisition module is configured to determine the audio data identifier and the performance time information corresponding to the drumbeat sub-audio based on the tempo and time signature of the target music;

[0043] Based on the tempo, time signature, and chord list of the target music, determine the audio data identifier and the performance time information corresponding to the chord sub-audio;

[0044] The audio data identifier and the performance time information corresponding to the drumbeat sub-audio, and the audio data identifier and the performance time information corresponding to the chord sub-audio together form the audio data identifier and the performance time information corresponding to the plurality of sub-audios.

[0045] Optionally, the acquisition module is configured to determine the audio data identifier corresponding to the time signature and tempo of the target music, and use the audio data identifier corresponding to the time signature and tempo of the target music as the audio data identifier corresponding to the drumbeat sub-audio;

[0046] Based on the time signature and tempo of the target music, determine the performance time information corresponding to the drumbeat sub-audio.

[0047] Optionally, the chord list includes a chord identifier and the performance time information corresponding to the chord identifier;

[0048] The acquisition module is configured to determine the audio data identifier corresponding to the chord identifier based on the tempo and time signature of the target music;

[0049] Determine the performance time information and the audio data identifier corresponding to the chord sub-audio as the performance time information and the audio data identifier corresponding to the chord sub-audio.

[0050] Optionally, the generating module is configured to perform a fusion process on each sub-audio based on the playing time information corresponding to each sub-audio to obtain an intermediate audio of the target music;

[0051] Perform a frequency-domain compression process on the intermediate audio of the target music to obtain a synthesized audio of the target music.

[0052] Optionally, the synthesizing module is configured to obtain a first sub-audio in a first frequency range and a second sub-audio in a second frequency range corresponding to the intermediate audio, where the frequency of the first frequency range is less than the frequency of the second frequency range;

[0053] Perform gain compensation on the first sub-audio based on a first gain coefficient to obtain a third sub-audio, and perform gain compensation on the second sub-audio based on a second gain coefficient to obtain a fourth sub-audio;

[0054] Perform a compression and frequency shift process on the fourth sub-audio to obtain a fifth sub-audio, where the lower limit of a third frequency range corresponding to the fifth sub-audio is equal to the lower limit of the second frequency range;

[0055] Perform a fusion process on the third sub-audio and the fifth sub-audio to obtain a synthesized audio of the target music.

[0056] Optionally, the generating module is configured to perform frequency compression on the fourth sub-audio at a target ratio to obtain a sixth sub-audio;

[0057] Perform a frequency upward shift on the sixth sub-audio by a target value to obtain the fifth sub-audio, where the target value is equal to the difference between the lower limit of the second frequency range and the lower limit of a fourth frequency range corresponding to the sixth sub-audio.

[0058] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory. At least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to enable the computer device to implement any one of the above audio synthesis methods.

[0059] On the other hand, a computer-readable storage medium is further provided. At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to enable a computer to implement any one of the above audio synthesis methods.

[0060] On the other hand, a computer program or a computer program product is further provided. At least one computer instruction is stored in the computer program or the computer program product, and the at least one computer instruction is loaded and executed by a processor to enable a computer to implement any one of the above audio synthesis methods.

[0061] The technical solution provided by the embodiment of the present application at least brings the following beneficial effects:

[0062] The technical solution provided by the embodiment of the present application re-composes the target music. When composing the music, the instrument timbre of the sub-audio used matches the hearing-impaired hearing timbre, so that the hearing-impaired patients can hear the sub-audio used in the composition. Furthermore, based on the sub-audio, the synthesized audio of the target music is obtained. When the hearing-impaired patients listen to the synthesized audio of the target music, there will be no problem of intermittent or occasional inaudibility, and there will also be no distortion. The hearing-impaired patients can hear the complete music, and the listening experience of the hearing-impaired patients is better, which can fundamentally solve the problems of poor sound quality and poor listening effect when the hearing-impaired patients listen to music. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0064] Figure 1 It is a schematic diagram of the implementation environment of an audio synthesis method provided by the embodiment of the present application;

[0065] Figure 2 It is a flowchart of an audio synthesis method provided by the embodiment of the present application;

[0066] Figure 3 It is a simple score diagram of the fourth, fifth, and sixth musical measures of the song "Paradise" provided by the embodiment of the present application;

[0067] Figure 4 It is a simple score diagram corresponding to the synthesized audio of the fourth, fifth, and sixth musical measures of the song "Paradise" provided by the embodiment of the present application;

[0068] Figure 5 It is a flowchart of an audio synthesis method provided by the embodiment of the present application;

[0069] Figure 6 It is a schematic structural diagram of an audio synthesis device provided by the embodiment of the present application;

[0070] Figure 7 It is a schematic structural diagram of a terminal device provided by the embodiment of the present application;

[0071] Figure 8 It is a schematic structural diagram of a server provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0073] The following will introduce in detail the terms involved in the embodiments of this application.

[0074] WDRC (Wide Dynamic Range Compressor), a dynamic range control algorithm, characterized by a low compression ratio / low compression threshold and supporting dynamic adjustment of compression metrics.

[0075] Cross-Fade: The overlapping parts at the beginning and end of two audio segments are spliced into a complete audio segment after being faded in and out alternately.

[0076] Nonlinear compression frequency shift: A method of compressing the high-frequency components of hearing-impaired patients and then shifting them to the low-frequency region of the residual hearing of hearing-impaired patients.

[0077] Figure 1 It is a schematic diagram of the implementation environment of an audio synthesis method provided by an embodiment of this application. As Figure 1 shown, the implementation environment includes: a computer device 101. The audio synthesis method provided by the embodiments of this application can be executed by the computer device 101. Exemplarily, the computer device 101 can be a terminal device or a server, and this application does not limit this.

[0078] The terminal device can be at least one of a smart phone, a game console, a desktop computer, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop computer.

[0079] The server can be a single server, a server cluster composed of multiple servers, or any one of a cloud computing platform and a virtualization center, and this application does not limit this. The server is communicatively connected to the terminal device through a wired network or a wireless network. The server can have the functions of data transceiver, data processing, and data storage. Of course, the server can also have other functions, and this application does not limit this.

[0080] Based on the above implementation environment, the embodiments of this application provide an audio synthesis method to Figure 2Taking the flowchart of an audio synthesis method provided by an embodiment of the present application as an example, this method can be executed by the Figure 1 computer device 101 in Figure 2 . As shown in

[0081] , this method includes the following steps:

[0082] In step 201, musical score data of the target music is obtained, where the musical score data includes audio data identifiers and performance time information of multiple sub-audios, and the musical instrument timbre corresponding to each sub-audio matches the hearing-impaired hearing timbre.

[0083] Optionally, in the spectrum of the musical instrument corresponding to each sub-audio, the ratio of the energy in the low-frequency band to the energy in the high-frequency band is greater than a ratio threshold. The low-frequency band is the band below the frequency threshold, and the high-frequency band is the band above the frequency threshold, where the ratio threshold is used to indicate the condition that the ratio of the energy in the low-frequency band to the energy in the high-frequency band in the spectrum of the audio that can be heard by hearing-impaired patients needs to meet.

[0084] Among them, the frequency threshold can be obtained based on experiments, and the embodiments of the present application do not limit this. For example, the frequency threshold is 2 kHz. The ratio threshold is the minimum value of the ratio of the energy in the low-frequency band to the energy in the high-frequency band in the spectrum of the audio that can be heard by hearing-impaired patients.

[0085] Exemplarily, there are multiple audios stored in the computer device, and the ratio of the energy in the low-frequency band to the energy in the high-frequency band corresponding to each audio is different, and the ratio of the energy in the low-frequency band to the energy in the high-frequency band corresponding to each audio differs by a certain value, for example, 2%. They are played in order from high to low according to the ratio of the energy in the low-frequency band to the energy in the high-frequency band for the hearing-impaired patients to listen to. In response to the hearing-impaired patients being able to hear the audio with the ratio of the energy in the low-frequency band to the energy in the high-frequency band being 50%, but the hearing-impaired patients being unable to hear the audio with the ratio of the energy in the low-frequency band to the energy in the high-frequency band being 48%, therefore, the ratio threshold is set to 50%.

[0086] Generally speaking, the frequency range of the sound that normal people can hear is roughly within 20,000 Hz, and the frequency range that hearing-impaired patients can hear is roughly within 8,000 Hz. The sounding frequency of the musical instrument corresponding to the sub-audio used in the embodiments of the present application is mainly within 8,000 Hz, which is designed for hearing-impaired patients. For hearing-impaired patients, they can hear more clearly, so the synthesized audio synthesized from these sub-audios can also be better listened to by hearing-impaired patients.

[0087] Optionally, the process of determining which musical instrument timbres match the hearing-impaired hearing timbre is as follows: Obtain the sounds corresponding to each musical instrument, play the sounds corresponding to each musical instrument so that the hearing-impaired patient can listen. Based on the feedback information of the hearing-impaired patient, determine which musical instrument timbres match the hearing-impaired hearing timbre.

[0088] If the feedback information indicates that the hearing-impaired patient can hear a certain sound, determine that the musical instrument timbre of the musical instrument corresponding to the sound that the hearing-impaired patient can hear matches the hearing-impaired hearing timbre. If the feedback information indicates that the hearing-impaired patient cannot hear a certain sound, determine that the musical instrument timbre of the musical instrument corresponding to the sound that the hearing-impaired patient cannot hear does not match the hearing-impaired hearing timbre.

[0089] Exemplarily, obtain Sound One, Sound Two, and Sound Three, where Sound One is the sound corresponding to a piano, Sound Two is the sound corresponding to a bass, and Sound Three is the sound corresponding to a snare drum. Play these three sounds separately so that the hearing-impaired patient can listen to these three sounds respectively. If the hearing-impaired patient can hear Sound Two and Sound Three but cannot hear Sound One, determine that the timbres of the bass and the snare drum match the hearing-impaired hearing timbre, while the timbre of the piano does not match the hearing-impaired hearing timbre.

[0090] It should be noted that the sounds corresponding to all musical instruments can be obtained for the hearing-impaired patient to listen to, and then the musical instrument timbres that match the hearing-impaired hearing timbre can be determined. In this embodiment of the application, only the above two musical instrument timbres that match the hearing-impaired hearing timbre are taken as examples for illustration. The musical instrument timbres that match the hearing-impaired hearing timbre can be more or less, and this embodiment of the application does not limit this.

[0091] Optionally, the sub-audio corresponding to the audio data identifier and the performance time information included in the score data of the target music can be a drumbeat sub-audio, or a chord sub-audio, or both a drumbeat sub-audio and a chord sub-audio. This embodiment of the application does not limit this. Since when the sub-audio corresponding to the audio data identifier and the performance time information included in the score data is only a drumbeat sub-audio or only a chord sub-audio, the synthesized audio of the target music obtained according to the score data, although the hearing-impaired patient can hear it, is relatively dull and single. Therefore, this embodiment of the application takes the sub-audio corresponding to the audio data identifier and the performance time information included in the score data as both a drumbeat sub-audio and a chord sub-audio as an example for illustration. The score data includes the audio data identifier and the performance time information corresponding to the drumbeat sub-audio, and the audio data identifier and the performance time information corresponding to the chord sub-audio.

[0092] It should be noted that when the sub-audio corresponding to the audio data identifier and the performance time information included in the score data of the target music is a drumbeat sub-audio or a chord sub-audio, the process of obtaining the synthesized audio of the target music is similar to the process of obtaining the synthesized audio of the target music when the sub-audio corresponding to the audio data identifier and the performance time information included in the score data of the target music is a drumbeat sub-audio and a chord sub-audio.

[0093] In a possible implementation manner, the process of obtaining the score data of the target music may be: based on the tempo, time signature, and chord list of the target music, determine the audio data identifiers and performance time information corresponding to multiple sub-audios.

[0094] Among them, before determining the audio data identifiers and performance time information corresponding to multiple sub-audios based on the tempo, time signature, and chord list of the target music, it is also necessary to first determine the tempo, time signature, and chord list of the target music. The methods for determining the tempo, time signature, and chord list of the target music include but are not limited to the following three: The first one: Obtain the audio corresponding to the target music, and use an audio analysis tool to process the audio corresponding to the target music to obtain the tempo, time signature, and chord list of the target music. The second one: Obtain the score corresponding to the target music, and based on the score corresponding to the target music, determine the tempo, time signature, and chord list of the target music. Among them, the score can be a staff notation or a numbered musical notation, and the embodiments of the present application do not limit this. The third one: Obtain the electronic score of the target music, and use a score analysis tool to process the electronic score of the target music to obtain the tempo, time signature, and chord list of the target music. Among them, the electronic score is composed of the notes corresponding to each beat included in the target music, and the electronic score may also include information such as tempo and time signature.

[0095] Optionally, the process of using an audio analysis tool to process the audio corresponding to the target music to obtain the tempo, time signature, and chord list of the target music is: input the audio corresponding to the target music into the audio analysis tool, and based on the output result of the audio analysis tool, obtain the tempo, time signature, and chord list of the target music. The audio analysis tool is used to analyze the audio, and then obtain the tempo, time signature, and chord list corresponding to the audio. Of course, when the audio analysis tool analyzes the audio, it can also obtain other information of the audio, and the embodiments of the present application do not limit this. The audio analysis tool can be a machine learning model, such as a neural network model, etc.

[0096] Optionally, the process of determining the tempo, time signature, and chord list of the target music based on the score corresponding to the target music is: a user with musical literacy determines the tempo, time signature, and chord list of the target music based on the score corresponding to the target music.

[0097] Optionally, the process of using a music score analysis tool to process the electronic music score of the target music to obtain the tempo, time signature, and chord list of the target music is as follows: Input the electronic music score corresponding to the target music into the music score analysis tool, and the music score analysis tool analyzes the electronic music score of the target music to obtain the tempo, time signature, and chord list of the target music. The specific process is as follows:

[0098] There is a chord library stored in the computer device, and the corresponding relationship between chord identifiers and electronic music scores of chords is stored in the chord library. The process by which the music score analysis tool analyzes the electronic music score of the target music to obtain the chord list of the target music is as follows: The music score analysis tool obtains an electronic music score segment corresponding to a certain music measure, searches for the electronic music score of the chord that matches this electronic music score segment in the above corresponding relationship, and determines the chord identifier corresponding to the found electronic music score of the chord as the chord identifier of this music measure. Furthermore, the performance time information of this music measure and the chord identifier corresponding to this music measure can be obtained. Traverse all the music measures of the target music according to this method, thereby obtaining the chord list of the target music. In addition, the music score analysis tool can directly obtain the tempo and time signature from the electronic music score of the target music.

[0099] Among them, the chord list includes chord identifiers and the performance time information corresponding to the chord identifiers. The chord identifier can be the chord name or a string composed of the notes that make up the chord. The embodiments of the present application do not limit this. Exemplarily, the chord name is C chord, and the notes that make up the C chord are 123. The chord identifier can be C chord or 123.

[0100] Optionally, the performance time information includes any two of the start beat, end beat, and duration beat. For example, the performance time information includes the start beat and the end beat. Exemplarily, the performance time information is (1, 4), that is, the performance time information is from the 1st beat to the 4th beat. Another example, the performance time information includes the start beat and the duration beat. Exemplarily, the performance time information is [1, 4], that is, the performance time information is from the 1st beat and lasts for 4 beats. Another example, the performance time information includes the duration beat and the end beat. Exemplarily, the performance time information is [4, 4], that is, the performance time information is to last for 4 beats and end at the 4th beat.

[0101] Exemplarily, the time signature of the target music is 4 / 4 beat, the tempo is 60 beats per minute, and the chord list is shown in Table 1 below. Among them, 4 / 4 beat means that a quarter note is one beat, and a music measure has 4 beats; 60 beats per minute means that there are 60 beats in one minute, and the time interval between each beat is 1 second.

[0102] Table 1

[0103]

[0104]

[0105] As shown in Table 1 above, among them, (1,4) is used to indicate starting from the 1st beat and ending at the 4th beat, and N.C is used to indicate no chord. The chord identifiers and the performance time information corresponding to the chord identifiers are shown in Table 1 above, and will not be elaborated here one by one.

[0106] It should be noted that the above is only an example of the chord identifiers included in the target music provided by the embodiments of the present application and the performance time information corresponding to the chord identifiers, and does not limit the chord identifiers included in the target music and the performance time information corresponding to the chord identifiers.

[0107] In a possible implementation manner, the multiple sub-audios include a drumbeat sub-audio and a chord sub-audio. The process of determining the audio data identifiers and performance time information corresponding to the multiple sub-audios based on the tempo, time signature, and chord list of the target music is as follows: Based on the tempo and time signature of the target music, determine the audio data identifiers and performance time information corresponding to the drumbeat sub-audio; based on the tempo, time signature, and chord list of the target music, determine the audio data identifiers and performance time information corresponding to the chord sub-audio. The audio data identifiers and performance time information corresponding to the drumbeat sub-audio, and the audio data identifiers and performance time information corresponding to the chord sub-audio, together constitute the audio data identifiers and performance time information corresponding to the multiple sub-audios.

[0108] Among them, the process of determining the audio data identifiers and performance time information corresponding to the drumbeat sub-audio based on the tempo and time signature of the target music is as follows: Determine the audio data identifiers corresponding to the time signature and tempo of the target music, and use the audio data identifiers corresponding to the time signature and tempo of the target music as the audio data identifiers corresponding to the drumbeat sub-audio; based on the time signature and tempo of the target music, determine the performance time information corresponding to the drumbeat sub-audio.

[0109] Optionally, before obtaining the audio data identifiers and performance time information corresponding to the drumbeat sub-audio, it is necessary to first determine the drum instrument. The process of determining the drum instrument can be to manually specify a drum instrument among multiple drum instruments, or a computer device can randomly determine a drum instrument. The embodiments of the present application do not limit this. It should be noted that whether it is a drum instrument specified manually or a drum instrument randomly determined by a computer device, the instrument timbre of the determined drum instrument matches the hearing-impaired hearing timbre.

[0110] Exemplarily, the determined drum instrument is a snare drum.

[0111] In a possible implementation, after determining the percussion instrument, multiple percussion sub-audios corresponding to the determined percussion instrument are obtained from the first audio library. Then, based on the tempo and time signature of the target music, the percussion sub-audio corresponding to the tempo and time signature of the target music is determined from the multiple percussion sub-audios, and the audio data identifier corresponding to the percussion sub-audio corresponding to the tempo and time signature of the target music is used as the audio data identifier corresponding to the percussion sub-audio included in the sheet music data.

[0112] Optionally, the computer device pre-stores a first audio library, which stores multiple percussion sub-audios, and the instrument timbres corresponding to the multiple percussion sub-audios stored in the first audio library match the hearing-impaired hearing timbre. Each percussion sub-audio in the first audio library corresponds to an audio data identifier.

[0113] Among them, the percussion sub-audios stored in the first audio library are audio segments in the MP3 (Moving Picture Experts Group Audio Layer III) format, or audio segments in other formats, which are not limited in this embodiment of the present application.

[0114] As shown in Table II below, it is a table of the corresponding relationship between the audio data identifier corresponding to the percussion sub-audio of the snare drum stored in the first audio library provided by the embodiment of the present application, and the tempo and time signature corresponding to the percussion sub-audio.

[0115] Table II

[0116] Time signature Tempo Audio data identifier 4 / 4 time 60 beats per minute A1 4 / 4 time 30 beats per minute A2 4 / 4 time 80 beats per minute A3 3 / 4 time 60 beats per minute A4 3 / 4 time 30 beats per minute A5 3 / 4 time 80 beats per minute A6

[0117] Based on Table II above, when the time signature is 4 / 4 and the tempo is 60 beats per minute, the audio data identifier corresponding to the percussion sub-audio is A1. When the time signature and tempo are other values, the audio data identifier corresponding to the percussion sub-audio can be seen in Table II above, which will not be elaborated here.

[0118] It should be noted that the percussion sub-audios corresponding to different audio data identifiers are different. For example, when the audio data identifier is A1, the corresponding percussion sub-audio is an audio of 4 beats with a time interval of one second between each beat. When the audio data identifier is A2, the corresponding percussion sub-audio is an audio of 4 beats with a time interval of 2 seconds between each beat.

[0119] It should also be noted that Table II above is only an example of the corresponding relationship between the audio data identifier corresponding to the percussion sub-audio and the tempo and time signature corresponding to the percussion sub-audio provided by the embodiment of the present application, and does not limit the first audio library. The first audio library includes various percussion instruments and the percussion sub-audios corresponding to various time signatures and various tempos.

[0120] Exemplarily, the determined percussion instrument is a snare drum, the tempo of the target music is 60 beats per minute, and the time signature is 4 / 4. Determine multiple drum point audio clips corresponding to the snare drum in the first audio library. Identify the audio data of the drum point audio clip corresponding to the tempo and time signature of the target music among the multiple drum point audio clips as the audio data identifier corresponding to the drum point audio clip included in the musical score data. That is, determine the audio data identifier A1 as the audio data identifier corresponding to the drum point audio clip included in the musical score data of the target music.

[0121] In a possible implementation, the process of determining the performance time information corresponding to the drum point audio clip based on the time signature and tempo of the target music is as follows: Determine the total number of beats included in the target music based on the tempo of the target music and the duration of the target music. Determine the number of musical measures included in the target music based on the time signature of the target music and the total number of beats included in the target music. Based on the number of musical measures included in the target music and the time signature of the target music, determine the performance time information corresponding to each musical measure, and use the performance time information corresponding to each musical measure as the performance time information corresponding to the drum point audio clip.

[0122] Exemplarily, if the tempo of the target music is 60 beats per minute and the duration is 1 minute, then the total number of beats included in the target music is 60 beats. If the time signature of the target music is 4 / 4, then based on the time signature of the target music and the total number of beats included in the target music, it is determined that there are 15 musical measures in the target music. Since each musical measure includes 4 beats and there are 15 musical measures in total, the performance time information corresponding to each musical measure can be determined, and then the performance time information corresponding to each musical measure is used as the performance time information corresponding to the drum point audio clip.

[0123] Exemplarily, taking the tempo of the target music as 60 beats per minute, the time signature as 4 / 4, the duration as 1 minute, and the performance time information including the starting beat and the continuous beats as an example, the total number of beats included in the target music is 60 beats, the number of musical measures included is 15, and the performance time information corresponding to each musical measure is: (1, 4), (5, 8), (9, 12), (13, 16), (17, 20), (21, 24), (25, 28), (29, 32), (33, 36), (37, 40), (41, 44), (45, 48), (49, 52), (53, 56), (57, 60). Therefore, the performance time information corresponding to the drum point audio clip is also (1, 4), (5, 8), (9, 12), (13, 16), (17, 20), (21, 24), (25, 28), (29, 32), (33, 36), (37, 40), (41, 44), (45, 48), (49, 52), (53, 56), (57, 60).

[0124] In a possible implementation manner, the process of determining the audio data identifier and performance time information corresponding to the chord sub-audio based on the tempo, time signature, and chord list of the target music is as follows: Based on the tempo and time signature of the target music, determine the audio data identifier corresponding to the chord identifier. Determine the performance time information and audio data identifier corresponding to the chord sub-audio as the performance time information and audio data identifier corresponding to the chord identifier.

[0125] Optionally, before obtaining the audio data identifier and performance time information corresponding to the chord sub-audio, it is necessary to first determine the chord instrument. The process of determining the chord instrument can be for a human to specify a chord instrument among multiple chord instruments, or for a computer device to randomly determine a chord instrument. The embodiments of the present application do not limit this. It should be noted that whether it is a chord instrument specified manually or a chord instrument randomly determined by a computer device, the instrument timbre of the determined chord instrument matches the hearing-impaired hearing timbre.

[0126] Exemplarily, the determined chord instrument is a bass.

[0127] Optionally, a second audio library is pre-stored in the computer device. The second audio library stores multiple chord sub-audios, and the instrument timbres of the multiple chord sub-audios stored in the second audio library match the hearing-impaired hearing timbre. Each chord sub-audio in the second audio library corresponds to an audio data identifier.

[0128] Among them, the chord sub-audios stored in the second audio library are audio segments in MP3 format or audio segments in other formats. The embodiments of the present application do not limit this.

[0129] As shown in Table 3 below, it is a table of the corresponding relationship between the audio data identifier corresponding to the chord sub-audio of the bass stored in the second audio library provided by the embodiments of the present application, and the tempo, time signature, and chord identifier corresponding to the chord sub-audio.

[0130] Table 3

[0131]

[0132] Based on Table 3 above, when the time signature is 4 / 4 and the tempo is 60 beats per minute, the audio data identifier corresponding to the chord sub-audio of the A chord is B1. When the time signature and tempo are other values, the audio data identifier corresponding to the chord sub-audio of the A chord can be seen in Table 3 above and will not be elaborated one by one here.

[0133] It should be noted that the chord sub-audio corresponding to different audio data identifiers is different. For example, the chord sub-audio corresponding to the audio data identifier B1 is an audio of A chord with 4 beats and a time interval of one second between each beat. The chord sub-audio corresponding to the audio data identifier B2 is an audio of A chord with 4 beats and a time interval of 2 seconds between each beat.

[0134] It should also be noted that Table 3 above is only an example table of the corresponding relationship between chord identifiers, tempos, time signatures, and audio data identifiers provided by the embodiments of the present application, and does not limit the second audio library. The second audio library includes various chord instruments and the chord sub-audio corresponding to various chord identifiers at various time signatures and various tempos.

[0135] In a possible implementation manner, since the performance time information corresponding to the chord identifier already exists in the chord list of the target music, and the audio data identifier corresponding to the chord identifier is determined based on Table 3 above, therefore, the performance time information corresponding to the chord identifier and the audio data identifier are determined as the performance time information and audio data identifier corresponding to the chord sub-audio included in the sheet music data.

[0136] Exemplarily, taking the tempo of the target music as 60 beats per minute, the time signature as 4 / 4, and the duration as 1 minute as an example, based on the above process, the sheet music data corresponding to the target music is shown in Table 4 below.

[0137] Table 4

[0138] Performance time information corresponding to the sub-audio Audio data identifier corresponding to the sub-audio (1,4) A1 (5,8) A1 (9,12) A1, B1 (13,16) A1, E1 (17,20) A1, C1 (21,24) A1, B1 … A1, E1 (57,60) A1, H1

[0139] As can be seen from Table 4 above, from the 1st beat to the 4th beat, the corresponding sub-audio is the drumbeat sub-audio corresponding to the audio data identifier A1. From the 5th beat to the 8th beat, the corresponding sub-audio is the drumbeat sub-audio corresponding to the audio data identifier A1. From the 9th beat to the 12th beat, the corresponding sub-audio is the drumbeat sub-audio corresponding to the audio data identifier A1 and the chord sub-audio corresponding to the audio data identifier B1. The audio data identifiers of the sub-audio corresponding to other performance time information are shown in Table 4 above and will not be elaborated here one by one.

[0140] Optionally, a user with musical literacy can also obtain the sheet music data of the target music based on the MIDI file of the target music. That is, the user determines the audio data identifier and performance time information corresponding to the drumbeat sub-audio, and / or the audio data identifier and performance time information corresponding to the chord sub-audio based on the MIDI file of the target music. Then, based on the input operation of the user on the computer device, the computer device obtains the sheet music data of the target music.

[0141] In step 202, the corresponding sub-audio is obtained based on each audio data identifier.

[0142] In a possible implementation, after determining the audio data identifiers corresponding to multiple sub-audios based on the above step 201, based on the audio data identifier corresponding to each sub-audio, the sub-audio corresponding to each audio data identifier is extracted from the audio library.

[0143] Optionally, the drumbeat sub-audio corresponding to the audio data identifier of the drumbeat sub-audio is extracted from the first audio library. For example, the drumbeat sub-audio corresponding to the audio data identifier A1 is extracted from the first audio library. The chord sub-audio corresponding to the audio data identifier of the chord sub-audio is extracted from the second audio library. For example, the chord sub-audio corresponding to the audio data identifier B1 is extracted from the second audio library.

[0144] In a possible implementation, when the number of beats included in the performance time information corresponding to the first audio data identifier is less than one musical measure, the sub-audio corresponding to the first audio data identifier is obtained from the audio library, and intercepted in the sub-audio corresponding to the first audio data identifier according to the number of beats included in the performance time information corresponding to the first audio data identifier, to obtain the sub-audio corresponding to the performance time information corresponding to the first audio data identifier, and the number of beats of the sub-audio corresponding to the performance time information corresponding to the first audio data identifier is the same as the number of beats included in the performance time information corresponding to the first audio data identifier.

[0145] Exemplarily, the first audio data identifier is B1, the performance time information corresponding to the first audio data identifier is (5,7) beats, and the number of beats included is 3 beats. Therefore, the sub-audio with the audio data identifier B1 is obtained from the audio library, and 3 / 4 is intercepted in the sub-audio with the audio data identifier B1 to obtain the sub-audio corresponding to B1 at (5,7) beats.

[0146] In step 203, based on the performance time information corresponding to each sub-audio, each sub-audio is fused to generate the synthesized audio of the target music.

[0147] In a possible implementation, based on the performance time information corresponding to each sub-audio, each sub-audio is fused to obtain the intermediate audio of the target music, and the intermediate audio of the target music is used as the synthesized audio of the target music.

[0148] Among them, there are the following two situations to fuse each sub-audio based on the performance time information corresponding to each sub-audio to obtain the intermediate audio of the target music.

[0149] Situation 1: In response to the non-existence of sub-audios with overlapping performance time information among multiple sub-audios, based on the performance time information corresponding to each sub-audio, the multiple sub-audios are spliced to obtain the intermediate audio of the target music.

[0150] Since the drumbeat audio needs to run through the entire music, when there are no sub-audios with overlapping performance time information among multiple sub-audios, it indicates that the target music only includes drumbeat audio and does not include chord sub-audios, or only includes chord sub-audios and does not include drumbeat audio, and each performance time information corresponds to only one chord sub-audio.

[0151] Optionally, when splicing multiple sub-audios to obtain the intermediate audio of the target music, each sub-audio can be processed with fade-in and fade-out first to obtain multiple sub-audios that have been processed with fade-in and fade-out, and then the multiple sub-audios that have been processed with fade-in and fade-out are spliced to obtain the intermediate audio of the target music. The purpose of the fade-in and fade-out processing is to prevent the spliced intermediate audio from being distorted, thereby making the intermediate audio more coherent.

[0152] The process of performing fade-in and fade-out processing on a sub-audio is as follows: perform fade-in processing on the head of the sub-audio and perform fade-out processing on the tail of the sub-audio to obtain a sub-audio that has been processed with fade-in and fade-out.

[0153] Among them, the duration of the fade-in processing and the duration of the fade-out processing need to be the same, and the duration of the fade-in processing and the fade-out processing are not limited in the embodiments of the present application. For example, if the duration of the fade-in processing and the fade-out processing is 50 milliseconds, then the first 50 milliseconds of the sub-audio are processed with fade-in, and the last 50 milliseconds of the sub-audio are processed with fade-out.

[0154] Exemplarily, the target music only includes drumbeat audio, and the performance time information corresponding to the drumbeat audio is respectively (1, 4), (5, 8), (9, 12), (13, 16). The drumbeat audio is processed with fade-in and fade-out to obtain a drumbeat audio that has been processed with fade-in and fade-out. The drumbeat audio that has been processed with fade-in and fade-out is spliced four times to obtain the intermediate audio of the target music. The intermediate audio includes four segments of drumbeat audio that have been processed with fade-in and fade-out.

[0155] Optionally, when splicing multiple sub-audios that have been processed with fade-in and fade-out, adjacent two sub-audios can also be cross-faded, that is, the tail of the sub-audio in the front position and the head of the sub-audio in the rear position are cross-mixed together, thereby obtaining the intermediate audio of the target music. Among them, the duration of the cross-mixed part of adjacent two sub-audios can be any value, and the embodiments of the present application do not limit this. For example, the duration of the cross-mixed part of adjacent two sub-audios is 200 milliseconds. That is, the last 200 milliseconds of the sub-audio in the front position and the first 200 milliseconds of the sub-audio in the rear position are cross-mixed together.

[0156] Case 2: In response to there being at least two first sub-audios corresponding to the same performance time information, mix the at least two first sub-audios to obtain a second sub-audio. The performance time information corresponding to the second sub-audio is the same as the performance time information corresponding to the at least two first sub-audios. Then, perform fade-in and fade-out processing on the second sub-audio and the third sub-audio respectively to obtain the second sub-audio after fade-in and fade-out processing and the third sub-audio after fade-in and fade-out processing, where the third sub-audio is a sub-audio with performance time information different from that of the second sub-audio. According to the performance time information corresponding to the second sub-audio and the performance time information corresponding to the third sub-audio, splice the second sub-audio after fade-in and fade-out processing and the third sub-audio after fade-in and fade-out processing to obtain the intermediate audio of the target music.

[0157] Exemplarily, the target music has 8 beats. There are drum sub-audios from the 1st beat to the 4th beat and from the 4th beat to the 8th beat, and there is a chord sub-audio from the 5th beat to the 8th beat. Therefore, mix the drum sub-audio from the 5th beat to the 8th beat and the chord sub-audio from the 5th beat to the 8th beat to obtain a second sub-audio. The performance time information corresponding to the second sub-audio is (5, 8). Then, perform fade-in and fade-out processing on the drum sub-audio from the 1st beat to the 4th beat to obtain the drum sub-audio from the 1st beat to the 4th beat after fade-in and fade-out processing. Perform fade-in and fade-out processing on the second sub-audio from the 5th beat to the 8th beat to obtain the second sub-audio from the 5th beat to the 8th beat after fade-in and fade-out processing. Then, splice the drum sub-audio from the 1st beat to the 4th beat after fade-in and fade-out processing and the second sub-audio from the 5th beat to the 8th beat after fade-in and fade-out processing to obtain the intermediate audio of the target music.

[0158] Optionally, when splicing the second sub-audio after fade-in and fade-out processing and the third sub-audio after fade-in and fade-out processing, cross-fading processing can also be performed on any two adjacent sub-audios among the second sub-audio after fade-in and fade-out processing and the third sub-audio after fade-in and fade-out processing. The process of cross-fading processing is as shown in Case 1 above and will not be elaborated here.

[0159] Optionally, after obtaining the intermediate audio of the target music, ambient sound can also be added to the intermediate audio to obtain the intermediate audio with ambient sound added, and use the intermediate audio with ambient sound added as the synthesized audio of the target music.

[0160] Among them, there is a third audio library stored in the computer device. The third audio library stores various types of ambient sounds, such as the sound of rain, the sound of cicadas, the sound of the coast, and so on. The duration of the ambient sounds stored in the third audio library is of any duration, which is not limited in this embodiment of the present application. The ambient sounds stored in the third audio library are sounds that hearing-impaired patients can hear. The ambient sounds stored in the third audio library are audio segments in MP3 format or audio segments in other formats, which are not limited in this embodiment of the present application.

[0161] Generally, ambient sound is added at the beginning of a piece of music. Of course, ambient sound can also be added at other positions in the music work. The type of the added ambient sound and the position where the ambient sound is added are both manually set, and the embodiments of the present application do not limit this.

[0162] Optionally, when adding a target ambient sound at a target position of the target music, it is determined whether the duration of the target ambient sound is consistent with the duration corresponding to the target position. If the duration of the target ambient sound is not consistent with the duration corresponding to the target position, the target ambient sound is first subjected to frame insertion / removal processing so that the duration of the target ambient sound after frame insertion / removal is consistent with the duration corresponding to the target position. Then, the target ambient sound after frame insertion / removal is mixed with the audio at the target position to obtain the target audio at the target position. Next, the target audio at the target position is spliced with the audio in the intermediate audio except for the audio at the target position to obtain the synthesized audio of the target music.

[0163] If the duration of the target ambient sound is consistent with the duration corresponding to the target position, the target ambient sound is mixed with the audio at the target position to obtain the target audio at the target position. Then, the target audio at the target position is spliced with the audio in the intermediate audio except for the audio at the target position to obtain the synthesized audio of the target music.

[0164] Exemplarily, when adding an ambient sound of "rain sound" from the 1st to the 3rd second in the intermediate audio of the target music, and the duration of the ambient sound of "rain sound" is 2 seconds, the ambient sound of "rain sound" is first subjected to frame insertion processing to obtain the ambient sound of "rain sound" after frame insertion processing. The duration of the ambient sound of "rain sound" after frame insertion processing is 3 seconds. The ambient sound of "rain sound" after frame insertion processing is mixed with the audio from the 1st to the 3rd second in the intermediate audio of the target music to obtain the target audio from the 1st to the 3rd second. Then, the target audio from the 1st to the 3rd second is spliced with the audio in the intermediate audio except for the audio from the 1st to the 3rd second to obtain the synthesized audio of the target music.

[0165] Optionally, the intermediate audio of the target music can also be subjected to frequency domain compression processing to obtain the synthesized audio of the target music.

[0166] Optionally, the process of performing frequency-domain compression processing on the intermediate audio of the target music to obtain the synthesized audio of the target music is as follows: Obtain the first sub-audio in the first frequency range and the second sub-audio in the second frequency range corresponding to the intermediate audio, where the frequency of the first frequency range is less than the frequency of the second frequency range. Based on the first gain coefficient, perform gain compensation on the first sub-audio to obtain the third sub-audio. Based on the second gain coefficient, perform gain compensation on the second sub-audio to obtain the fourth sub-audio. Perform compression frequency shift processing on the fourth sub-audio to obtain the fifth sub-audio, where the lower limit of the third frequency range corresponding to the fifth sub-audio is equal to the lower limit of the second frequency range. Perform fusion processing on the third sub-audio and the fifth sub-audio to obtain the synthesized audio of the target music.

[0167] Among them, the intermediate audio can be analyzed based on the analysis filter in the quadrature mirror filter bank to obtain the first sub-audio in the first frequency range and the second sub-audio in the second frequency range. Alternatively, the intermediate audio can be processed based on a frequency divider to obtain the first sub-audio in the first frequency range and the second sub-audio in the second frequency range. Of course, the first sub-audio and the second sub-audio can also be obtained in other ways, and the embodiments of the present application do not limit this.

[0168] Each frequency range includes one or more frequency bands, and each frequency band corresponds to a gain coefficient. Based on the gain coefficient corresponding to each frequency band, determine the decibel compensation value corresponding to each frequency band. Based on the decibel compensation value corresponding to each frequency band, perform gain compensation on the audio corresponding to each frequency band to obtain the audio after gain compensation for this frequency range.

[0169] Exemplarily, the first frequency range is from 0 to 1 kHz, the first frequency range includes only one frequency band, and the gain coefficient corresponding to the frequency band from 0 to 1 kHz is 2. Based on the gain coefficient 2 corresponding to the frequency band from 0 to 1 kHz, determine the decibel compensation value corresponding to the frequency band from 0 to 1 kHz. Perform gain compensation on the first sub-audio based on the decibel compensation value corresponding to the frequency band from 0 to 1 kHz to obtain the third sub-audio.

[0170] For another example, the second frequency range is from 1 kHz to 8 kHz. The second frequency range includes three frequency bands, namely: the first frequency band: from 1 kHz to 2 kHz, the second frequency band: from 2 kHz to 4 kHz, and the third frequency band: from 4 kHz to 8 kHz. The gain coefficient corresponding to the first frequency band is 2.5, the gain coefficient corresponding to the second frequency band is 3, and the gain coefficient corresponding to the third frequency band is 3.5. Therefore, based on the gain coefficient corresponding to the first frequency band, the decibel compensation value corresponding to the first frequency band is determined; based on the gain coefficient corresponding to the second frequency band, the decibel compensation value corresponding to the second frequency band is determined; and based on the gain coefficient corresponding to the third frequency band, the decibel compensation value corresponding to the third frequency band is determined. The audio in the first frequency band is subjected to gain compensation according to the decibel compensation value corresponding to the first frequency band, the audio in the second frequency band is subjected to gain compensation according to the decibel compensation value corresponding to the second frequency band, and the audio in the third frequency band is subjected to gain compensation according to the decibel compensation value corresponding to the third frequency band, thereby obtaining the fourth sub-audio.

[0171] Optionally, the process of performing compression frequency shift processing on the fourth sub-audio to obtain the fifth sub-audio is as follows: performing frequency compression on the fourth sub-audio at a target ratio to obtain a sixth sub-audio, and performing frequency upward shift on the sixth sub-audio by a target value to obtain the fifth sub-audio, where the target value is equal to the difference between the lower limit of the second frequency range and the lower limit of the fourth frequency range corresponding to the sixth sub-audio.

[0172] Since there is an overlap between the frequency range of the sixth sub-audio obtained by performing frequency compression on the fourth sub-audio at a target ratio and the first frequency range corresponding to the third sub-audio, it is necessary to perform frequency upward shift on the sixth sub-audio by a target value to obtain the fifth sub-audio, so that there is no overlap between the frequency range corresponding to the fifth sub-audio and the first frequency range corresponding to the third sub-audio, thereby making the listening experience of the subsequent synthesized audio better.

[0173] Among them, the target ratio can be any value, and the embodiments of the present application do not limit this. For example, the target ratio is 50%.

[0174] Exemplarily, the target ratio is 50%, the second frequency range corresponding to the fourth sub-audio is from 1 kHz to 8 kHz. After performing frequency compression on the fourth sub-audio at the target ratio, a sixth sub-audio is obtained, and the fourth frequency range corresponding to the sixth sub-audio is from 500 Hz to 4 kHz. Based on the lower limit of the fourth frequency range and the lower limit of the second frequency range, the target value is determined to be 500. Therefore, the frequency of the sixth sub-audio is shifted upward by 500 Hz to obtain the fifth sub-audio, and the third frequency range corresponding to the fifth sub-audio is from 1 kHz to 4.5 kHz.

[0175] Optionally, the method for fusing the third sub-audio and the fifth sub-audio to obtain the synthesized audio of the target music includes, but is not limited to: processing the third sub-audio and the fifth sub-audio through the synthesis filter of the quadrature mirror filter bank to obtain the synthesized audio of the target music. Or, mixing the third sub-audio and the fifth sub-audio to obtain the synthesized audio of the target music.

[0176] When mixing the third sub-audio and the fifth sub-audio, there is an easy problem of popping. Therefore, a limiter can also be used to process the audio after mixing the third sub-audio and the fifth sub-audio, and then obtain the synthesized audio of the target music.

[0177] Optionally, after obtaining the synthesized audio of the target music, the synthesized audio of the target music can also be played for the hearing-impaired patients to listen to. In response to receiving a modification instruction for the timbre of the target sub-audio in the synthesized audio from the hearing-impaired patients, an interactive page is displayed, and a drum control, a chord control, and an ambient sound control are displayed on the interactive page. In response to receiving a selection instruction for any control, a plurality of sub-controls included in the control are displayed, and each sub-control corresponds to a sub-audio. In response to a selection instruction for any one of the plurality of sub-controls, the sub-audio corresponding to the selected sub-control is played. In response to receiving a confirmation instruction for the selected sub-control, the sub-audio corresponding to the selected sub-control is used to replace the target sub-audio, and then the synthesized audio of the modified target music is obtained.

[0178] For example, in response to a selection instruction for the drum control, the drum sub-controls are displayed, and each drum sub-control corresponds to a drum sub-audio. In response to a selection instruction for any one of the plurality of drum sub-controls, the drum sub-audio corresponding to the selected drum sub-control is played. In response to receiving a confirmation instruction for the selected drum sub-control, the sub-audio corresponding to the selected drum sub-control is used to replace the target sub-audio, and then the synthesized audio of the modified target music is obtained.

[0179] The above method re-composes the target music. When composing the music, the instrument timbre of the sub-audio used matches the hearing-impaired hearing timbre, so that the hearing-impaired patients can hear the sub-audio used in the composition, and then obtain the synthesized audio of the target music based on the sub-audio. When the hearing-impaired patients listen to the synthesized audio of the target music, there will be no problems of intermittency or occasional inaudibility, and there will also be no distortion. The hearing-impaired patients can hear smooth music, and the listening experience of the hearing-impaired patients is better, which can fundamentally solve the problems of poor sound quality and poor listening effect when the hearing-impaired patients listen to music.

[0180] Since the duration of a song is relatively long, the number of musical measures it contains is relatively large, and the number of beats it contains is also relatively large. Here, taking the fourth, fifth, and sixth musical measures in the song "Paradise" as the target music as an example, the process of obtaining the synthesized audio of the target music is described. Figure 3 The following shows the simplified score diagrams of the fourth, fifth, and sixth musical measures of the song "Paradise".

[0181] Obtain the electronic score of the target music, input the electronic score into a score analysis tool, and then obtain the tempo, time signature, and chord list of the target music. Among them, the tempo of the target music is 70 beats per minute, the time signature is 4 / 4, and the chord list is shown in Table 5 below.

[0182] Table 5

[0183] Performance time information Chord identifier (13,16) D chord (17,20) Dm chord (21,24) Am chord

[0184] Pre-set the instrument timbre of the drum beat audio used in the synthesized audio of the target music as drums, and the instrument timbre of the chord audio as rock bass. Since the tempo of the target music is 70 and the time signature is 4 / 4, therefore, determine the audio data identifier N1 in the first audio library, and use the drum beat audio corresponding to the audio data identifier N1 as the drum beat audio in the synthesized audio. Based on the tempo, time signature, and chord list of the target music, determine the audio data identifiers M1, M2, and M3 in the second audio library. Among them, the audio data identifier M1 corresponds to the chord audio of the D chord, the audio data identifier M2 corresponds to the chord audio of the Dm chord, and the audio data identifier M3 corresponds to the chord audio of the Am chord. Use the chord audio corresponding to the audio data identifiers M1, M2, and M3 respectively as the chord audio in the synthesized audio. Then obtain the score data of the target music, and the score data is shown in Table 6 below.

[0185] Table 6

[0186] Performance time information corresponding to the sub-audio Audio data identifier corresponding to the sub-audio (13,16) N1, M1 (17,20) N1, M2 (21,24) N1, M3

[0187] Next, extract the drum beat audio with the audio data identifier N1 in the first audio library, and extract the chord audio with the audio data identifiers M1, M2, and M3 in the second audio library. Since there are both drum beat audio and chord audio at the playing time information (13,16), (17,20), and (21,24), therefore, it is necessary to mix the drum beat audio and chord audio corresponding to each playing time information to obtain the mixed audio corresponding to each playing time information, that is, obtain the first mixed audio, the second mixed audio, and the third mixed audio.

[0188] Among them, the first mixed sub-audio is obtained based on the drum sub-audio with audio data identifier N1 and the chord sub-audio with audio data identifier M1, and the playing time information of the first mixed sub-audio is (13, 16). The second mixed sub-audio is obtained based on the drum sub-audio with audio data identifier N1 and the chord sub-audio with audio data identifier M2, and the playing time information of the second mixed sub-audio is (17, 20). The third mixed sub-audio is obtained based on the drum sub-audio with audio data identifier N1 and the chord sub-audio with audio data identifier M3, and the playing time information of the third mixed sub-audio is (21, 24).

[0189] After that, fade-in and fade-out processing is performed on each mixed sub-audio to obtain the mixed sub-audio after fade-in and fade-out processing. Immediately afterwards, two adjacent mixed sub-audio with playing time information in the mixed sub-audio after fade-in and fade-out processing are spliced to obtain the intermediate audio of the target music.

[0190] Optionally, when splicing two adjacent mixed sub-audio with playing time information, cross-fading processing can be performed on the two mixed sub-audio to be spliced, and then the intermediate audio of the target music is obtained.

[0191] Optionally, the intermediate audio of the target music is used as the synthesized audio of the target music. As Figure 4 shown is the simplified score diagram corresponding to the synthesized audio of the fourth, fifth, and sixth musical measures of the song "Paradise" generated through the above processing. Among them, the mark numbered 1 represents the drumbeat, and there is one drumbeat in each musical measure, located at the first beat of the musical measure.

[0192] Optionally, the intermediate audio of the target music is analyzed to obtain a first sub-audio and a second sub-audio. Gain compensation is performed on the first sub-audio to obtain a third sub-audio, and gain compensation is performed on the second sub-audio to obtain a fourth sub-audio. The frequency of the fourth sub-audio is compressed by 50% to obtain a sixth sub-audio. The frequency of the sixth sub-audio is shifted up by 500 Hz to obtain a fifth sub-audio. Furthermore, based on the third sub-audio and the fifth sub-audio, the synthesized audio of the target music is obtained.

[0193] Figure 5 Shown is the flowchart of an audio synthesis method provided by an embodiment of the present application. In Figure 5Among them, the target music is obtained, and by analyzing the target music, the score data of the target music is obtained. Based on the score data of the target music and the pre-stored audio library (the audio library includes the first audio library, the second audio library, and the third audio library, where multiple drumbeat audio clips are stored in the first audio library, multiple chord audio clips are stored in the second audio library, and multiple environmental sound audio clips are stored in the third audio library), it is determined which drumbeat audio clips, chord audio clips, and environmental sound audio clips are included in the target music. Since there may be a situation where there are at least two audio clips with the same performance time information, it is necessary to perform mixing processing on the at least two audio clips with the same performance time information. For example, Figure 5 in the Mth performance time information, there are tracks 1, 2... N. Among them, tracks 1, 2, and N respectively correspond to an audio clip. Based on a multi-channel mixer, the audio clips corresponding to tracks 1, 2, and N are mixed respectively to obtain the mixed audio clip. Fade-in and fade-out processing is performed on the mixed audio clip and the other audio clips among the multiple audio clips except for the audio clips with the same performance time information to obtain the audio clip after fade-in and fade-out processing. Then, the mixed audio clip after fade-in and fade-out processing and the other audio clips after fade-in and fade-out processing are spliced to obtain the intermediate audio of the target music.

[0194] At this time, the intermediate audio of the target music can be used as the synthesized audio of the target music. It is also possible to further process the intermediate audio of the target music to obtain the synthesized audio of the target music.

[0195] The process of further processing is as follows: within an orthogonal mirror filter bank, the first audio clip and the second audio clip are obtained. Gain compensation is performed on the first audio clip within a two-channel wide dynamic range compressor to obtain the third audio clip. Gain compensation is performed on the second audio clip to obtain the fourth audio clip. Nonlinear compression and frequency shift processing is performed on the fourth audio clip to obtain the fifth audio clip. Based on the third audio clip and the fifth audio clip, the synthesized audio of the target music is obtained.

[0196] Figure 6 The following shows a schematic structural diagram of an audio synthesis device provided by an embodiment of the present application, as Figure 6 shown, the device includes:

[0197] An acquisition module 601, configured to acquire the score data of the target music, where the score data includes audio data identifiers corresponding to multiple audio clips and performance time information, and the musical instrument timbre corresponding to each audio clip matches the hearing-impaired hearing timbre;

[0198] The acquisition module 601 is configured to acquire the corresponding audio clip based on each audio data identifier;

[0199] A generation module 602, configured to perform a fusion process on each sub-audio based on the performance time information corresponding to each sub-audio, so as to generate a synthesized audio of the target music.

[0200] Optionally, in the spectrum of the instrument corresponding to each sub-audio, the ratio of the energy of the low-frequency band to the energy of the high-frequency band is greater than a ratio threshold. The low-frequency band is the band below the frequency threshold, and the high-frequency band is the band above the frequency threshold. The ratio threshold is used to indicate the condition that the ratio of the energy of the low-frequency band to the energy of the high-frequency band in the spectrum of the audio that can be heard by the hearing-impaired patient needs to meet.

[0201] Optionally, an acquisition module 601, configured to determine the audio data identifier and the performance time information corresponding to a plurality of sub-audios based on the tempo, time signature, and chord list of the target music.

[0202] Optionally, the plurality of sub-audios include a drumbeat sub-audio and a chord sub-audio;

[0203] The acquisition module 601 is configured to determine the audio data identifier and the performance time information corresponding to the drumbeat sub-audio based on the tempo and time signature of the target music;

[0204] Based on the tempo, time signature, and chord list of the target music, determine the audio data identifier and the performance time information corresponding to the chord sub-audio;

[0205] The audio data identifier and the performance time information corresponding to the drumbeat sub-audio, and the audio data identifier and the performance time information corresponding to the chord sub-audio together constitute the audio data identifier and the performance time information corresponding to the plurality of sub-audios.

[0206] Optionally, the acquisition module 601 is configured to determine the audio data identifier corresponding to the time signature and tempo of the target music, and use the audio data identifier corresponding to the time signature and tempo of the target music as the audio data identifier corresponding to the drumbeat sub-audio;

[0207] Based on the time signature and tempo of the target music, determine the performance time information corresponding to the drumbeat sub-audio.

[0208] Optionally, the chord list includes a chord identifier and the performance time information corresponding to the chord identifier;

[0209] The acquisition module 601 is configured to determine the audio data identifier corresponding to the chord identifier based on the tempo and time signature of the target music;

[0210] Determine the performance time information and the audio data identifier corresponding to the chord sub-audio based on the performance time information and the audio data identifier corresponding to the chord identifier.

[0211] Optionally, a generation module 602 is configured to perform a fusion process on each sub-audio based on the playing time information corresponding to each sub-audio to obtain an intermediate audio of the target music;

[0212] Perform a frequency-domain compression process on the intermediate audio of the target music to obtain a synthesized audio of the target music.

[0213] Optionally, a synthesis module 602 is configured to obtain a first sub-audio in a first frequency range and a second sub-audio in a second frequency range corresponding to the intermediate audio, where the frequency of the first frequency range is less than the frequency of the second frequency range;

[0214] Based on a first gain coefficient, perform gain compensation on the first sub-audio to obtain a third sub-audio, and based on a second gain coefficient, perform gain compensation on the second sub-audio to obtain a fourth sub-audio;

[0215] Perform a compression and frequency shift process on the fourth sub-audio to obtain a fifth sub-audio, where the lower limit of a third frequency range corresponding to the fifth sub-audio is equal to the lower limit of the second frequency range;

[0216] Perform a fusion process on the third sub-audio and the fifth sub-audio to obtain a synthesized audio of the target music.

[0217] Optionally, the generation module 602 is configured to perform frequency compression on the fourth sub-audio at a target ratio to obtain a sixth sub-audio;

[0218] Perform a frequency upward shift on the sixth sub-audio by a target value to obtain a fifth sub-audio, where the target value is equal to the difference between the lower limit of the second frequency range and the lower limit of a fourth frequency range corresponding to the sixth sub-audio.

[0219] The above device re-composes the target music. When re-composing, the instrument timbre of the sub-audio used matches the hearing-impaired hearing timbre, enabling hearing-impaired patients to hear the sub-audio used in the re-composition. Furthermore, based on the sub-audio, a synthesized audio of the target music is obtained. When hearing-impaired patients listen to the synthesized audio of the target music, there will be no problems such as intermittency or occasional inaudibility, and there will also be no distortion. This enables hearing-impaired patients to hear smooth music, providing a better listening experience for hearing-impaired patients and fundamentally solving the problems of poor sound quality and poor listening effect when hearing-impaired patients listen to music.

[0220] It should be understood that the above Figure 6When the provided device implements its functions, only the division of the above-mentioned functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0221] Figure 7 The block diagram of a terminal device 700 provided by an exemplary embodiment of the present application is shown. The terminal device 700 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal device 700 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0222] Generally, the terminal device 700 includes: a processor 701 and a memory 702.

[0223] The processor 701 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0224] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction for being executed by the processor 701 to implement the audio synthesis method provided in the method embodiments of the present application.

[0225] In some embodiments, the terminal device 700 may further optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, a positioning assembly 708, and a power supply 709.

[0226] The peripheral device interface 703 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 may be implemented on a separate chip or circuit board, and the present embodiment does not limit this.

[0227] The radio frequency circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 may communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 704 may further include a circuit related to NFC (Near Field Communication), and the present application does not limit this.

[0228] The display screen 705 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input to the processor 701 as control signals for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 705, which is disposed on the front panel of the terminal device 700; in other embodiments, there may be at least two display screens 705, which are respectively disposed on different surfaces of the terminal device 700 or are in a foldable design; in other embodiments, the display screen 705 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal device 700. Even, the display screen 705 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0229] The camera module 706 is used to collect images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal device 700, and the rear camera is disposed on the back of the terminal device 700. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting and VR (Virtual Reality) shooting by fusing the main camera and the wide-angle camera, or other fused shooting functions. In some embodiments, the camera module 706 may further include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, and can be used for light compensation under different color temperatures.

[0230] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 701 for processing, or input to the radio frequency circuit 704 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal device 700. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into audible sound waves for humans, but also convert the electrical signal into inaudible sound waves for humans for uses such as ranging. In some embodiments, the audio circuit 707 may further include a headphone jack.

[0231] The positioning component 708 is used to locate the current geographical location of the terminal device 700 to implement navigation or LBS (Location Based Service). The positioning component 708 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.

[0232] The power supply 709 is used to supply power to each component in the terminal device 700. The power supply 709 may be alternating current, direct current, a primary battery, or a rechargeable battery. When the power supply 709 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0233] In some embodiments, the terminal device 700 further includes one or more sensors 170. The one or more sensors 170 include but are not limited to: an acceleration sensor 711, a gyroscope sensor 712, a pressure sensor 713, a fingerprint sensor 714, an optical sensor 715, and a proximity sensor 716.

[0234] The acceleration sensor 711 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal device 700. For example, the acceleration sensor 711 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 701 can control the display screen 705 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 711. The acceleration sensor 711 can also be used for game or collection of the user's motion data.

[0235] The gyroscope sensor 712 can detect the body direction and rotation angle of the terminal device 700. The gyroscope sensor 712 can cooperate with the acceleration sensor 711 to collect the 3D actions of the user on the terminal device 700. Based on the data collected by the gyroscope sensor 712, the processor 701 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0236] The pressure sensor 713 can be disposed on the side frame of the terminal device 700 and / or the lower layer of the display screen 705. When the pressure sensor 713 is disposed on the side frame of the terminal device 700, it can detect the holding signal of the user on the terminal device 700, and the processor 701 can perform left / right hand recognition or shortcut operations based on the holding signal collected by the pressure sensor 713. When the pressure sensor 713 is disposed on the lower layer of the display screen 705, the processor 701 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 705. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0237] The fingerprint sensor 714 is used to collect the fingerprint of the user. The processor 701 can identify the user's identity based on the fingerprint collected by the fingerprint sensor 714, or the fingerprint sensor 714 can identify the user's identity based on the collected fingerprint. When the identity of the user is identified as a trusted identity, the processor 701 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 714 can be disposed on the front, back, or side of the terminal device 700. When there are physical buttons or manufacturer logos on the terminal device 700, the fingerprint sensor 714 can be integrated with the physical buttons or manufacturer logos.

[0238] The optical sensor 715 is used to collect the ambient light intensity. In one embodiment, the processor 701 can control the display brightness of the display screen 705 according to the ambient light intensity collected by the optical sensor 715. Specifically, when the ambient light intensity is high, the display brightness of the display screen 705 is increased; when the ambient light intensity is low, the display brightness of the display screen 705 is decreased. In another embodiment, the processor 701 can also dynamically adjust the shooting parameters of the camera module 706 according to the ambient light intensity collected by the optical sensor 715.

[0239] A proximity sensor 716, also known as a distance sensor, is typically disposed on the front panel of the terminal device 700. The proximity sensor 716 is used to collect the distance between the user and the front of the terminal device 700. In one embodiment, when the proximity sensor 716 detects that the distance between the user and the front of the terminal device 700 is gradually decreasing, the processor 701 controls the display screen 705 to switch from the lit state to the off state; when the proximity sensor 716 detects that the distance between the user and the front of the terminal device 700 is gradually increasing, the processor 701 controls the display screen 705 to switch from the off state to the lit state.

[0240] Those skilled in the art can understand that Figure 7 the structure shown in does not constitute a limitation on the terminal device 700, and may include more or fewer components than shown, or combine certain components, or adopt a different component layout.

[0241] Figure 8 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 801 and one or more memories 802. Among them, at least one program code is stored in the one or more memories 802, and the at least one program code is loaded and executed by the one or more processors 801 to implement the audio synthesis methods provided by the above various method embodiments. Of course, the server 800 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server 800 may also include other components for implementing the functions of the device, which will not be elaborated here.

[0242] In an exemplary embodiment, a computer-readable storage medium is also provided. At least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to enable a computer to implement any of the above audio synthesis methods.

[0243] Optionally, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0244] In an exemplary embodiment, a computer program or a computer program product is further provided. At least one computer instruction is stored in the computer program or the computer program product. The at least one computer instruction is loaded and executed by a processor to enable a computer to implement any of the above audio synthesis methods.

[0245] It should be understood that "a plurality of" as mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0246] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0247] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An audio synthesis method, characterized in that, the method includes: Based on the tempo and time signature of the target music, determine the audio data identifier and performance time information corresponding to the drum beat sub-audio; Based on the tempo, time signature and chord list of the target music, determine the audio data identifier and performance time information corresponding to the chord sub-audio; The audio data identifier and performance time information corresponding to the drum beat sub-audio, and the audio data identifier and performance time information corresponding to the chord sub-audio, constitute the audio data identifier and performance time information corresponding to multiple sub-audios included in the sheet music data, wherein the instrument timbre corresponding to each sub-audio matches the hearing-impaired hearing timbre; Obtain the corresponding sub-audio based on each audio data identifier; Based on the performance time information corresponding to each sub-audio, perform a fusion process on each sub-audio to generate the synthesized audio of the target music.

2. The method according to claim 1, characterized in that, In the spectrum of the instrument corresponding to each sub-audio, the ratio of the energy of the low-frequency band to the energy of the high-frequency band is greater than the ratio threshold, the low-frequency band is the band below the frequency threshold, and the high-frequency band is the band above the frequency threshold, wherein the ratio threshold is used to indicate the ratio of the energy of the low-frequency band to the energy of the high-frequency band in the spectrum of the audio that can be heard by hearing-impaired patients. The conditions that need to be met.

3. The method according to claim 1, characterized in that, The determining the audio data identifier and performance time information corresponding to the drum beat sub-audio based on the tempo and time signature of the target music includes: Determine the audio data identifier corresponding to the time signature and tempo of the target music, and use the audio data identifier corresponding to the time signature and tempo of the target music as the audio data identifier corresponding to the drum beat sub-audio; Based on the time signature and tempo of the target music, determine the performance time information corresponding to the drum beat sub-audio.

4. The method according to claim 1, characterized in that, The chord list includes a chord identifier and the performance time information corresponding to the chord identifier; The determining the audio data identifier and performance time information corresponding to the chord sub-audio based on the tempo, time signature and chord list of the target music includes: Based on the tempo and time signature of the target music, determine the audio data identifier corresponding to the chord identifier; Determine the performance time information and audio data identifier corresponding to the chord sub-audio as the performance time information and audio data identifier corresponding to the chord identifier.

5. The method according to any one of claims 1 to 4, characterized in that, The performing a fusion process on each sub-audio based on the performance time information corresponding to each sub-audio to generate the synthesized audio of the target music includes: Based on the performance time information corresponding to each sub-audio, perform a fusion process on each sub-audio to obtain the intermediate audio of the target music; Perform a frequency domain compression process on the intermediate audio of the target music to obtain the synthesized audio of the target music.

6. The method according to claim 5, characterized in that, Performing frequency-domain compression processing on the intermediate audio of the target music to obtain the synthesized audio of the target music includes: Obtaining a first sub-audio in a first frequency range and a second sub-audio in a second frequency range corresponding to the intermediate audio, where the frequency of the first frequency range is less than the frequency of the second frequency range; Based on a first gain coefficient, performing gain compensation on the first sub-audio to obtain a third sub-audio, and based on a second gain coefficient, performing gain compensation on the second sub-audio to obtain a fourth sub-audio; Performing compression frequency shift processing on the fourth sub-audio to obtain a fifth sub-audio, where the lower limit of a third frequency range corresponding to the fifth sub-audio is equal to the lower limit of the second frequency range; Performing fusion processing on the third sub-audio and the fifth sub-audio to obtain the synthesized audio of the target music.

7. The method according to claim 6, wherein, Performing compression frequency shift processing on the fourth sub-audio to obtain a fifth sub-audio includes: Performing frequency compression on the fourth sub-audio at a target ratio to obtain a sixth sub-audio; Performing frequency upward shift on the sixth sub-audio by a target value to obtain the fifth sub-audio, where the target value is equal to the difference between the lower limit of the second frequency range and the lower limit of a fourth frequency range corresponding to the sixth sub-audio.

8. A computer device, wherein, The computer device includes a processor and a memory, and at least one program code is stored in the memory. The at least one program code is loaded and executed by the processor to enable the computer device to implement the audio synthesis method according to any one of claims 1 to 7.

9. A computer-readable storage medium, wherein, At least one program code is stored in the computer-readable storage medium. The at least one program code is loaded and executed by a processor to enable a computer to implement the audio synthesis method according to any one of claims 1 to 7.

10. A computer program product, wherein, At least one computer instruction is stored in the computer program product. The at least one computer instruction is loaded and executed by a processor to enable a computer to implement the audio synthesis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Personalized real-time audio generation based on user physiological response

    US10790919B1