Audio signal processing method and device, electronic equipment and storage medium

By customizing sound effects specifically for elderly users, and employing high-frequency sound enhancement, human voice enhancement, and audio dynamic range compression effects, the problem of hearing loss and health fragility among elderly users when watching multimedia resources is solved, improving the viewing experience and reducing health risks.

CN118921521BActive Publication Date: 2025-12-16BEIJING YOUKU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411171023.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-12-16
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

Existing technologies fail to optimize the sound effects of multimedia resources for elderly users who have hearing loss and physical fragility, resulting in a poor viewing experience and potentially adverse health effects.

Method used

An audio signal processing method is provided that customizes sound effects for elderly users, including high-frequency sound enhancement, voice enhancement, and audio dynamic range compression effects, to increase the amplitude of high-frequency and voice frequency band signals and compress the dynamic range. These effects are then presented when playing multimedia resources according to the user's selection.

Benefits of technology

It improves the viewing experience of multimedia resources for elderly users, reduces adverse effects on their health, enhances the hearing of high-frequency and human voice bands, and reduces the impact of dynamic range on health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118921521B_ABST
    Figure CN118921521B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of audio data processing, and proposes an audio signal processing method and device, electronic equipment and storage medium. The method comprises: showing a sound effect mode supported by a target multimedia resource to a user; in response to the user selecting a senior exclusive sound effect, causing the audio of the target multimedia resource to present a special effect corresponding to the senior exclusive sound effect when playing the target multimedia resource; the special effect corresponding to the senior exclusive sound effect comprises at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect and an audio dynamic range compression special effect, the high-frequency sound enhancement special effect is used to enhance the amplitude of signals belonging to a high-frequency band, the human voice enhancement special effect is used to enhance the amplitude of signals belonging to a human voice band, and the audio dynamic range compression special effect is used to compress the dynamic range of signals. The method supports the senior exclusive sound effect customized for the elderly users, can improve the viewing experience of the elderly users on the multimedia resources, and reduces the adverse effects of the elderly users watching the multimedia resources on the physical health.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of audio data processing, and particularly relates to an audio signal processing method and device, an electronic device, and a storage medium. BACKGROUND

[0002] Hearing is an important bridge for people to communicate with the outside world. Therefore, optimizing the audio playback effect of multimedia resources has become an important way to improve the viewing experience of users for multimedia resources.

[0003] The sound effect optimization that can be achieved by the prior art is mostly set from the perspective of the scene, such as simulating the sound effect of a stage play or the sound effect of a concert, without considering the sound preference of users of different ages. Compared with young users, the hearing of old users is poorer, and the sensitivity to sound is significantly decreased. At the same time, the heart of old users is relatively fragile, and some sound effects can cause old users to feel unwell.

[0004] Therefore, how to improve the viewing experience of old users for multimedia resources and reduce the adverse effects of old users watching multimedia resources on physical health has become a research hotspot in the field. SUMMARY

[0005] Therefore, the present disclosure provides an audio signal processing method, device, electronic device, and storage medium. The method supports customized sound effects for old users, and makes the audio of the target multimedia resource present special effects corresponding to the sound effects for old users, which can improve the viewing experience of old users for multimedia resources and reduce the adverse effects of old users watching multimedia resources on physical health.

[0006] According to an aspect of the present disclosure, an audio signal processing method is provided, which includes: showing a user a sound effect mode supported by a target multimedia resource; in response to the user selecting a sound effect for old users in the sound effect mode, making the audio of the target multimedia resource present special effects corresponding to the sound effect for old users when playing the target multimedia resource; wherein the special effects corresponding to the sound effect for old users include at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect, and an audio dynamic range compression special effect, the high-frequency sound enhancement special effect is used to increase the amplitude of signals belonging to the high-frequency band, the human voice enhancement special effect is used to increase the amplitude of signals belonging to the human voice band, and the audio dynamic range compression special effect is used to compress the dynamic range of signals.

[0007] In a possible implementation, in response to the user selecting the long-elder-specific sound effect in the sound effect mode, the audio of the target multimedia resource plays the special effect corresponding to the long-elder-specific sound effect when playing the target multimedia resource, including: in response to the user selecting the long-elder-specific sound effect in the sound effect mode, processing the original audio of the target multimedia resource based on the special effect corresponding to the long-elder-specific sound effect to obtain to-be-played audio; and using the to-be-played audio to replace the original audio when playing the target multimedia resource.

[0008] In a possible implementation, the processing of the original audio of the target multimedia resource based on the special effect corresponding to the long-elder-specific sound effect to obtain to-be-played audio includes: sequentially using one or more special effects corresponding to the long-elder-specific sound effect on the original audio to obtain the to-be-played audio; and when the long-elder-specific sound effect corresponds to multiple special effects, the original audio is used as a signal to be rendered by a first special effect, a rendering result corresponding to a previous special effect is used as a signal to be rendered by a next special effect, and a rendering result corresponding to a last special effect is used as the to-be-played audio.

[0009] In a possible implementation, the rendering parameter of the high-frequency sound enhancement special effect includes a first threshold and a first coefficient, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the long-elder-specific sound effect to obtain to-be-played audio includes: when the high-frequency sound enhancement special effect is used to render a signal to be rendered, the signal to be rendered is divided into a first signal and a second signal, the first signal is a signal in a high-frequency band; the amplitude of the first signal is increased according to the first coefficient to obtain a third signal; the amplitudes of the third signal and the first threshold are compared, and when the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain a fourth signal; and the fourth signal and the second signal are combined as a rendering result corresponding to the high-frequency sound enhancement special effect.

[0010] In a possible implementation, the rendering parameter of the human voice enhancement special effect includes a second coefficient and a third coefficient, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the long-elder-specific sound effect to obtain to-be-played audio includes: when the human voice enhancement special effect is used to render a signal to be rendered, the signal to be rendered is divided into a fifth signal and a sixth signal, the fifth signal is a signal in a human voice frequency band; the amplitude of the fifth signal is increased according to the second coefficient to obtain a seventh signal; the amplitude of a noise signal in the sixth signal is reduced according to the third coefficient to obtain an eighth signal; and the seventh signal and the eighth signal are combined as a rendering result corresponding to the human voice enhancement special effect.

[0011] In a possible implementation, the rendering parameter of the audio dynamic range compression special effect includes a second threshold, a third threshold, a fourth coefficient, and a fifth coefficient, the second threshold is greater than the third threshold, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the elder-specific sound effect to obtain the audio to be played includes: when the audio to be rendered is rendered using the audio dynamic range compression special effect, comparing the amplitude of the audio to be rendered with the second threshold and the third threshold; when the amplitude of the audio to be rendered is greater than the second threshold, reducing the amplitude of the audio to be rendered according to the fourth coefficient; and when the amplitude of the audio to be rendered is less than the third threshold, increasing the amplitude of the audio to be rendered according to the fifth coefficient.

[0012] In a possible implementation, the method further includes: receiving a feedback signal from a user when the target multimedia resource is played; and adjusting the rendering parameter of at least one special effect corresponding to the elder-specific sound effect according to the feedback signal.

[0013] In a possible implementation, the rendering parameter of any special effect corresponding to the elder-specific sound effect is a rendering parameter pre-set on the current terminal device or a rendering parameter pre-set on a user account logged in on the current terminal device.

[0014] In a possible implementation, the method further includes: in response to the start of playing of the target multimedia resource, querying an interface display mode of the current terminal device; in response to the interface display mode of the current terminal device being an elder mode, determining whether a playing state of the target multimedia resource meets a first preset condition; in response to the interface display mode of the current terminal device being other than the elder mode, querying a playing record of a user on the current terminal device or querying user identity information recorded by the current terminal device; in a case where the number of times that the user selects the elder-specific sound effect when playing a multimedia resource on the current terminal device is greater than a fourth threshold according to the playing record or in a case where the age of the user is greater than a fifth threshold according to the user identity information, determining whether the playing state of the target multimedia resource meets the first preset condition; and in response to the playing state of the target multimedia resource meeting the first preset condition, causing the audio of the target multimedia resource to present the special effect corresponding to the elder-specific sound effect when the target multimedia resource is played.

[0015] In a possible implementation, the first preset condition includes one or more of the following: the duration of full-screen playing of the target multimedia resource exceeds a sixth threshold; the audio played is the original audio of the target multimedia resource; the target multimedia resource is played in full screen on the interface of the current terminal device; the target multimedia resource is in a playing state; and the duration of content of the target multimedia resource that has not been played is greater than a seventh threshold.

[0016] In a possible implementation, in a case where the current terminal device plays the target multimedia resource in an audio mode other than the seniority-specific audio mode, the method further includes: receiving an audio mode selection signal from the associated device, the audio mode selection signal indicating an audio mode; and causing the audio of the target multimedia resource to present a special effect corresponding to the seniority-specific audio mode when the audio mode selection signal indicates the seniority-specific audio mode.

[0017] In a possible implementation, the method further includes: storing the to-be-played audio, and directly replacing the original audio with the stored to-be-played audio when the user plays the target multimedia resource again and is in the audio mode of the seniority-specific audio mode.

[0018] According to another aspect of the present disclosure, an audio signal processing apparatus is provided, which includes: a display module configured to display, to a user, audio modes supported by a target multimedia resource; and a first audio processing module configured to, in response to the user selecting a seniority-specific audio mode from the audio modes, cause audio of the target multimedia resource to present a special effect corresponding to the seniority-specific audio mode when the target multimedia resource is played; wherein the special effect corresponding to the seniority-specific audio mode includes at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect, and an audio dynamic range compression special effect, the high-frequency sound enhancement special effect is configured to increase an amplitude of a signal belonging to a high-frequency band, the human voice enhancement special effect is configured to increase an amplitude of a signal belonging to a human voice band, and the audio dynamic range compression special effect is configured to compress a dynamic range of a signal.

[0019] In a possible implementation, in response to the user selecting the seniority-specific audio mode from the audio modes, causing the audio of the target multimedia resource to present the special effect corresponding to the seniority-specific audio mode when the target multimedia resource is played includes: in response to the user selecting the seniority-specific audio mode from the audio modes, processing original audio of the target multimedia resource based on the special effect corresponding to the seniority-specific audio mode to obtain to-be-played audio; and using the to-be-played audio to replace the original audio when the target multimedia resource is played.

[0020] In a possible implementation, processing the original audio of the target multimedia resource based on the special effect corresponding to the seniority-specific audio mode to obtain the to-be-played audio includes: sequentially using one or more special effects corresponding to the seniority-specific audio mode to render on the basis of the original audio to obtain the to-be-played audio; wherein, when the seniority-specific audio mode corresponds to multiple special effects, the original audio is used as a signal to be rendered for a first special effect, a rendering result corresponding to a previous special effect is used as a signal to be rendered for a next special effect, and a rendering result corresponding to a last special effect is used as the to-be-played audio.

[0021] In a possible implementation, the rendering parameter of the high-frequency sound enhancement special effect includes a first threshold and a first coefficient, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the seniority-specific sound effect to obtain the audio to be played includes: when the high-frequency sound enhancement special effect is used to render the signal to be rendered, the signal to be rendered is divided into a first signal and a second signal, the first signal is a signal in a high-frequency band; the amplitude of the first signal is increased according to the first coefficient to obtain a third signal; the amplitudes of the third signal and the first threshold are compared, and when the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain a fourth signal; and the fourth signal and the second signal are combined as a rendering result corresponding to the high-frequency sound enhancement special effect.

[0022] In a possible implementation, the rendering parameter of the high-frequency sound enhancement special effect includes a first threshold and a first coefficient, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the seniority-specific sound effect to obtain the audio to be played includes: when the high-frequency sound enhancement special effect is used to render the signal to be rendered, the signal to be rendered is divided into a first signal and a second signal, the first signal is a signal in a high-frequency band; the amplitude of the first signal is increased according to the first coefficient to obtain a third signal; the amplitudes of the third signal and the first threshold are compared, and when the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain a fourth signal; and the fourth signal and the second signal are combined as a rendering result corresponding to the high-frequency sound enhancement special effect.

[0023] In a possible implementation, the rendering parameter of the audio dynamic range compression special effect includes a second threshold, a third threshold, a fourth coefficient and a fifth coefficient, the second threshold is greater than the third threshold, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the seniority-specific sound effect to obtain the audio to be played includes: when the audio dynamic range compression special effect is used to render the signal to be rendered, the amplitude of the signal to be rendered is compared with the second threshold and the third threshold; when the amplitude of the signal to be rendered is greater than the second threshold, the amplitude of the signal to be rendered is reduced according to the fourth coefficient; and when the amplitude of the signal to be rendered is less than the third threshold, the amplitude of the signal to be rendered is increased according to the fifth coefficient.

[0024] In a possible implementation, the apparatus further includes a first receiving module configured to receive a feedback signal from a user when the target multimedia resource is played, and an adjusting module configured to adjust the rendering parameter of at least one special effect corresponding to the seniority-specific sound effect according to the feedback signal.

[0025] In a possible implementation, the rendering parameter of any one of the special effects corresponding to the elder-specific sound effect is a rendering parameter pre-set on the current terminal device, or a rendering parameter pre-set on a user account logged in on the current terminal device.

[0026] In a possible implementation, the apparatus further includes: a first query module, configured to query an interface display mode of the current terminal device in response to the target multimedia resource starting to be played; a first judgment module, configured to judge whether a playing state of the target multimedia resource meets a first preset condition in response to the interface display mode of the current terminal device being an elder mode; a second query module, configured to query a playing record of a user on the current terminal device, or query user identity information recorded by the current terminal device in response to the interface display mode of the current terminal device being other than the elder mode; a second judgment module, configured to judge whether the playing state of the target multimedia resource meets the first preset condition according to a condition that the number of times that the user selects the elder-specific sound effect when playing a multimedia resource on the current terminal device is greater than a fourth threshold value, or according to a condition that the age of the user is greater than a fifth threshold value according to the user identity information; and a second audio processing module, configured to cause the audio of the target multimedia resource to present the special effect corresponding to the elder-specific sound effect when the target multimedia resource is played in response to the playing state of the target multimedia resource meeting the first preset condition.

[0027] In a possible implementation, the first preset condition includes one or more of the following: a time length of full-screen playing of the target multimedia resource exceeds a sixth threshold value; the audio played is original audio of the target multimedia resource; the target multimedia resource is played in full screen on the interface of the current terminal device; the target multimedia resource is in a playing state; and a time length of content of the target multimedia resource that has not been played is greater than a seventh threshold value.

[0028] In a possible implementation, the apparatus further includes: a second receiving module, configured to receive a sound effect mode selection signal from an associated device in a case that the current terminal device plays a target multimedia resource in a sound effect mode other than the elder-specific sound effect mode; and a third audio processing module, configured to cause the audio of the target multimedia resource to present the special effect corresponding to the elder-specific sound effect in response to the sound effect mode selection signal indicating the elder-specific sound effect.

[0029] In a possible implementation, the apparatus further includes: a storage module, configured to store the to-be-played audio; and a replacement module, configured to directly use the stored to-be-played audio to replace the original audio in a case that the user plays the target multimedia resource again and is in the sound effect mode of the elder-specific sound effect.

[0030] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0031] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium having stored thereon computer program instructions is provided, wherein the computer program instructions, when executed by a processor, implement the above method.

[0032] According to another aspect of the present disclosure, a computer program product is provided, comprising computer-readable code, or a non-transitory computer-readable storage medium carrying computer-readable code, which, when run in a processor of an electronic device, causes the processor in the electronic device to perform the above method.

[0033] According to the audio signal processing method of the embodiments of the present disclosure, by showing the sound effect mode supported by the target multimedia resource to the user, in response to the user selecting the long-elder exclusive sound effect in the sound effect mode, the audio of the target multimedia resource is made to present the special effect corresponding to the long-elder exclusive sound effect when playing the target multimedia resource, so that the user can actively select the sound effect mode, and the target multimedia resource can present the special effect corresponding to the long-elder exclusive sound effect; the special effect corresponding to the long-elder exclusive sound effect includes at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect, and an audio dynamic range compression special effect, so that the optimization mode of the long-elder exclusive sound effect for the audio is more diversified; the high-frequency sound enhancement special effect is used to enhance the amplitude of the signal belonging to the high-frequency band, so that the old user has a better listening experience for the signal of the high-frequency band. The human voice enhancement special effect is used to enhance the amplitude of the signal belonging to the human voice frequency band, so that the old user has a better listening experience for the signal of the human voice frequency band. The audio dynamic range compression special effect is used to compress the dynamic range of the signal, to reduce the adverse effects of the old user watching the multimedia resource on the health. In summary, the audio signal processing method of the present disclosure supports the long-elder exclusive sound effect customized for the old user, and makes the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect, which can improve the viewing experience of the old user for the multimedia resource and reduce the adverse effects of the old user watching the multimedia resource on the health.

[0034] Other features and aspects of the present disclosure will become apparent from the following detailed description of example embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0035] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate example embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.

[0036] Figure 1An example application scenario of the audio signal processing method according to an embodiment of the present disclosure is shown.

[0037] Figure 2 A schematic diagram showing a flow of the audio signal processing method according to an embodiment of the present disclosure is shown.

[0038] Figure 3 A schematic diagram showing processing of original audio of a target multimedia resource based on a plurality of special effects corresponding to a long elder exclusive sound effect according to an embodiment of the present disclosure is shown.

[0039] Figure 4 An example frequency band division manner according to an embodiment of the present disclosure is shown.

[0040] Figure 5 An example of amplitude change of a signal before and after compression using an audio dynamic range compression special effect according to an embodiment of the present disclosure is shown.

[0041] Figure 6 A schematic diagram showing the structure of an audio signal processing apparatus according to an embodiment of the present disclosure is shown.

[0042] Figure 7 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0043] Various exemplary embodiments, features and aspects of the present disclosure will be explained in detail below with reference to the accompanying drawings. The same reference numbers in different drawings denote the same or similar elements unless specified otherwise. Although various aspects of the embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.

[0044] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0045] In addition, for the purpose of better illustrating the present disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will appreciate that the present disclosure can be practiced without some or all of the specific details. In some instances, well-known methods, apparatuses, elements and circuits have not been described in detail in order to avoid unnecessarily obscuring the present disclosure.

[0046] The following first introduces the terms that can appear in this document.

[0047] Dynamic Range Compression (DRC): Dynamic range refers to the difference between the maximum amplitude and the minimum amplitude of a signal. In view of the transient characteristics of the audio signal, the amplitude of the signal with too high amplitude is compressed, and the amplitude of the signal with too low amplitude is increased, so that the maximum amplitude of the signal is reduced, the minimum amplitude of the signal is increased, and the difference between the maximum amplitude and the minimum amplitude of the signal is reduced, thereby achieving the effect of dynamic range compression.

[0048] Figure 1 An exemplary application scenario of the audio signal processing method according to an embodiment of the present disclosure is shown.

[0049] As Figure 1 shown, the application scenario can include a terminal device and a server. A user can view a playable multimedia resource on the terminal device and select a target multimedia resource to be watched. The terminal device can send an identifier of the target multimedia resource to the server.

[0050] The server includes an attribute storage module, a data generation module, and a multimedia service module. The attribute storage module is configured to store attributes of the multimedia resource, such as sound effects supported by the multimedia resource, supported languages, a playback address of the multimedia resource, and the like. The data generation module is configured to generate data of the multimedia resource in multiple formats, such as mp4 format, hls format, cmaf format, and the like. The audio of the multimedia resource generated by the data generation module is the original audio without added special effects.

[0051] After receiving the identifier of the target multimedia resource, the server sends a sound effect acquisition request to the attribute storage module via the multimedia service module and sends a data acquisition request to the data generation module. The sound effect acquisition request and the data acquisition request can include the identifier of the target multimedia resource.

[0052] In response to the sound effect acquisition request, the attribute storage module can query the sound effects supported by the target multimedia resource and return them to the multimedia service module. Exemplarily, the sound effects supported by the target multimedia resource can include one or more of texture sound effects, live live sound effects, sleep aid sound effects, theater sound effects, ultra-low bass sound effects, and senior exclusive sound effects. The present disclosure does not limit the number and types of sound effects supported by the target multimedia resource.

[0053] In response to the sound effect acquisition request, the data generation module can generate data of the target multimedia resource and return it to the multimedia service module. The format of the generated data can be pre-set or randomly selected, and the present disclosure does not limit this.

[0054] The multimedia service module can encapsulate the sound effect supported by the target multimedia resource and the data of the target multimedia resource, and send the identification of the target multimedia resource and the address of the encapsulated data packet to the terminal device. In the encapsulated data packet, the identification of the sound effect supported by the target multimedia resource can be included, so that the terminal device can perform subsequent rendering processing according to the sound effect indicated by the identification.

[0055] In this case, after the terminal device receives the identification of the target multimedia resource and the address of the encapsulated data packet, it can determine that the data packet stored at the address is a data packet obtained by encapsulating the sound effect supported by the target multimedia resource and the data of the target multimedia resource.

[0056] The terminal device can include a download module, a decoding module, and a rendering module. The download module downloads the data packet from the received address, and decodes the data packet through the decoding module to obtain the sound effect supported by the target multimedia resource and the data of the target multimedia resource. Then, the terminal device can use a physical player (not shown) included therein to start playing the data of the target multimedia resource, and start executing the audio signal processing method of the present disclosure to show the user the sound effect mode supported by the target multimedia resource. At this time, the played audio is the original audio. The audio signal processing method can be executed by any data processing unit (not shown) of the terminal device, which is not limited by the present disclosure.

[0057] The terminal device can communicate with the server through multiple interfaces (not shown), and different interfaces are used to transmit different information. For example, the terminal device can transmit information such as user starting to play the multimedia resource, user closing the multimedia resource, and audio of the multimedia resource being rendered to the server through different interfaces, and receive data packet addresses and data packets from the server.

[0058] In one example, the download module can download data packets in formats such as rtp, ykfile, hls, cmaf, cmaf-dolby, etc., and the decoding module can use decoders such as ffmpeg, omx, atb, etc. to implement decoding. The data packets in formats such as rtp, ykfile, hls, cmaf can be decoded using the ffmpeg decoder, and the data packets in the cmaf-dolby format can be decoded using the omx decoder or the atb decoder.

[0059] In one example, when the terminal device uses an Android system, the physical player can include opensl, audiotrack, and the like. When the terminal device uses an iOS system, the physical player can include audio queue, audio unit, and the like. When the terminal device uses a Windows system, the physical player can include direct sound, and the like. The present disclosure does not limit the type of system used by the terminal device and the type of physical player included by the terminal device.

[0060] When the user selects a certain sound effect mode, the terminal device can render the audio of the target multimedia resource through the rendering module. The special effect used in rendering corresponds to the sound effect mode selected by the user. Different sound effect modes correspond to different special effects, so the rendering processes corresponding to different sound effect modes are different. The terminal device can replace the original audio of the target multimedia resource with the rendered audio, and play the data of the multimedia resource after the audio is replaced using the physical player. In this case, the audio heard by the user is the audio of the sound effect mode selected by the user.

[0061] The above is only an example description, and the present application does not limit the specific division of the structures and functions implemented by the server and the terminal device.

[0062] Figure 2 A schematic diagram showing the flow of the audio signal processing method according to an embodiment of the present disclosure.

[0063] As Figure 2 shown, in one possible implementation, the method includes:

[0064] Step S21: showing the user the sound effect modes supported by the target multimedia resource;

[0065] Step S22: in response to the user selecting an elder-specific sound effect in the sound effect modes, causing the audio of the target multimedia resource to present the special effect corresponding to the elder-specific sound effect when playing the target multimedia resource;

[0066] The special effect corresponding to the elder-specific sound effect includes at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect, and an audio dynamic range compression special effect. The high-frequency sound enhancement special effect is used to increase the amplitude of signals belonging to the high-frequency band. The human voice enhancement special effect is used to increase the amplitude of signals belonging to the human voice band. The audio dynamic range compression special effect is used to compress the dynamic range of signals.

[0067] For example, the present disclosure supports the user actively selecting a sound effect mode. That is, the terminal device can first show the user the sound effect modes supported by the target multimedia resource, and then, according to the user's selection, cause the audio of the target multimedia resource to present the special effect corresponding to the selected sound effect when playing the target multimedia resource, thereby changing the audio presentation effect when playing the target multimedia resource.

[0068] Old users have poor sensitivity to high-frequency sounds. When the amplitude of signals in the high-frequency band is too low, the old users can not effectively obtain the information included in the signals, or even completely miss the signals. Therefore, the special effects corresponding to the senior exclusive sound effects can include a high-frequency sound enhancement effect. The high-frequency sound enhancement effect is used to increase the amplitude of signals in the high-frequency band. After the high-frequency sound enhancement effect is added, the amplitude of signals in the high-frequency band in the audio is higher, so that the old users can have a better listening experience for signals in the high-frequency band.

[0069] Due to hearing loss, it is more difficult for old users to identify signals with low volume. Since vocals carry most of the information in audio, if the amplitude of signals in the vocal frequency band is too low, the viewing experience of the old users will be greatly reduced. Therefore, the special effects corresponding to the senior exclusive sound effects can include a vocal enhancement effect. The vocal enhancement effect is used to increase the amplitude of signals in the vocal frequency band. After the vocal enhancement effect is added, the amplitude of signals in the vocal frequency band in the audio is higher, so that the old users have a better listening experience for signals in the vocal frequency band.

[0070] Due to poor adaptability, audio signals with too large a dynamic range can cause physical discomfort to old users. Therefore, the special effects corresponding to the senior exclusive sound effects can include an audio dynamic range compression effect. The audio dynamic range compression effect is used to compress the dynamic range of signals. After the audio dynamic range compression effect is added, the dynamic range of signals in all frequency bands of the audio is reduced, thereby reducing the adverse effects of watching multimedia resources on the health of old users.

[0071] In actual applications, the special effects corresponding to the senior exclusive sound effects can include at least one of the high-frequency sound enhancement effect, the vocal enhancement effect, and the audio dynamic range compression effect. Those skilled in the art should understand that the senior exclusive sound effects can also include more special effects suitable for old users, as long as the sound effects are more suitable for old users after the corresponding special effects are added. The present disclosure does not limit the types of special effects included in the senior exclusive sound effects.

[0072] According to the audio signal processing method provided in the embodiments of the present disclosure, by showing the user the sound effect modes supported by the target multimedia resource, in response to the user selecting the long-elder-specific sound effect in the sound effect modes, the audio of the target multimedia resource is caused to present the special effect corresponding to the long-elder-specific sound effect when playing the target multimedia resource, so that the user can actively select the sound effect mode, and the target multimedia resource can present the special effect corresponding to the long-elder-specific sound effect; the special effect corresponding to the long-elder-specific sound effect includes at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect, and an audio dynamic range compression special effect, so that the optimization mode of the long-elder-specific sound effect for the audio is more diversified; the high-frequency sound enhancement special effect is used to enhance the amplitude of the signal belonging to the high-frequency band, so that the old user has a better listening experience for the signal of the high-frequency band. The human voice enhancement special effect is used to enhance the amplitude of the signal belonging to the human voice band, so that the old user has a better listening experience for the signal of the human voice band. The audio dynamic range compression special effect is used to compress the dynamic range of the signal, so as to reduce the adverse effects of the old user watching the multimedia resource on the health. In summary, the audio signal processing method provided in the present disclosure supports the long-elder-specific sound effect customized for the old user, and causes the audio of the target multimedia resource to present the special effect corresponding to the long-elder-specific sound effect, which can improve the viewing experience of the old user for the multimedia resource and reduce the adverse effects of the old user watching the multimedia resource on the health.

[0073] In a possible implementation, in response to the user selecting the long-elder-specific sound effect in the sound effect modes, the audio of the target multimedia resource is caused to present the special effect corresponding to the long-elder-specific sound effect when playing the target multimedia resource, including:

[0074] In response to the user selecting the long-elder-specific sound effect in the sound effect modes, the original audio of the target multimedia resource is processed based on the special effect corresponding to the long-elder-specific sound effect to obtain the to-be-played audio;

[0075] When playing the target multimedia resource, the original audio is replaced with the to-be-played audio.

[0076] For example, the original audio of the multimedia resource is the audio without adding the special effect, therefore, when the user actively selects the long-elder-specific sound effect, the original audio can be taken as the signal processing object, and the original audio is processed based on the special effect corresponding to the long-elder-specific sound effect to obtain the to-be-played audio adding the special effect corresponding to the long-elder-specific sound effect. When playing the multimedia resource, the original audio can be replaced with the to-be-played audio, and the picture of the multimedia resource remains unchanged.

[0077] The process of processing the original audio of the target multimedia resource based on the special effect corresponding to the long-elder-specific sound effect can be performed in the terminal device, for example, in the rendering module shown in Figure 1 .

[0078] The following introduces an exemplary method of processing the original audio to obtain the to-be-played audio.

[0079] In one possible implementation, the original audio of the target multimedia resource is processed based on special effects corresponding to the elder's exclusive sound effects to obtain the audio to be played, including:

[0080] Based on the original audio, one or more special effects corresponding to the elder's exclusive sound effects are applied sequentially to render the audio to be played.

[0081] In the case of multiple special effects corresponding to the elder-specific sound effects, the original audio is used as the signal to be rendered for the first special effect, the rendering result of the previous special effect is used as the signal to be rendered for the next special effect, and the rendering result of the last special effect is used as the audio to be played.

[0082] For example, a sound effect specifically for elders might correspond to one or more special effects. By applying one or more of these special effects sequentially to the original audio, the audio to be played can be obtained.

[0083] In cases where the exclusive sound effect for elders corresponds to only one special effect, the original audio is used as the signal to be rendered for that special effect, and the rendering result of that special effect is used as the audio to be played.

[0084] When the elder-specific sound effect corresponds to multiple special effects, the original audio can be used as the signal to be rendered for the first special effect, the rendering result of the previous special effect can be used as the signal to be rendered for the next special effect, and the rendering result of the last special effect can be used as the audio to be played. Figure 3 This diagram illustrates the processing of the original audio of a target multimedia resource based on various special effects corresponding to the elder-specific sound effects according to an embodiment of this disclosure.

[0085] like Figure 3 As shown, when the elder-specific sound effects correspond to the high-frequency sound enhancement effect (Effect A), the vocal enhancement effect (Effect B), and the audio dynamic range compression effect (Effect C), the high-frequency sound enhancement effect can be used as the first effect, the vocal enhancement effect as the second effect, and the audio dynamic range compression effect as the last effect. The original audio can be used as the signal to be rendered for the high-frequency sound enhancement effect, the rendering result of the high-frequency sound enhancement effect can be used as the signal to be rendered for the vocal enhancement effect, the rendering result of the vocal enhancement effect can be used as the signal to be rendered for the audio dynamic range compression effect, and the rendering result of the audio dynamic range compression effect can be used as the audio to be played.

[0086] Those skilled in the art should understand that the human voice enhancement special effect or the audio dynamic range compression special effect can also be the first special effect, the high frequency sound enhancement special effect or the audio dynamic range compression special effect can also be the second special effect, and the high frequency sound enhancement special effect or the human voice enhancement special effect can also be the last special effect. If the rendering process corresponding to a special effect does not affect the rendering process corresponding to another special effect, the two special effects can also be rendered together. For example, assuming that the rendering process corresponding to the high frequency sound enhancement special effect does not affect the rendering process corresponding to the human voice enhancement special effect, the original audio can be rendered using the high frequency sound enhancement special effect and the human voice enhancement special effect, and the rendering result of the high frequency sound enhancement special effect and the human voice enhancement special effect can be rendered using the audio dynamic range compression special effect to obtain the to-be-played audio. As long as the to-be-played audio is rendered using all the special effects corresponding to the seniority exclusive mode, the specific rendering order of the special effects is not limited in the present disclosure.

[0087] The seniority exclusive sound effect relates to the processing of signals in the high frequency band and the human voice band, and therefore, accurate division of the frequency band is the key to ensuring the effect of the seniority exclusive sound effect. Figure 4 An exemplary frequency band division manner according to an embodiment of the present disclosure is shown.

[0088] As shown in Figure 4 , the frequency band of 31 Hz to 62 Hz can be divided into a sub-bass frequency band affecting the spatial sense of tone color, and the sound belonging to this frequency band can include the sound of instruments such as bass drum and bass, and thunder.

[0089] The frequency band of 62 Hz to 125 Hz can be divided into a bass frequency band affecting the fullness of sound, and the representative sound belonging to this frequency band can include the sound of instruments such as bass drum, cello, and trombone.

[0090] The frequency band of 125 Hz to 250 Hz can be divided into a bass frequency band affecting the strength of sound, and the representative sound belonging to this frequency band can include male human voice and nasal sound.

[0091] The frequency band of 250 Hz to 500 Hz can be divided into a medium-low frequency band affecting the thickness and strength of human voice, and the representative sound belonging to this frequency band can include human voice.

[0092] The frequency band of 1 KHz to 2 KHz can be divided into a medium frequency band having the most sensitive frequency of human ear and throat sound sense, and the representative sound belonging to this frequency band can include human voice and percussion.

[0093] The frequency band of 4 KHz to 8 KHz can be divided into a medium-high frequency band having penetrating power, hierarchical influence, brightness, and clarity, and the representative sound belonging to this frequency band can include the sound of wind instruments.

[0094] The frequency band of 8KHz-16KHz can be divided into a high frequency band affecting the sound level and improving the background noise, and representative sounds belonging to this frequency band can include the sounds of flutes and trumpets.

[0095] The frequency band of 16KHz or above can be divided into a super high frequency band affecting the overall color of the sound, and representative sounds belonging to this frequency band can include the sounds of cymbals and triangles.

[0096] In the example shown in Figure 4 , the frequency band of 125Hz-500Hz and the frequency band of 1KHz-2KHz can be a human voice frequency band, and the frequency band of 8KHz or above can be a high frequency band. It should be understood by those skilled in the art that in actual applications, the frequency band can also have more division methods, and the present disclosure does not limit the division method of the frequency band.

[0097] The following will introduce an exemplary rendering process of the high frequency sound enhancement special effect in combination with the frequency band division method shown in Figure 4 .

[0098] In one possible implementation, the rendering parameter of the high frequency sound enhancement special effect includes a first threshold and a first coefficient,

[0099] The original audio of the target multimedia resource is processed based on the special effect corresponding to the long-elders-only sound effect to obtain the to-be-played audio, including:

[0100] When the high frequency sound enhancement special effect is used to render the to-be-rendered signal,

[0101] The to-be-rendered signal is divided into a first signal and a second signal, and the first signal is a signal of a high frequency band;

[0102] The amplitude of the first signal is increased according to the first coefficient to obtain a third signal;

[0103] The amplitudes of the third signal and the first threshold are compared, and when the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain a fourth signal;

[0104] The fourth signal and the second signal are merged as a rendering result corresponding to the high frequency sound enhancement special effect.

[0105] For example, the rendering parameters of different special effects are different. For example, the rendering parameters of the high frequency sound enhancement special effect can include a first threshold and a first coefficient. The first threshold determines the upper limit of the amplitude of the signal in the high frequency band, and the first coefficient determines the enhancement degree of the signal in the high frequency band. The first threshold and the first coefficient can be set and adjusted according to the application scene requirements, and the specific values of the first threshold and the first coefficient are not limited in the present disclosure. The high frequency sound enhancement special effect can also include more rendering parameters, and the specific rendering parameters used by the high frequency sound enhancement special effect are not limited in the present disclosure.

[0106] Since the high frequency sound enhancement special effect is used to enhance the amplitude of the signal in the high frequency band, and the signal to be rendered can also include signals that do not belong to the high frequency band, such as signals in the low frequency band and the intermediate frequency band, the signal to be rendered can be divided into a first signal and a second signal first. The first signal is the signal in the high frequency band, and the amplitude of the first signal can be enhanced according to the first coefficient to obtain a third signal. This process can be implemented based on existing technology, and will not be described here. The second signal is a signal that does not belong to the high frequency band, and the amplitude of the second signal does not have to be enhanced in the rendering process of the high frequency sound enhancement special effect.

[0107] The value of the first coefficient can be adjusted according to the amplitude of the first signal. For example, when the amplitude of the first signal is already relatively large, the value of the first coefficient can be small; when the amplitude of the first signal is small, the value of the first coefficient can be large. In this way, the amplitude of the enhanced high frequency band signal is more stable.

[0108] If the amplitude of the first signal is large itself, after enhancing the amplitude of the first signal according to the first coefficient, the third signal obtained may be distorted due to the excessive amplitude. To avoid this, the amplitude of the third signal can be compared with the first threshold, and when the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain a fourth signal. Therefore, the upper limit of the amplitude of the fourth signal is the first threshold.

[0109] Finally, the fourth signal and the second signal can be combined as the rendering result corresponding to the high frequency sound enhancement special effect. The fourth signal and the second signal are frequency domain signals, so the combination of the fourth signal and the second signal does not change the amplitude of the fourth signal, that is, the upper limit of the amplitude of the signal in the high frequency band in the combined signal is still the first threshold. The combination of signals can be implemented based on existing technology, and will not be described here.

[0110] In this way, high-frequency signals can be enhanced. This disclosure addresses the common high-frequency hearing loss among the elderly by performing frequency analysis on the signal to be rendered and enhancing the high-frequency signals, making high-frequency sounds clearer and more audible. This allows the elderly to clearly hear everyday reminders, ringtones, and details in speech, improving their hearing perception. Simultaneously, intelligent automatic adjustment of high-frequency gain ensures better audio performance in different environments, allowing the elderly to better perceive consonants in words, thereby improving speech recognition, enhancing communication effectiveness, and improving communication quality. Fine-tuning is applied to prevent distortion that may occur during high-frequency sound enhancement, maintaining the naturalness and purity of the sound quality, making the enhanced high-frequency signal more comfortable to listen to and providing a natural listening experience.

[0111] Those skilled in the art should understand that the rendering of high-frequency sound enhancement effects can also be achieved through other methods of existing technology, as long as the amplitude of signals belonging to the high-frequency band can be increased. This disclosure does not limit the specific rendering method corresponding to high-frequency sound enhancement effects.

[0112] The following is combined Figure 4 The frequency band division method shown illustrates an exemplary rendering process corresponding to the human voice enhancement effect in this disclosure.

[0113] In one possible implementation, the rendering parameters for the voice enhancement effect include a second coefficient and a third coefficient.

[0114] Based on the special effects corresponding to the elder's exclusive sound effects, the original audio of the target multimedia resource is processed to obtain the audio to be played, including:

[0115] When rendering a signal using voice enhancement effects.

[0116] The signal to be rendered is divided into the fifth signal and the sixth signal. The fifth signal is the signal in the human voice frequency band.

[0117] The amplitude of the fifth signal is increased based on the second coefficient to obtain the seventh signal;

[0118] The eighth signal is obtained by reducing the amplitude of the noise signal in the sixth signal using the third coefficient.

[0119] The seventh and eighth signals are merged to form the rendering result corresponding to the voice enhancement effect.

[0120] For example, the rendering parameters of the vocal enhancement special effect can include a second coefficient and a third coefficient. The second coefficient determines the degree of enhancement of the signal in the vocal frequency band, and the third coefficient determines the degree of noise suppression. The second coefficient and the third coefficient can be set and adjusted according to the requirements of the application scenario, and the specific values of the second coefficient and the third coefficient are not limited in the present disclosure. The vocal enhancement special effect can also include more rendering parameters, and the specific rendering parameters used by the vocal enhancement special effect are not limited in the present disclosure.

[0121] Since the vocal special effect is used to enhance the amplitude of the signal in the vocal frequency band, and the signal to be rendered can also include signals not in the vocal frequency band, the signal to be rendered can be divided into a fifth signal and a sixth signal first. The fifth signal is the signal in the vocal frequency band, and the amplitude of the fifth signal can be enhanced according to the second coefficient to obtain a seventh signal. The sixth signal is a signal not in the vocal frequency band and includes a noise signal, and thus the amplitude of the noise signal in the sixth signal can be reduced according to the third coefficient to obtain an eighth signal. Finally, the seventh signal and the eighth signal can be combined as the rendering result corresponding to the vocal enhancement special effect. The enhancement, reduction, and combination of signal amplitudes can be implemented based on existing technologies, and thus will not be described here.

[0122] The value of the second coefficient can be adjusted according to the amplitude of the fifth signal. The exemplary adjustment manner is the same as the adjustment manner of the value of the first coefficient described above, and thus will not be described here. The process of extracting the noise signal from the sixth signal and reducing the amplitude of the noise signal can be implemented based on related technologies. The third coefficient can determine the degree of reduction of the amplitude of the noise signal.

[0123] Alternatively, the rendering parameters of the vocal enhancement special effect can also include an eighth threshold value for determining the upper limit of the amplitude of the vocal signal. If the amplitude of the fifth signal is large, the seventh signal obtained after the amplitude of the fifth signal is enhanced according to the second coefficient can be distorted due to the excessively large amplitude. To avoid this, the amplitude of the seventh signal can be compared with the eighth threshold value, and when the amplitude of the seventh signal exceeds the eighth threshold value, the amplitude of the seventh signal is reduced to the eighth threshold value to ensure that the upper limit of the amplitude of the seventh signal is the eighth threshold value. The value of the eighth threshold value can be set according to the requirements of the application scenario, and can be the same as or different from the first threshold value. The specific value of the eighth threshold value is not limited in the present disclosure. The seventh signal and the eighth signal are signals in the frequency domain, and thus the combination of the seventh signal and the eighth signal will not change the amplitude of the seventh signal, that is, the upper limit of the amplitude of the signal in the vocal frequency band in the combined signal is still the eighth threshold value.

[0124] In this way, the signal of the human voice frequency band can be enhanced. The method of the present disclosure optimizes the human voice frequency band, making the audio signal of the human voice part clearer and more prominent, improving the clarity and recognition of the dialogue, and enabling the elderly user to communicate more freely. At the same time, the background noise suppression technology is applied to reduce the amount of background noise and improve the signal-to-noise ratio of the speech signal. The optimization of the human voice reduces the fatigue of needing to concentrate to listen, making long-term communication more relaxed and enjoyable. The human voice enhancement technology significantly improves the speech clarity in noisy environments, enabling confident listening and participation in conversations in public places or family gatherings, and improving the adaptability of the method to various application scenarios.

[0125] Those skilled in the art should understand that the rendering corresponding to the human voice enhancement special effect can also be implemented in other ways by prior art, for example, a timbre recognition step can also be added to distinguish human voice signals of different roles and enhance the amplitude of the human voice signal of the corresponding role according to the user's selected role. As long as the amplitude of the signal belonging to the human voice frequency band can be improved, the present disclosure does not limit the specific rendering mode of the human voice enhancement special effect.

[0126] The following describes an exemplary rendering process corresponding to the audio dynamic range compression special effect in the present disclosure.

[0127] In one possible implementation, the rendering parameters of the audio dynamic range compression special effect include a second threshold, a third threshold, a fourth coefficient, and a fifth coefficient, the second threshold is greater than the third threshold,

[0128] The original audio of the target multimedia resource is processed based on the special effect corresponding to the long-elder exclusive sound effect to obtain the to-be-played audio, including:

[0129] When the audio dynamic range compression special effect is used to render the to-be-rendered signal,

[0130] The amplitudes of the to-be-rendered signal and the second threshold and the third threshold are compared;

[0131] When the amplitude of the to-be-rendered signal is greater than the second threshold, the amplitude of the to-be-rendered signal is reduced according to the fourth coefficient;

[0132] When the amplitude of the to-be-rendered signal is less than the third threshold, the amplitude of the to-be-rendered signal is increased according to the fifth coefficient.

[0133] For example, the rendering parameters for audio dynamic range compression effects may include a second threshold, a third threshold, a fourth coefficient, and a fifth coefficient. The second threshold determines when the amplitude of the signal to be rendered needs to be reduced, the fourth coefficient determines the degree of amplitude reduction, the third threshold determines when the amplitude of the signal to be rendered needs to be increased, and the fifth coefficient determines the degree of amplitude increase. The second threshold, third threshold, fourth coefficient, and fifth coefficient can be set and adjusted according to the application scenario requirements, and this disclosure does not limit the specific values ​​of the second threshold, third threshold, fourth coefficient, and fifth coefficient. Audio dynamic range compression effects may also include more rendering parameters, such as attack time (how long it takes for dynamic range compression to take effect) and release time (how long it takes for dynamic range compression to recover), etc., and this disclosure does not limit the specific rendering parameters used for audio dynamic range compression effects.

[0134] For example, different rendering parameters can be preset for different sound sources to adaptively achieve dynamic range compression, ensuring that different sound sources can be dynamically adjusted appropriately, making volume changes smoother and more natural.

[0135] Since audio dynamic range compression effects are used to compress the dynamic range of a signal, and the essence of dynamic range compression is to reduce the amplitude of signals with excessively high amplitudes and increase the amplitude of signals with excessively low amplitudes, we can compare the amplitude of the signal to be rendered with the second and third thresholds. If the amplitude of the signal to be rendered is greater than the second threshold, it is considered that the amplitude of the signal to be rendered is too high and needs to be reduced, which can be done according to a fourth coefficient. If the amplitude of the signal to be rendered is less than the third threshold, it is considered that the amplitude of the signal to be rendered is too low and needs to be increased, which can be done according to a fifth coefficient. The reduction and increase of signal amplitude can be achieved based on existing technologies, which will not be elaborated here.

[0136] Figure 5 An example is shown of the amplitude change of the signal before and after compression using an audio dynamic range compression effect according to an embodiment of this disclosure.

[0137] like Figure 5 As shown, when the signal amplitude is greater than the second threshold, the signal amplitude begins to decrease. The degree of decrease is proportional to the difference between the signal amplitude and the second threshold, and the ratio can be equal to the fourth coefficient. When the signal amplitude is less than the third threshold, the signal amplitude begins to increase. The degree of increase is proportional to the difference between the third threshold and the signal amplitude, and the ratio can be equal to the fifth coefficient.

[0138] In this way, the dynamic range of the signal can be compressed. The method of the present disclosure intelligently adjusts the dynamic range of the audio signal according to the transient characteristics of the audio signal, and compresses the excessively high volume and raises the excessively low volume to achieve uniform processing of the sound. By finely controlling the rendering parameters of the audio dynamic range compression effect, better dynamic compression effects can be achieved while ensuring sound quality. Dynamic range compression effectively prevents discomfort caused by sudden volume changes in audio, allowing the elderly to comfortably listen to various audio content. At the same time, it avoids the impact of sudden high volume on hearing, which is particularly suitable for the sensitive hearing needs of the elderly and provides a safer listening environment. After dynamic compression, the volume of the audio is more balanced, allowing the audio signal to achieve stable volume output without affecting the sound quality, improving overall auditory comfort and experience.

[0139] Those skilled in the art should understand that the rendering corresponding to the audio dynamic range compression effect can also be implemented in other ways by existing technologies, as long as the dynamic range of the compressible signal is compressed. The present disclosure does not limit the specific rendering mode of the audio dynamic range compression effect.

[0140] The audio signal processing method of the embodiments of the present disclosure can also be used with existing audio quality detection algorithms. Existing audio quality detection algorithms can detect the quality of the audio to be played in multiple dimensions, such as audio-visual synchronization detection, audio volume detection, physical player capability detection, and physical player output audio smoothness detection. When the detection is unqualified, the rendering parameters of the effect, the physical player used, the data processing unit executing the audio signal processing method, etc. can be automatically adjusted.

[0141] For example, when it is detected that the volume of the audio to be played is low, the first threshold, the first coefficient, and the second coefficient described above can be increased, or other physical players can be replaced. When it is detected that the audio to be played is not smooth enough, the second threshold described above can be reduced, the third threshold described above can be increased, or other physical players can be replaced. The present disclosure does not limit the specific audio quality detection method and the adjustment method when the audio quality is unqualified.

[0142] In one possible implementation, the method further includes:

[0143] When playing the target multimedia resource, a feedback signal from the user is received;

[0144] According to the feedback signal, the rendering parameters of at least one effect corresponding to the senior exclusive sound effect are adjusted.

[0145] For example, the disclosure also supports adjusting the rendering parameter according to the feedback of the user. When playing the target multimedia resource, the terminal device can further receive a feedback signal from the user, which can indicate that the noise is too large / small, the voice is too high / low, etc. The terminal device can adjust the rendering parameter of the at least one special effect corresponding to the seniority-specific sound effect according to the feedback signal. For example, when the feedback signal indicates that the voice is too low, the second coefficient described above can be increased.

[0146] In this way, the audio playing effect of the multimedia resource can be more adapted to the needs of the user.

[0147] In a possible implementation, the rendering parameter of any one special effect corresponding to the seniority-specific sound effect is a rendering parameter pre-set on the current terminal device or a rendering parameter pre-set on a user account logged in on the current terminal device.

[0148] For example, the user can set the rendering parameter of the special effect corresponding to the seniority-specific sound effect on the current terminal device, and the user account logged in on the current terminal device can also pre-set the rendering parameter of the special effect corresponding to the seniority-specific sound effect. In this case, when the rendering of any one special effect corresponding to the seniority-specific sound effect is performed, the pre-set rendering parameter can be directly used.

[0149] In this way, the audio playing effect of the multimedia resource can be more adapted to the habits of the user.

[0150] In a possible implementation, the method further includes:

[0151] In response to the target multimedia resource starting to be played, querying an interface display mode of the current terminal device;

[0152] In response to the interface display mode of the current terminal device being the seniority mode, determining whether a playing state of the target multimedia resource meets a first preset condition;

[0153] In response to the interface display mode of the current terminal device being other than the seniority mode, querying a playing record of the user on the current terminal device, or querying user identity information recorded by the current terminal device;

[0154] In a case where it is determined according to the playing record that the number of times that the user selects the seniority-specific sound effect when playing the multimedia resource on the current terminal device is greater than a fourth threshold value, or in a case where it is determined according to the user identity information that the age of the user is greater than a fifth threshold value, determining whether the playing state of the target multimedia resource meets the first preset condition;

[0155] In response to the playing state of the target multimedia resource meeting the first preset condition, causing the audio of the target multimedia resource to present the special effect corresponding to the seniority-specific sound effect when the target multimedia resource is played.

[0156] For example, the terminal device determines whether the user is an elderly user. When it is determined that the user is an elderly user and the playing state of the target multimedia resource meets certain conditions, the terminal device automatically causes the audio of the target multimedia resource to present the special effect corresponding to the special sound effect for the elderly.

[0157] For example, the interface display mode of some terminal devices includes an elderly mode. When the interface display mode of the terminal device is already in the elderly mode, it can be directly determined that the user is an elderly user.

[0158] When the interface display mode of the terminal device is other than the elderly mode, whether the user is an elderly user can also be determined according to the playing record of the user on the current terminal device or according to the user identity information recorded by the current terminal device. For example, a fourth threshold value and a fifth threshold value can be preset. When it is determined according to the playing record that the number of times that the user selects the special sound effect for the elderly when playing multimedia resources on the current terminal device is greater than the fourth threshold value, it can be determined that the user is an elderly user. When it is determined according to the user identity information that the age of the user is greater than the fifth threshold value, it can be determined that the user is an elderly user. The values of the fourth threshold value and the fifth threshold value can be set according to the requirements of the application scenario, and the disclosure does not limit the specific values of the fourth threshold value and the fifth threshold value.

[0159] After it is determined that the user is an elderly user, it can also be determined whether the playing state of the target multimedia resource meets a first preset condition. When the playing state of the target multimedia resource meets the first preset condition, it can be considered that causing the audio of the target multimedia resource to present the special effect corresponding to the special sound effect for the elderly has high value, and in this case, the audio of the target multimedia resource is caused to present the special effect corresponding to the special sound effect for the elderly.

[0160] An example of the first preset condition is given below.

[0161] In one possible implementation, the first preset condition includes one or more of the following:

[0162] The duration of full-screen playing of the target multimedia resource exceeds a sixth threshold value;

[0163] The played audio is the original audio of the target multimedia resource;

[0164] The target multimedia resource is played in full screen on the interface of the current terminal device;

[0165] The target multimedia resource is in a playing state;

[0166] The duration of the content in the target multimedia resource that has not been played is greater than a seventh threshold value.

[0167] For example, the first preset condition can include that a duration of full-screen playing of the target multimedia resource exceeds a sixth threshold value. When this condition is met, it can be considered that the user has a greater probability of continuing to play the target multimedia resource, and it has higher value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect. The sixth threshold value can be set according to the application scenario requirement, for example, set to 10 seconds, and the specific value of the sixth threshold value is not limited in the present disclosure.

[0168] The first preset condition can also include that the played audio is the original audio of the target multimedia resource. If the played audio is not the original audio of the target multimedia resource, but the audio to which the special effect has been added, it can be considered that the user actively selects other sound effect modes, and the user's selection should be given priority. At this time, it has no value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect. Therefore, when the condition that the played audio is the original audio of the target multimedia resource is met, it has higher value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect.

[0169] The first preset condition can also include that the target multimedia resource is played full screen on the interface of the current terminal device. When the user does not play the target multimedia resource full screen, it can be considered that the user has a greater probability of switching to other multimedia resources. At this time, it has lower value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect. Therefore, when the condition that the target multimedia resource is played full screen on the interface of the current terminal device is met, it has higher value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect.

[0170] The first preset condition can also include that the target multimedia resource is in a playing state. When the target multimedia resource is in a paused state, even if the special effect is added to the original audio, it will not be played immediately. At this time, it has lower value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect. Therefore, when the condition that the target multimedia resource is in a playing state is met, it has higher value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect.

[0171] The first preset condition can also include that a duration of content in the target multimedia resource that has not been played is greater than a seventh threshold value. If the target multimedia resource is about to end soon, the duration of the audio to be rendered is too short. At this time, it has lower value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect. Therefore, when the condition that the duration of the content in the target multimedia resource that has not been played is greater than the seventh threshold value is met, it has higher value to make the audio of the target multimedia resource present the special effect corresponding to the long-elder exclusive sound effect. The seventh threshold value can be set according to the application scenario requirement, for example, set to 5 minutes, and the specific value of the seventh threshold value is not limited in the present disclosure.

[0172] Those skilled in the art shall understand that the first preset condition can further include more or less conditions, for example, the target multimedia resource is played at the top layer of the interface, a new function guide component of the player is not started, and the like, as long as the playing state of the target multimedia resource meets the first preset condition, it has high value to make the audio of the target multimedia resource present the special effect corresponding to the special sound effect for the elderly, and the disclosure does not limit the specific content of the first preset condition.

[0173] The terminal device can be a smart device, and it can be difficult for the elderly user to operate the smart device to actively select the special sound effect for the elderly. Based on this, the disclosure also supports remotely selecting the sound effect mode by the associated device of the terminal device when the terminal device plays the target multimedia resource.

[0174] In a possible implementation, in the case where the current terminal device plays the target multimedia resource in a sound effect mode other than the special sound effect for the elderly, the method further includes:

[0175] receiving a sound effect mode selection signal from the associated device, the sound effect mode selection signal indicating the sound effect mode;

[0176] in the case where the sound effect mode selection signal indicates the special sound effect for the elderly, making the audio of the target multimedia resource present the special effect corresponding to the special sound effect for the elderly.

[0177] For example, the current terminal device can have at least one associated device. In the case where the current terminal device plays the target multimedia resource in a sound effect mode other than the special sound effect for the elderly, if the terminal device receives a sound effect mode selection signal from the associated device, and the sound effect mode selection signal indicates the sound effect mode, it can be considered that the associated device remotely selects the sound effect mode of the target multimedia resource on the current terminal device. In this case, the terminal device can make the audio of the target multimedia resource present the special effect corresponding to the special sound effect for the elderly.

[0178] Further, a verification mechanism can be added to exclude the case where the associated device mistakenly sends the sound effect mode selection signal. For example, when the current terminal device receives the sound effect mode selection signal from the associated device, the user of the current terminal device can be prompted through a pop-up window or the like, and the user is waited for confirmation on whether to make the audio of the target multimedia resource present the special effect corresponding to the special sound effect for the elderly, and when the confirmation of the user is received within a preset time period, the audio of the target multimedia resource is made to present the special effect corresponding to the special sound effect for the elderly. When the confirmation of the user is not received within the preset time period, or the information that the user refuses to make the audio of the target multimedia resource present the special effect corresponding to the special sound effect for the elderly is received, the target multimedia resource is still played in the current sound effect mode.

[0179] In this way, the use experience of the elderly user can be improved.

[0180] In a possible implementation, the method further includes:

[0181] storing the to-be-played audio, and directly using the stored to-be-played audio to replace the original audio when the user plays the target multimedia resource again and is in the sound effect mode of the sound effect special for the elderly.

[0182] For example, the user may play the same multimedia resource multiple times. If the audio is re-rendered every time it is played, the data processing cost of the terminal device will be increased. Therefore, after the user plays the target multimedia resource for the first time, the to-be-played audio obtained by rendering can be stored. When the user plays the target multimedia audio again and the user selects the sound effect special for the elderly, or the terminal device determines that the user is an elderly user and the playing condition of the target multimedia resource meets the first preset condition, the terminal device can directly use the stored to-be-played audio to replace the original audio.

[0183] The selection of the sound effect special for the elderly can also be bound to the user account. That is, after the user account logs in the terminal device (for example, logs in an application for playing the target multimedia resource), in response to the sound effect mode corresponding to the user account being set as the sound effect special for the elderly, the audio of the target multimedia resource is caused to present the special effect corresponding to the sound effect special for the elderly when the target multimedia resource is played. In this way, the user can log in the application (for example, a video APP) for playing the target multimedia resource through the user account, set the sound effect mode as the sound effect special for the elderly, and thus, all multimedia resources can be played in the sound effect mode of the sound effect special for the elderly during the login period of the user account, so as to further simplify the operation of the user and provide higher use convenience for the elderly user.

[0184] In this way, the data processing cost of the terminal device can be reduced.

[0185] The present disclosure also provides an audio signal processing apparatus, Figure 6 FIG. 1 shows a schematic diagram of the structure of an audio signal processing apparatus according to an embodiment of the present disclosure.

[0186] As Figure 6 shown, in a possible implementation, the apparatus includes:

[0187] The display module 61 is configured to display the sound effect modes supported by the target multimedia resource to the user.

[0188] The first audio processing module 62 is configured to, in response to the user selecting the sound effect special for the elderly from the sound effect modes, cause the audio of the target multimedia resource to present the special effect corresponding to the sound effect special for the elderly when the target multimedia resource is played.

[0189] The special effect corresponding to the senior-specific sound effect includes at least one of a high-frequency sound enhancement special effect, a human voice enhancement special effect, and an audio dynamic range compression special effect. The high-frequency sound enhancement special effect is used to increase the amplitude of a signal in a high-frequency band. The human voice enhancement special effect is used to increase the amplitude of a signal in a human voice band. The audio dynamic range compression special effect is used to compress the dynamic range of a signal.

[0190] In a possible implementation, in response to the user selecting the senior-specific sound effect in the sound effect mode, the audio of the target multimedia resource is caused to present the special effect corresponding to the senior-specific sound effect when the target multimedia resource is played, including: in response to the user selecting the senior-specific sound effect in the sound effect mode, processing the original audio of the target multimedia resource based on the special effect corresponding to the senior-specific sound effect to obtain to-be-played audio; and using the to-be-played audio to replace the original audio when the target multimedia resource is played.

[0191] In a possible implementation, the processing of the original audio of the target multimedia resource based on the special effect corresponding to the senior-specific sound effect to obtain to-be-played audio includes: sequentially using one or more special effects corresponding to the senior-specific sound effect to render on the basis of the original audio to obtain the to-be-played audio; and when the senior-specific sound effect corresponds to multiple special effects, the original audio is used as a signal to be rendered for a first special effect, a rendering result corresponding to a previous special effect is used as a signal to be rendered for a next special effect, and a rendering result corresponding to a last special effect is used as the to-be-played audio.

[0192] In a possible implementation, the rendering parameter of the high-frequency sound enhancement special effect includes a first threshold and a first coefficient, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the senior-specific sound effect to obtain to-be-played audio includes: when the high-frequency sound enhancement special effect is used to render a signal to be rendered, the signal to be rendered is divided into a first signal and a second signal, the first signal being a signal in a high-frequency band; the amplitude of the first signal is increased according to the first coefficient to obtain a third signal; the amplitudes of the third signal and the first threshold are compared, and when the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain a fourth signal; and the fourth signal and the second signal are merged as a rendering result corresponding to the high-frequency sound enhancement special effect.

[0193] In a possible implementation, the rendering parameter of the human voice enhancement special effect includes a second coefficient and a third coefficient, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the seniority special sound effect to obtain the audio to be played includes: when the human voice enhancement special effect is used to render the signal to be rendered, the signal to be rendered is divided into a fifth signal and a sixth signal, the fifth signal is a signal of a human voice frequency band; the amplitude of the fifth signal is raised according to the second coefficient to obtain a seventh signal; the amplitude of a noise signal in the sixth signal is reduced according to the third coefficient to obtain an eighth signal; and the seventh signal and the eighth signal are combined as a rendering result corresponding to the human voice enhancement special effect.

[0194] In a possible implementation, the rendering parameter of the audio dynamic range compression special effect includes a second threshold, a third threshold, a fourth coefficient, and a fifth coefficient, the second threshold is greater than the third threshold, and the processing of the original audio of the target multimedia resource based on the special effect corresponding to the seniority special sound effect to obtain the audio to be played includes: when the audio dynamic range compression special effect is used to render the signal to be rendered, the amplitudes of the signal to be rendered and the second threshold and the third threshold are compared; when the amplitude of the signal to be rendered is greater than the second threshold, the amplitude of the signal to be rendered is reduced according to the fourth coefficient; and when the amplitude of the signal to be rendered is less than the third threshold, the amplitude of the signal to be rendered is raised according to the fifth coefficient.

[0195] In a possible implementation, the apparatus further includes a first receiving module configured to receive a feedback signal from a user when the target multimedia resource is played, and an adjusting module configured to adjust the rendering parameter of at least one special effect corresponding to the seniority special sound effect according to the feedback signal.

[0196] In a possible implementation, the rendering parameter of any special effect corresponding to the seniority special sound effect is a rendering parameter pre-set on a current terminal device or a rendering parameter pre-set on a user account logged in on the current terminal device.

[0197] In a possible implementation, the apparatus further includes: a first query module, configured to query an interface display mode of the current terminal device in response to the target multimedia resource starting to play; a first judgment module, configured to judge whether the playing state of the target multimedia resource meets a first preset condition in response to the interface display mode of the current terminal device being an elder mode; a second query module, configured to query a playing record of the user on the current terminal device or query user identity information recorded by the current terminal device in response to the interface display mode of the current terminal device being other modes than the elder mode; a second judgment module, configured to judge whether the playing state of the target multimedia resource meets the first preset condition according to a condition that the number of times that the user selects the elder exclusive sound effect when playing a multimedia resource on the current terminal device is greater than a fourth threshold value according to the playing record or according to a condition that the age of the user is greater than a fifth threshold value according to the user identity information; and a second audio processing module, configured to cause the audio of the target multimedia resource to present a special effect corresponding to the elder exclusive sound effect when playing the target multimedia resource in response to the playing state of the target multimedia resource meeting the first preset condition.

[0198] In a possible implementation, the first preset condition includes one or more of the following: the duration of full-screen playing of the target multimedia resource exceeds a sixth threshold value; the audio played is the original audio of the target multimedia resource; the target multimedia resource is played in full screen on the interface of the current terminal device; the target multimedia resource is in a playing state; and the duration of content in the target multimedia resource that has not been played is greater than a seventh threshold value.

[0199] In a possible implementation, the apparatus further includes: a second receiving module, configured to receive an audio mode selection signal from an associated device in a case that the current terminal device plays a target multimedia resource in an audio mode other than the elder exclusive sound effect, the audio mode selection signal indicating an audio mode; and a third audio processing module, configured to cause the audio of the target multimedia resource to present a special effect corresponding to the elder exclusive sound effect in response to the audio mode selection signal indicating the elder exclusive sound effect.

[0200] In a possible implementation, the apparatus further includes: a storage module, configured to store the to-be-played audio; and a replacement module, configured to directly use the stored to-be-played audio to replace the original audio in a case that the user plays a target multimedia resource again and is in the audio mode of the elder exclusive sound effect.

[0201] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.

[0202] The embodiments of the present disclosure further provide a computer readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the method described above. The computer readable storage medium can be a volatile or non-volatile computer readable storage medium.

[0203] The embodiments of the present disclosure further provide an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the method described above when executing the instructions stored in the memory.

[0204] The embodiments of the present disclosure further provide a computer program product, comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in the processor of an electronic device, the processor in the electronic device executes the method described above.

[0205] Figure 7 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the method described above.

[0206] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input output interface 1958 (I / O interface). The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0207] In exemplary embodiments, a non-volatile computer readable storage medium is also provided, such as the memory 1932 including computer program instructions executable by the processing component 1922 of the electronic device 1900 to complete the method described above.

[0208] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0209] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0210] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0211] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0212] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0213] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0214] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0215] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0216] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of terms is intended to best describe the principles of the embodiments, practical application, or technical improvements in the art, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An audio signal processing method, characterized in that, The method is applied to a terminal device, the terminal device including a download module, a decoding module, and a physical player, and the method includes: Send the identifier of the target multimedia resource selected by the user to the server; The address to receive the encapsulated data packet; Use the download module to download the data packet from the received address; The decoding module decodes the data packet to obtain the sound effect mode supported by the target multimedia resource and the data of the target multimedia resource; The target multimedia resource is started playing using the physical player, and the sound effect modes supported by the target multimedia resource are displayed to the user. In response to the user selecting the senior-exclusive sound effect in the sound effect mode, the original audio of the target multimedia resource is processed based on the special effect corresponding to the senior-exclusive sound effect to obtain the audio to be played; when playing the target multimedia resource, the original audio is replaced with the audio to be played so that the audio of the target multimedia resource presents the special effect corresponding to the senior-exclusive sound effect. In response to the current terminal device's interface display mode being set to "Elderly Mode," or in response to the current terminal device's interface display mode being set to another mode besides "Elderly Mode," and if, based on the user's playback history on the current terminal device, the number of times the user selects the "Elderly Exclusive Sound Effect" when playing multimedia resources on the current terminal device exceeds a fourth threshold, or if, based on the user's identity information recorded on the current terminal device, the user's age exceeds a fifth threshold, it is determined whether the playback status of the target multimedia resource meets a first preset condition. In response to the target multimedia resource's playback status meeting the first preset condition, when playing the target multimedia resource, the audio of the target multimedia resource is made to display the special effects corresponding to the "Elderly Exclusive Sound Effect." The special effects corresponding to the exclusive sound effects for elders include high-frequency sound enhancement effects, human voice enhancement effects, and audio dynamic range compression effects. The high-frequency sound enhancement effect is used to increase the amplitude of signals belonging to the high-frequency band, the human voice enhancement effect is used to increase the amplitude of signals belonging to the human voice frequency band, and the audio dynamic range compression effect is used to compress the dynamic range of the signal. The first preset condition includes one or more of the following: the duration of full-screen playback of the target multimedia resource exceeds the sixth threshold; the played audio is the original audio of the target multimedia resource; the target multimedia resource is played in full-screen mode on the interface of the current terminal device; the target multimedia resource is in playback mode; and the duration of unplayed content in the target multimedia resource exceeds the seventh threshold.

2. The method according to claim 1, characterized in that, Based on the special effects corresponding to the elder-specific sound effects, the original audio of the target multimedia resource is processed to obtain the audio to be played, including: Based on the original audio, various special effects corresponding to the elder-specific sound effects are applied sequentially to render the audio to be played. The original audio is used as the signal to be rendered for the first special effect, the rendering result of the previous special effect is used as the signal to be rendered for the next special effect, and the rendering result of the last special effect is used as the audio to be played.

3. The method according to claim 2, characterized in that, The rendering parameters for high-frequency sound enhancement effects include a first threshold and a first coefficient. The process of processing the original audio of the target multimedia resource based on the special effects corresponding to the elder's exclusive sound effects to obtain the audio to be played includes: When rendering a signal using high-frequency sound enhancement effects. The signal to be rendered is divided into a first signal and a second signal, where the first signal is a high-frequency signal. The amplitude of the first signal is increased by the first coefficient to obtain the third signal; The amplitude of the third signal is compared with the magnitude of the first threshold. When the amplitude of the third signal exceeds the first threshold, the amplitude of the third signal is reduced to the first threshold to obtain the fourth signal. The fourth signal and the second signal are combined to form the rendering result corresponding to the high-frequency sound enhancement effect.

4. The method according to claim 2, characterized in that, The rendering parameters for voice enhancement effects include a second coefficient and a third coefficient. The process of processing the original audio of the target multimedia resource based on the special effects corresponding to the elder's exclusive sound effects to obtain the audio to be played includes: When rendering a signal using voice enhancement effects. The signal to be rendered is divided into a fifth signal and a sixth signal, wherein the fifth signal is a signal in the human voice frequency band; The amplitude of the fifth signal is increased based on the second coefficient to obtain the seventh signal; The amplitude of the noise signal in the sixth signal is reduced by the third coefficient to obtain the eighth signal; The seventh signal and the eighth signal are combined to form the rendering result corresponding to the human voice enhancement effect.

5. The method according to claim 2, characterized in that, The rendering parameters for the audio dynamic range compression effect include a second threshold, a third threshold, a fourth coefficient, and a fifth coefficient, where the second threshold is greater than the third threshold. The process of processing the original audio of the target multimedia resource based on the special effects corresponding to the elder's exclusive sound effects to obtain the audio to be played includes: When rendering a signal using audio dynamic range compression effects. Compare the amplitude of the signal to be rendered with the magnitudes of the second threshold and the third threshold; When the amplitude of the signal to be rendered is greater than the second threshold, the amplitude of the signal to be rendered is reduced according to the fourth coefficient; When the amplitude of the signal to be rendered is less than the third threshold, the amplitude of the signal to be rendered is increased according to the fifth coefficient.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: When playing the target multimedia resource, receive feedback signals from the user; Based on the feedback signal, adjust the rendering parameters of at least one special effect corresponding to the elder-specific sound effect.

7. The method according to any one of claims 1-5, characterized in that, The rendering parameters for any of the special effects corresponding to the elder-specific sound effects are the rendering parameters preset on the current terminal device, or the rendering parameters preset by the user account logged in on the current terminal device.

8. The method according to claim 1, characterized in that, The method further includes: In response to the start of playback of the target multimedia resource, query the current interface display mode of the terminal device; In response to the current terminal device's interface display mode being the elder mode, it is determined whether the playback status of the target multimedia resource meets the first preset condition. In response to the current terminal device's interface display mode being any mode other than the senior mode, query the user's playback history on the current terminal device, or query the user's identity information recorded on the current terminal device.

9. The method according to claim 1, characterized in that, When the target multimedia resource is played on the current terminal device using a sound effect mode other than the elderly-specific sound effect, the method further includes: Receive a sound effect mode selection signal from an associated device, the sound effect mode selection signal indicating a sound effect mode; When the sound effect mode selection signal indicates an audio effect exclusive to elders, the audio of the target multimedia resource will display the special effects corresponding to the audio effect exclusive to elders.

10. The method according to claim 1, characterized in that, The method further includes: The audio to be played is stored. When the user plays the target multimedia resource again and is in the sound effect mode of the elder-specific sound effect, the stored audio to be played is directly used to replace the original audio.

11. An audio signal processing device, characterized in that, The device is applied to a terminal device, which includes a download module, a decoding module, and a physical player. The terminal device is used to: send the identifier of the target multimedia resource selected by the user to the server; receive the address of the encapsulated data packet; download the data packet from the received address using the download module; and decode the data packet using the decoding module to obtain the sound effect mode supported by the target multimedia resource and the data of the target multimedia resource. Use the physical player to start playing the target multimedia resource; The device includes: The display module is used to show users the sound effect modes supported by the target multimedia resource; The first audio processing module is used to respond to the user's selection of the elder-specific sound effect in the sound effect mode, process the original audio of the target multimedia resource based on the special effect corresponding to the elder-specific sound effect to obtain the audio to be played; when playing the target multimedia resource, replace the original audio with the audio to be played, so that the audio of the target multimedia resource presents the special effect corresponding to the elder-specific sound effect; The first judgment module is used to determine whether the playback status of the target multimedia resource meets the first preset condition in response to the current terminal device's interface display mode being the elder mode; the second judgment module is used to determine whether the playback status of the target multimedia resource meets the first preset condition when the current terminal device's interface display mode is in a mode other than the elder mode, and when the user's playback history on the current terminal device determines that the number of times the user selects the elder-specific sound effect when playing multimedia resources on the current terminal device is greater than a fourth threshold, or when the user's age is determined to be greater than a fifth threshold based on the user's identity information recorded on the current terminal device; the second audio processing module is used to, in response to the target multimedia resource's playback status meeting the first preset condition, make the audio of the target multimedia resource present the special effects corresponding to the elder-specific sound effect when playing the target multimedia resource; The special effects corresponding to the exclusive sound effects for elders include high-frequency sound enhancement effects, human voice enhancement effects, and audio dynamic range compression effects. The high-frequency sound enhancement effect is used to increase the amplitude of signals belonging to the high-frequency band, the human voice enhancement effect is used to increase the amplitude of signals belonging to the human voice frequency band, and the audio dynamic range compression effect is used to compress the dynamic range of the signal. The first preset condition includes one or more of the following: the duration of full-screen playback of the target multimedia resource exceeds the sixth threshold; the played audio is the original audio of the target multimedia resource; the target multimedia resource is played in full-screen mode on the interface of the current terminal device; the target multimedia resource is in playback mode; and the duration of unplayed content in the target multimedia resource exceeds the seventh threshold.

12. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 10 when executing instructions stored in the memory.

13. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Sound effect processing method and apparatus

    CN105405448A

  • Sound effect setting method, device, electronic device and storage medium

    CN109271128A

  • Volume adjustment method and system, display terminal and readable storage medium

    CN110996159A