Sound effect processing method, sound effect processing device, terminal and storage medium
By obtaining the current context type of the target user and using mapping information to generate matching audio effects, the problem of terminal sound effects failing to meet user needs is solved, achieving a clear and comfortable audio experience in different environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-07-22
- Publication Date
- 2026-07-31
AI Technical Summary
The sound effects played by existing terminals cannot meet users' high requirements, and it is difficult to provide a clear and comfortable sound experience in different environments.
By obtaining the target user's current context type, using mapping information to determine contextual audio effects, processing the audio to be processed, and generating audio that matches the user's current context.
It achieves the matching of the audio heard by the target user in different environments with their current situation, providing a clear and comfortable sound experience and reducing the need for users to manually adjust the volume.
Smart Images

Figure CN115696170B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sound effects processing, specifically to a sound effects processing method, a sound effects processing device, a terminal, and a storage medium. Background Technology
[0002] In recent years, with the advancement of technology, people have increasingly higher requirements for the sound effects of the sound played on the terminal. Various devices or software related to sound effect processing have appeared on the market to improve the stereo effect and spatial layering of the sound playback.
[0003] However, the sound effects played on existing terminals cannot meet users' needs. Therefore, the current sound effect processing methods still have the problem of failing to meet users' needs. Summary of the Invention
[0004] This application provides a sound effect processing method, a sound effect processing device, a terminal, and a storage medium, which can make the audio to be played adaptable to the user's current environment.
[0005] This application provides a sound effect processing method, including: Get audio files to be processed from other users, where other users are users other than the target user; Determine the scenario type of the target user's current situation; Obtain mapping relationship information, which includes the mapping relationship between scene type and preset audio effects; Based on the mapping relationship information, determine the scene audio effects corresponding to the scene type; Use contextual audio effects to process the audio of other users, and obtain the contextual audio of other users. Play contextual audio from other users so that the contextual audio listened to by the target user matches the current context of the target user.
[0006] This application embodiment also provides a sound effect processing device, including: The audio acquisition unit is used to acquire audio to be processed from other users, which are users other than the target user. The scenario type determination unit is used to determine the scenario type of the target user's current situation; The mapping relationship acquisition unit is used to acquire mapping relationship information, which includes the mapping relationship between scene type and preset audio effects; The target sound effect determination unit is used to determine the scene audio effects corresponding to the scene type based on the mapping relationship information. The audio effects processing unit is used to apply contextual audio effects to the audio to be processed of other users, thereby obtaining the contextual audio of other users. The audio playback unit is used to play contextual audio from other users so that the contextual audio listened to by the target user matches the current context of the target user.
[0007] In some embodiments, the scenario type determination unit is configured to: Obtain the target user's current geographical location information; Get the current time and location of the target user at their current geographical location; The scenario type is determined based on the current geographical location information and the current time.
[0008] In some embodiments, the contextual audio effects have corresponding first-channel pitch operators and second-channel pitch operators, which are used to calculate the pitch difference between the left and right channels. Other users' contextual audio includes first-channel audio and second-channel audio. The audio effects processing unit is used for: The first channel audio is obtained by convolving the audio from other users with the first channel pitch operator. The second-channel audio is obtained by convolving the audio from other users using the second-channel pitch operator.
[0009] In some embodiments, after obtaining the contextual audio from other users, the contextual audio effect also has corresponding first channel audio intensity and second channel audio intensity, and the device is further used for: The amplitude of the first channel audio is adjusted according to the first channel audio intensity to obtain the processed first channel intensity audio. The amplitude of the second channel audio is adjusted according to the intensity of the second channel audio to obtain the processed second channel intensity audio.
[0010] In some embodiments, the contextual audio effect also includes reverberation information, and after obtaining contextual audio from other users, the device is further used to: The reverberation information is used to process the background audio of other users to obtain the reverberated audio of other users.
[0011] In some embodiments, the reverberation information includes corresponding direct phonon information, early reflected phonon information, and late reflected phonon information. The device performs reverberation processing on the context audio of other users based on the reverberation information to obtain the reverberated audio of other users. The device also includes: Based on the direct phonon information, the contextual audio of other users is processed to obtain the direct phonon audio; Early reflection phonon information is used to process the contextual audio of other users into early reflection phonons to obtain early reflection audio. Based on the post-reflection phonon information, the contextual audio of other users is processed by post-reflection sound to obtain post-reflection sound audio. The direct sound audio, early reflected sound audio, and late reflected sound audio are superimposed to obtain the reverberation audio of other users.
[0012] In some embodiments, the device performs direct sound processing on the contextual audio of other users based on direct sound phonon information to obtain direct sound audio. The direct sound phonon information includes a direct sound operator. The device is also used to: The direct sound operator is used to process the contextual audio of other users to obtain direct sound audio.
[0013] In some embodiments, the device performs early reflection processing on the contextual audio of other users based on early reflection phonon information to obtain early reflection audio. The early reflection phonon information includes a frequency delay filter type and an early reflection operator. The device is also used to: The contextual audio of other users is delayed based on the type of frequency delay filter to obtain frequency-delayed audio. Early reflection audio is obtained by using an early reflection operator to process the frequency-delayed audio.
[0014] In some embodiments, post-reflection phonon processing is performed on the context audio of other users based on post-reflection phonon information to obtain post-reflection audio. The post-reflection phonon information includes frequency delay filter type, frequency filter type, phase delay filter type, and post-reflection operator, including: Based on the frequency delay filter type, the contextual audio of other users is subjected to audio delay processing to obtain frequency-delayed audio; The frequency-delayed audio is filtered according to the type of frequency filter to obtain the frequency-filtered audio. The frequency-filtered audio is phase-filtered according to the phase delay filter type to obtain the phase-delayed audio. The post-reflection sound operator is used to process the phase-delayed audio with post-reflection sound to obtain the post-reflection sound audio.
[0015] This application also provides a terminal, including a memory storing multiple instructions; the processor loads the instructions from the memory to execute the steps in any of the sound effect processing methods provided in this application.
[0016] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the sound effect processing methods provided in this application.
[0017] This application embodiment can acquire audio from other users (other users are those other than the target user); determine the scenario type of the target user's current situation; acquire mapping relationship information, which includes the mapping relationship between scenario type and preset audio effects; determine the scenario audio effect corresponding to the scenario type based on the mapping relationship information; apply the scenario audio effect to the audio from other users to obtain the scenario audio from other users; and play the scenario audio from other users so that the scenario audio listened to by the target user matches the scenario currently in the target user's situation.
[0018] In this application, the scenario type of the current situation of the target user can be determined first. Based on the scenario type, the scenario audio effect that maps to it can be found from the preset audio effects. In this way, the terminal can process the audio to be processed of other users according to the scenario audio effect to obtain the scenario audio of other users. Then, the scenario audio of other users can be played through electronic devices. Thus, this solution proposes a new way of audio processing, so that the scenario audio listened to by the target user matches the scenario in which the target user is currently in, which is beneficial for the target user to directly obtain clear and comfortable sound. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1a This is a schematic diagram of a scene for the sound effect processing method provided in the embodiments of this application; Figure 1b This is a schematic flowchart of the sound effect processing method provided in the embodiments of this application; Figure 1c This is a schematic diagram of stereo generation provided in an embodiment of this application; Figure 1d This is a schematic diagram of sound effect processing provided in an embodiment of this application; Figure 1e This is a schematic diagram of dual-channel audio generation provided in an embodiment of this application; Figure 1f This is a schematic diagram of the reverberator result corresponding to the reverberation processing provided in the embodiment itself; Figure 2 This is a schematic diagram illustrating the application of the sound effect processing method provided in this application embodiment in a server scenario; Figure 3 This is a schematic diagram of the sound processing device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the mobile terminal provided in the embodiments of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] This application provides a sound effect processing method, a sound effect processing device, a terminal, and a storage medium.
[0023] Specifically, the audio processing device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet, smart Bluetooth device, laptop, or personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.
[0024] In some embodiments, the sound processing device may also be integrated into multiple electronic devices, such as multiple servers, with multiple servers implementing the sound processing method of this application.
[0025] In some embodiments, the server may also be implemented as a terminal.
[0026] The following sections provide detailed descriptions of each example. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0027] For example, see Figure 1a Taking the integration of audio processing devices into electronic devices as an example, the electronic device acquires audio to be processed from other users (other users are those other than the target user), determines the scenario type of the target user's current situation, obtains mapping relationship information, which includes the mapping relationship between scenario type and preset audio effects, and, based on the mapping relationship information, uses scenario audio effects to process the audio to be processed from other users according to the scenario type, obtains the scenario audio of other users, and plays the scenario audio of other users so that the scenario audio listened to by the target user matches the scenario of the target user's current situation.
[0028] Before processing the audio of other users, the system obtains the context type of the current situation of the target user and determines the context audio effect corresponding to the context type based on the mapping relationship information. The system then processes the audio of other users according to the context audio effect and plays the context audio of the other user. In this way, the context audio heard by the target user can match the current situation of the target user, which is beneficial for the target user to directly obtain clear and comfortable sound.
[0029] In this embodiment, a sound effect processing method is provided, such as... Figure 1b and Figure 1c As shown, the specific process of this sound effect processing method can be as follows: 101. Obtain audio files from other users, where other users are those other than the target user.
[0030] The audio to be processed by other users refers to audio transmitted by other users through their audio terminals to the target user's listening terminal, which is then subject to audio effects processing. This audio can be unprocessed or, although processed, requires further processing. The audio to be processed by other users can originate from video calls, voice calls, recorded audio, video conferencing, etc.
[0031] For example, in some embodiments, the audio processing algorithm or audio processing device can acquire the audio to be processed from other users in real time. For instance, during a real-time voice call, the audio processing device can directly acquire the audio to be processed from other users in the voice call.
[0032] For example, in some embodiments, the audio processing algorithm or audio processing device can obtain the audio to be processed from other users locally. For instance, the audio to be processed from other users is stored on the local storage of the target user's self-owned terminal, and the audio processing software or audio processing device on the terminal can obtain the audio to be processed from other users on the local storage.
[0033] For example, in some embodiments, the audio processing algorithm or audio processing device can obtain the audio to be processed from other users in a remote location. For instance, the audio to be processed from other users is stored on a server, or the audio to be processed from other users needs to be transmitted to the target user's terminal through the server. In this case, the audio processing algorithm or audio processing device on the target user's terminal can obtain the audio to be processed from other users on the server.
[0034] 102. Determine the scenario type of the target user's current situation.
[0035] The target user is the person listening to the audio. The current scenario refers to the target user's geographical location and current time while listening to the audio. Scenario types can be categorized by geographical location within a specific time period, such as a bar between 8 PM and 2 AM the next day, or a forest between 5 PM and 8 PM. They can also be categorized by ambient noise, such as low noise, relatively low noise, medium noise, relatively high noise, and high noise. Low noise scenarios could be forests or bedrooms between midnight and 8 AM the next day. Relatively low noise scenarios could be concert halls between 9 AM and 10 PM, or lecture halls between 9 AM and 6 PM. Medium noise scenarios could be cars between 7 AM and 10 PM, or restaurants during mealtimes. Relatively high noise scenarios could be roadsides between 7 AM and 10 PM, subways between 7 AM and 10 PM, or airports between 7 AM and 10 PM. High noise scenarios could be bars between 8 PM and 2 AM the next day, or concerts between 8 PM and midnight. Alternatively, a combination of specific geographical location and ambient noise can be used.
[0036] For example, in some embodiments, the scenario type of the target user's current situation can be determined locally. For instance, the sound effect processing algorithm or sound effect processing device can determine the scenario type of the current situation after obtaining the target user's current situation.
[0037] For example, in some embodiments, the scenario type of the target user's current situation can be determined remotely. For instance, the calculation process for determining the scenario type of the target user's current situation is performed on a server, that is, the sound effect processing algorithm is loaded into the server, and the sound effect processing device obtains the calculation result from the server.
[0038] In some embodiments, step 102 may include the following steps: In some embodiments, the current geographic location information of the target user is obtained.
[0039] The current geographic location information can be the target user's location area on a map, such as a supermarket, cinema, or bar. Alternatively, it can be the location selected and set by the target user within the sound processing device or server; this location can be a specific spatial location, such as a supermarket, cinema, or bar. Finally, it can be input location information, either text-based or voice-based.
[0040] For example, in some embodiments, the current geographic location information of the target user can be generated by location software or location device. After the geographic location information is generated, the sound effect processing device retrieves the geographic location information from the location software or location device.
[0041] For example, in some embodiments, the geographic location information may be pre-stored in the sound processing device or server. The target user selects the geographic location information corresponding to their location locally or remotely. For example, when the target user is in a bar, the geographic location information selected by the target user may be the bar.
[0042] For example, in some embodiments, the target user inputs geolocation information to the sound processing device via text or voice.
[0043] In some embodiments, the current time of the target user's current geographical location is obtained.
[0044] The current time is the time when the target user is at the current geographical location. The current time can be read from the Internet or from the timekeeping system of the target user's terminal.
[0045] In some embodiments, the scenario type is determined based on the current geographic location information and the current time.
[0046] The scenario types can be categorized by geographical location within a specific time period. Specific scenario types can include bars from 8 p.m. to 2 a.m. the next day, forests from 5 p.m. to 8 p.m., etc.
[0047] For example, in some embodiments, the noise environment corresponding to a geographical location within a specific time period may also be different. For example, a bar with geographical location information of 8 pm to 2 am the next day is in operation, so the current noise environment corresponding to the bar is a high noise environment.
[0048] 103. Obtain mapping relationship information, which includes the mapping relationship between scene type and preset audio effects.
[0049] The preset audio effects can be pre-recorded in the audio processing algorithm or stored locally or remotely. The audio effects are used to modify the sound effects of other users' audio to be processed. The parameters of the audio effects can specifically include the number of channels, pitch, audio intensity, reverb, etc.
[0050] The mapping relationship between scene types and preset audio effects can be a pre-defined mapping relationship based on actual needs, or it can be a mapping relationship calculated by collecting big data on audio effects and scene types. The mapping relationship information can be pre-recorded in the sound effect processing algorithm or pre-stored in electronic devices such as sound effect processing devices and servers. When obtaining the mapping relationship, it can be retrieved from local memory or from other electronic devices via the network.
[0051] For example, in some embodiments, the mapping relationship information can be recorded in the sound effect processing algorithm. Specifically, it can be recorded in the sound effect processing algorithm in the form of a mapping relationship logic language. This sound effect processing algorithm can correspond to a conventional music player, or it can be loaded onto the algorithm corresponding to social software or online video software.
[0052] For example, in some embodiments, the mapping relationship information can also be stored locally. Specifically, the mapping relationship information can be stored on a local storage unit. When the audio processing device needs to perform audio processing on the audio to be processed by other users, the mapping relationship information in the storage unit can be directly retrieved locally.
[0053] For example, in some embodiments, the mapping relationship information can also be stored in a remote location. When the local machine performs sound effect processing on the audio to be processed of other users, the mapping relationship information can be retrieved from the remote location. Alternatively, the local machine can receive a user's request, perform sound effect processing on the audio to be processed on the server, and directly retrieve the mapping relationship information from the remote location.
[0054] 104. Based on the mapping relationship information, determine the scene audio effects corresponding to the scene type.
[0055] Among them, the contextual audio effects can be preset audio effects associated with the context type of the current context.
[0056] For example, in some embodiments, when performing sound effect processing on other users' audio to be processed locally, mapping relationship information stored locally or remotely is retrieved. Since the mapping relationship information includes the mapping relationship between scene type and preset audio effect, the preset audio feature corresponding to the scene type can be determined by the scene type. The corresponding preset audio feature is the scene audio effect.
[0057] 105. Use contextual audio effects to process the audio of other users, and obtain the contextual audio of other users.
[0058] Among them, reference Figure 1dThe parameters of contextual audio effects can specifically include the number of channels, pitch, audio intensity, reverb, etc. The audio to be processed from other users is processed according to the contextual audio effects to obtain the contextual audio of other users. The contextual audio of other users can include changes in the number of channels, pitch, audio intensity (volume), reverb, etc. In this way, the audio processing of the audio to be processed from other users is based on the context type of the target user's current situation, which helps the user to obtain clear and comfortable sound in the current situation.
[0059] In some embodiments, the sound effects processing steps can be performed locally.
[0060] In some embodiments, the sound effects processing steps can be performed remotely.
[0061] In some embodiments, the sound effect processing steps can be performed on social media software, and the sound effect processing algorithm on the social media software can be stored on the terminal or on the server corresponding to the social media software.
[0062] In some embodiments, step 105 may include the following steps: The contextual audio effects have corresponding first-channel pitch operators and second-channel pitch operators. These operators are used to calculate the pitch difference between the left and right channels. Other users' contextual audio includes first-channel and second-channel audio. The audio effects processing unit is used for: The first channel audio is obtained by convolving the audio from other users with the first channel pitch operator. The second-channel audio is obtained by convolving the audio from other users using the second-channel pitch operator.
[0063] Among them, reference Figure 1e The first and second channel pitch operators can be obtained using the Head-Related Transfer Function (HRTF). HRTF is a source location function that integrates the time delay difference (ITD), sound pressure difference (IID), and somatic acoustic reflection spectrum characteristics; in other words, it represents the response of the sound transmission path. The excitation response data corresponding to the HRTF transfer function is HRIR. The most commonly used HRIR datasets are those from CIPIC and MIT. For example, the CIPIC experimental data collected time-domain measurements of binaural listening signals from 45 measurement subjects, each at 25 different horizontal orientations and 50 different vertical orientations, totaling 1250 orientations.
[0064] In some embodiments, the contextual audio effect carries pitch parameters for the left channel and pitch parameters for the right channel. The first channel pitch operator and the second channel pitch operator are obtained from the pitch parameters carried by the contextual audio effect based on the head association transfer function.
[0065] In some embodiments, the first channel pitch operator is the convolution of the audio to be processed from other users with the left channel time-domain measurement data from the corresponding location, as shown in the following formula:
[0066] Where n is the label for audio quantization, It's convolution. u ( n () represents the mono audio signal labeled n. For the left channel time-domain measurement data labeled n, This is the first channel audio (left channel audio).
[0067] In some embodiments, the first channel pitch operator is the convolution of the audio to be processed from other users with the right channel time-domain measurement data from the corresponding location, as shown in the following formula:
[0068] Where n is the label for audio quantization, It's convolution. For the mono audio signal labeled n, For the right channel time-domain measurement data labeled n, This is the second channel audio (right channel audio).
[0069] As can be seen from the above formula, mono audio is processed by the first channel pitch operator and the second channel pitch operator to obtain the first channel audio and the second channel audio.
[0070] Since the Head-Related Transfer Function (HRTF) is mainly frequency-dependent, the frequency of the audio affects the pitch of the sound. Therefore, after the mono audio passes through the first channel pitch operator and the second channel pitch operator, there is a pitch difference between the left and right channels. This allows the mono audio to be converted into stereo audio after sound effects processing, thus obtaining stereo sound.
[0071] In some embodiments, after obtaining the contextual audio from other users, the contextual audio effect also has corresponding first channel audio intensity and second channel audio intensity, and the device is further used for: The amplitude of the first channel audio is adjusted according to the first channel audio intensity to obtain the processed first channel intensity audio. The amplitude of the second channel audio is adjusted according to the intensity of the second channel audio to obtain the processed second channel intensity audio.
[0072] The closer the sound source is to the listener, the higher the sound pressure level; conversely, the farther the sound source is from the listener, the lower the sound pressure level. This creates different auditory perception effects at different distances, which are then mapped to corresponding audio values. For example, sound pressure level can be adjusted by changing the amplitude of the audio signal. The amplitude affects the intensity of the sound, which in practical use is reflected in the volume.
[0073] The formula for the relationship between sound pressure and distance is:
[0074] in, It is a specific distance from the sound source. It is the target distance from the sound source. It is the sound pressure level at a specific distance. It is the sound pressure level at the target distance.
[0075] For example, When the value is 2, the sound pressure decreases by 6 dB.
[0076] The formula relating sound pressure level and audio frequency is:
[0077] in, is the sound pressure level, n is the label for audio quantization, N is a non-zero natural number, and LPC is the offset, which is used to correct the audio amplitude.
[0078] In some embodiments, the ambient audio effect carries sound intensity parameters corresponding to the left and right channels, respectively. The left channel sound intensity parameter represents the first channel audio intensity, and the right channel sound intensity parameter represents the second channel audio intensity. The first channel audio intensity corresponds to the first channel audio amplitude, and the audio amplitude of the first channel audio is adjusted according to this amplitude parameter. The second channel audio intensity corresponds to the second channel audio amplitude, and the audio amplitude of the second channel audio is adjusted according to this amplitude parameter. In summary, this achieves the adjustment of the sound intensity of the left and right channels.
[0079] In some embodiments, the contextual audio effect also includes reverberation information. After obtaining the contextual audio of other users, the device is further configured to: perform reverberation processing on the contextual audio of other users based on the reverberation information to obtain the contextual audio of other users with reverberation.
[0080] Among them, see Figure 1fReverberation is a disordered state formed after sound undergoes countless reflections in space. The direct sound, early reflections, and late reflections are the three elements that form a sound field. In human auditory perception, reverberation is usually used to create a chaotic sense of space or to judge the size of the space, which helps to make the played audio more three-dimensional.
[0081] In some embodiments, depending on the specific parameters of the contextual audio effect carrying reverberation information, the contextual audio of the other user can be the first channel audio and the second channel audio, or it can be the first channel intensity audio and the second channel intensity audio, which increases the spatial sense of the audio and makes the played audio more three-dimensional.
[0082] In some embodiments, the reverberation information includes corresponding direct phonon information, early reflected phonon information, and late reflected phonon information. The device is further configured to perform reverberation processing on the context audio of other users based on the reverberation information to obtain reverberated context audio for other users. In some embodiments, direct sound processing is performed on the contextual audio of other users based on direct sound phonon information to obtain direct sound audio. The direct sound phonon information includes direct sound operators. Direct sound processing is performed on the contextual audio of other users using direct sound operators to obtain direct sound audio.
[0083] Direct audio / video signal =
[0084] in, It can be the first channel audio, the second channel audio, the first channel intensity audio, or the second channel intensity audio. The attenuation factor is used to ensure that the direct audio frequency is the original audio frequency.
[0085] The technical term for reverberation decay time is reverberation time or reverberation decay. The reverberation time is calculated as the time required for the reverberation level to decay by 60 dB from the time the sound stops.
[0086] Therefore, the attenuation factor here This can be represented as reverberation time.
[0087] In some embodiments, the device performs early reflection processing on the context audio of other users based on early reflection phonon information to obtain early reflection audio. The early reflection phonon information includes a frequency delay filter type and an early reflection operator. The device is also used to: The contextual audio of other users is delayed based on the type of frequency delay filter to obtain frequency-delayed audio. Early reflection audio is obtained by performing early reflection processing on frequency-delayed audio using an early reflection operator.
[0088] Among them, the frequency delay filter type can be a finite impulse response (FIR) filter, specifically an 18-point finite impulse response filter, used to delay the frequency of audio.
[0089]
[0090] in, It can be the frequency-delayed audio from the first channel, the frequency-delayed audio from the second channel, the frequency-delayed audio from the first channel intensity, or even the audio from the second channel intensity used to select other users' audio for processing. The early reflection attenuation factor refers to the sound that arrives after only a few reflections. The specific response of the early reflection sound is that the sound is relatively loud.
[0091] As can be seen from the above, the early reflection attenuation factor can also be called the early reflection attenuation time.
[0092] In some embodiments, the device performs post-reflection sound processing on the context audio of other users based on post-reflection phonon information to obtain post-reflection sound audio. The post-reflection phonon information includes frequency delay filter type, frequency filter type, phase delay filter type, and post-reflection sound operator. The device is also used for: Based on the frequency delay filter type, the contextual audio of other users is subjected to audio delay processing to obtain frequency-delayed audio; The frequency-delayed audio is filtered according to the type of frequency filter to obtain the frequency-filtered audio. The frequency-filtered audio is phase-filtered according to the phase delay filter type to obtain the phase-delayed audio. The post-reflection sound operator is used to process the phase-delayed audio with post-reflection sound to obtain the post-reflection sound audio.
[0093] The frequency delay filter can be a finite impulse response (FIR) filter, specifically an 18-point FIR filter, used to delay the frequency of audio. Frequency filters are used to filter low, mid, or high-frequency audio; specific frequency filters can be low-pass comb filters. The phase delay filter can be an all-pass filter used to correct the phase of audio.
[0094]
[0095] in, It can be the frequency-delayed audio from the first channel audio after sequentially undergoing frequency delay, frequency filtering, and phase delay; it can also be the frequency-delayed audio from the second channel audio after sequentially undergoing frequency delay, frequency filtering, and phase delay; it can also be the frequency-delayed audio from the first channel intensity audio after sequentially undergoing frequency delay, frequency filtering, and phase delay; and it can also be the frequency-delayed audio from the second channel intensity audio after sequentially undergoing frequency delay, frequency filtering, and phase delay. The early reflection attenuation factor is the sound that has been reflected multiple times in the later stage. The later emitted sound is a continuous sound.
[0096] As can be seen from the above, the late reflection attenuation factor can also be called the late reflection attenuation time.
[0097] In some embodiments, the direct sound audio, the early reflected sound audio, and the later reflected sound audio are superimposed to obtain the reverberated background audio of other users, thereby giving the reverberated background audio of other users multiple reflected sounds, which helps to make the reverberated background audio of other users more spatial and three-dimensional when played.
[0098] In some embodiments, step 105 may further include: The contextual audio for other users may include processing the audio (mono) of other users to generate stereo (two-channel) audio, then adjusting the audio amplitude of the stereo by the relationship between sound pressure and distance. Since the audio amplitude is related to the sound intensity, the intensity of the stereo is adjusted. After the intensity-adjusted stereo is reverberated, the reverberated audio is obtained, which is the contextual audio for other users.
[0099] 106. Play contextual audio from other users so that the contextual audio listened to by the target user matches the current context of the target user.
[0100] In some embodiments, the target user's terminal plays contextual audio from other users. Since the contextual audio from other users is generated based on the target user's current context, the contextual audio from other users can match the target user's current context when it is played.
[0101] The audio processing solution provided in this application can be applied to various audio playback scenarios. For example, taking voice calls as an example, when a user answers a voice call, they are affected by surrounding environmental factors and will adjust the volume of the voice call according to the environment. The solution provided in this application can more easily make the audio played by the terminal suitable for the user's current situation, allowing the user to directly obtain clear and comfortable sound, further reducing the user's manual adjustment of the volume, so that the user's auditory perception of the audio is like a face-to-face conversation with the other party.
[0102] The method provided in this application embodiment can make the audio to be played suitable for the user's current situation. For example, this application embodiment can know the situational audio effects according to the type of situation the user is in, and perform sound effect situational processing on the audio to be processed of other users according to the situational audio effects, so that the situational audio of other users can be suitable for the user's current situation when played.
[0103] As can be seen from the above, the embodiments of this application can reduce the impact of the user's situation on the audio to be played. Therefore, this solution can meet the user's needs for audio effects in different situations, making the audio to be played adaptable to the user's current environment.
[0104] The method described in the above embodiments will be further described in detail below.
[0105] The sound effect processing method provided in this application can be applied to various applications such as voice calls, voice messages, and voice-based interactions. This application proposes a spatial sound effect voice interaction method that matches the current time and geographical coordinates. This application is a technology that combines real-world environment with virtual spatial acoustics. Through this technology, it can address users' different auditory needs and provide a novel auditory experience in different acoustic scenarios. It differs from existing technologies that rely on detection devices to detect user location and movement trajectory and present an actual on-site auditory experience through spatial sound effects, and also from spatial sound effect methods that rely entirely on virtual stereo movement trajectories based on a design script.
[0106] This application is based on obtaining the target user's current time and current geographical location information. Mobile devices such as smartphones and tablets have this detection capability. The obtained current time and current geographical location information will be mapped to a contextual audio effect. This mapping table can be stored locally or in the cloud, as shown in Table 1 below. If stored in the cloud, the target user's local device will upload the retrieved current time and current geographical location information to the cloud. The cloud will then search for the contextual audio effect using the mapping table and distribute it to the local device.
[0107] As shown in Table 1 below: Mapping table
[0108] The following are the audio effects for the scene: 1) Default sound effect template: Speaker is 1 meter away, facing each other; 2) Ear-to-ear sound effect template: The speaker is close to the listener's ear to communicate; 3) Walking and moving sound effect template: The speaker is within a certain distance range and communicates with the listener while moving slowly along a random or preset trajectory; 4) Speech-style sound effect template: The speaker is at a medium to long distance, with a loud voice and a certain reverberation effect; 5) Surprising sound effect template: The speaker's position is not fixed and the movement trajectory is random. For example, the speaker appears to the left front of the listener in the first sentence, behind the listener in the second sentence, and close to the listener's ear in the next sentence, giving people a surprising auditory experience. 6) Surround sound effect template: The speaker maintains a certain distance from the listener and communicates by rotating 360 degrees around the listener in the horizontal direction; 7) Fly-in-fly-out sound effect template: The speaker moves from a distance toward the listener at a high speed, or moves away from a position close to the listener at a high speed. Each scene audio effect includes a series of sound image location, distance, reverberation parameters, etc. Based on these parameters, virtual stereo is generated using relevant technologies, and the generated stereo is played through headphones or multiple speakers.
[0109] In real life, target users have different auditory experience needs depending on the environment they are in. For example: in a noisy environment, such as a bar, target users prefer a close-up or whispered conversation to avoid interference from ambient noise; in a quiet environment, such as a grove of trees, they need a more natural and casual conversation, hoping that the other person is also moving around freely; in an open environment, such as a large stadium, they want the sound to have a slight reverberation effect to match the acoustic environment.
[0110] The method described in the above embodiments will be further described in detail below.
[0111] In this embodiment, a server will be used as an example to describe the method of this application embodiment in detail.
[0112] like Figure 2 As shown, the specific process of a sound effect processing method is as follows: 201. The server retrieves audio files from other users, excluding the target user.
[0113] For example, a target user can directly upload other users' unprocessed audio to the server. The server directly obtains the unprocessed audio from other users, recognizing that it is already stored on the server, and then retrieves it from the target user's specified location. If the unprocessed audio from other users is a live call, the target user can first establish communication with the server, and the live call is directly transmitted from the target user's terminal to the server. If the unprocessed audio from other users is audio pre-stored on the server by the target user's terminal, the target user specifies the audio on the server through their terminal. Alternatively, the unprocessed audio from other users can be stored on the target user's terminal, and when the target user needs to process the audio, they upload it to the server through their terminal.
[0114] 202. The server determines the scenario type of the target user's current situation.
[0115] In some embodiments, step 202 may include the following steps: In some embodiments, the server obtains the target user's current geographic location information.
[0116] For example, the target user's terminal obtains the target user's geographical location information through location software or a location device, and then uploads the geographical location information to the server. When performing sound effects processing on other users' audio, the geographical location information can be uploaded manually by the target user or automatically by the terminal.
[0117] In some embodiments, the server obtains the current time at which the target user is located in the current geographical location.
[0118] For example, after the target user's terminal locates the target user, the target user's terminal obtains the Internet time that generates the current location location through the Internet and uploads the Internet time to the server. Alternatively, after receiving the current geographical location information, the server generates the time of receiving the current geographical location information as the current time. The current time can also be uploaded manually by the target user or automatically by the terminal.
[0119] In some embodiments, the server determines the scenario type based on the current geographic location information and the current time.
[0120] For example, after receiving the current geographic location information, the server can carry the location name of the target user on the map. At the same time, the server also receives the current time when the target user is at the current geographic location. The server can determine the scenario type of the current situation of the target user based on the location name corresponding to the current time.
[0121] 203. The server obtains mapping relationship information, which includes the mapping relationship between scene type and preset audio effects.
[0122] For example, mapping relationship information can be stored in the server. When the server obtains the audio to be processed from other users, the current geographical location information of the target user, and the current time at the current geographical location, the server can retrieve the mapping relationship information stored in the server after determining the scenario type based on the geographical location information.
[0123] 204. Based on the mapping relationship information, determine the scene audio effects corresponding to the scene type.
[0124] For example, based on the mapping information, the audio effects corresponding to this scenario type can be found.
[0125] 205. The server uses contextual audio effects to process the audio of other users, thus obtaining the contextual audio of other users.
[0126] For example, the server processes the audio of other users based on the parameters in the contextual audio effects, and then uploads the contextual audio of other users to the target user's terminal.
[0127] In some embodiments, step 205 may include the following steps: The contextual audio effects have corresponding first-channel pitch operators and second-channel pitch operators. These operators are used to calculate the pitch difference between the left and right channels. Other users' contextual audio includes first-channel and second-channel audio. The audio effects processing unit is used for: The first channel audio is obtained by convolving the audio from other users with the first channel pitch operator. The second-channel audio is obtained by convolving the audio from other users using the second-channel pitch operator.
[0128] For example, when other users' audio to be processed is mono, the server retrieves the first and second channel operators carried by the contextual audio effects. The server then performs audio effects processing on the other users' audio based on the first and second channel operators, thereby obtaining stereo audio, so that the other users' contextual audio is played in stereo.
[0129] In some embodiments, after obtaining the contextual audio from other users, the contextual audio effect also has corresponding first channel audio intensity and second channel audio intensity, and the device is further used for: The amplitude of the first channel audio is adjusted according to the first channel audio intensity to obtain the processed first channel intensity audio. The amplitude of the second channel audio is adjusted according to the intensity of the second channel audio to obtain the processed second channel intensity audio.
[0130] For example, the server retrieves the first and second channel audio intensities carried by the contextual audio effects. The server adjusts the amplitude of the first channel audio based on the first channel audio intensity and adjusts the amplitude of the second channel audio based on the second channel audio intensity, thereby making the sound intensities of the first and second channel audio different. This enhances the stereo effect of the contextual audio for other users and also allows the target user to perceive the effect of changes in the distance of the sound source when the contextual audio for other users is played.
[0131] In some embodiments, the contextual audio effect also includes reverberation information. After obtaining the contextual audio of other users, the device is further configured to: perform reverberation processing on the contextual audio of other users based on the reverberation information to obtain the contextual audio of other users with reverberation.
[0132] In some embodiments, the reverberation information includes corresponding direct phonon information, early reflected phonon information, and late reflected phonon information. The device is further configured to perform reverberation processing on the context audio of other users based on the reverberation information to obtain reverberated context audio for other users. In some embodiments, direct sound processing is performed on the contextual audio of other users based on direct sound phonon information to obtain direct sound audio. The direct sound phonon information includes direct sound operators. Direct sound processing is performed on the contextual audio of other users using direct sound operators to obtain direct sound audio.
[0133] In some embodiments, the device performs early reflection processing on the context audio of other users based on early reflection phonon information to obtain early reflection audio. The early reflection phonon information includes a frequency delay filter type and an early reflection operator. The device is also used to: The contextual audio of other users is delayed based on the type of frequency delay filter to obtain frequency-delayed audio. Early reflection audio is obtained by using an early reflection operator to process the frequency-delayed audio.
[0134] In some embodiments, the device performs post-reflection sound processing on the context audio of other users based on post-reflection phonon information to obtain post-reflection sound audio. The post-reflection phonon information includes frequency delay filter type, frequency filter type, phase delay filter type, and post-reflection sound operator. The device is also used for: Based on the frequency delay filter type, the contextual audio of other users is subjected to audio delay processing to obtain frequency-delayed audio; The frequency-delayed audio is filtered according to the type of frequency filter to obtain the frequency-filtered audio. The frequency-filtered audio is phase-filtered according to the phase delay filter type to obtain the phase-delayed audio. The post-reflection sound operator is used to process the phase-delayed audio with post-reflection sound to obtain the post-reflection sound audio.
[0135] For example, the server retrieves the reverberation information carried by the contextual audio effects and performs reverberation processing on the contextual audio of other users. The reverberation information carries direct phonon information, early reflection phonon information, and late reflection phonon information. The server processes the contextual audio of other users based on the direct phonon information to generate direct sound audio, processes the contextual audio of other users based on the early reflection phonon information to generate early reflection sound audio, and processes the contextual audio of other users based on the late reflection phonon information to generate late reflection sound audio. Then, the direct sound audio, early reflection sound audio, and late reflection sound audio are superimposed, so that the server gives the contextual audio of other users a reverberation effect.
[0136] 206. The server transmits other users' contextual audio to the target user's terminal, and the target user's terminal plays the other users' contextual audio so that the contextual audio listened to by the target user matches the target user's current context.
[0137] For example, after processing the audio of other users with sound effects and context, the server obtains the contextual audio of other users and transmits the contextual audio of other users to the terminal of the target user. Since the contextual audio of other users is obtained based on the context type of the current context of the target user, the contextual audio heard by the target user can match the current context of the target user.
[0138] As can be seen from the above, in this embodiment, the server obtains the audio to be processed from other users, and obtains the current geographical location information and the current time of the target user at the current geographical location. Based on the current geographical location information and the current time, the server determines the scenario type of the target user's current situation, and then obtains the scenario audio effect corresponding to the scenario type from the preset audio effects. Based on the scenario audio effect, the server performs sound effect scenario processing on the audio to be processed from other users, so that the scenario audio of other users is suitable for the target user's current situation when played. This avoids the situation where the target user may not be able to hear the audio and needs to adjust the volume manually. At the same time, it also avoids the situation where the sound intensity of the audio is insufficient when played, which is beneficial for the target user to directly obtain clear and comfortable sound, so that the audio to be played can be adapted to the user's current situation.
[0139] To better implement the above methods, this application also provides an audio processing device, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers.
[0140] For example, in this embodiment, the method of this application embodiment will be described in detail by taking the sound effect processing device specifically integrated into a terminal as an example.
[0141] For example, such as Figure 3 As shown, the sound processing device may include: (a) Audio acquisition unit 301; The audio acquisition unit 301 is used to acquire audio to be processed from other users, where other users are users other than the target user.
[0142] For example, the audio acquisition unit 301 can be used to acquire the audio to be processed from other users locally, or to acquire the audio to be processed from other users remotely, or to acquire the audio to be processed from other users in real time.
[0143] (II) Scenario Type Determination Unit 302; The scenario type determination unit 302 is used to determine the scenario type of the current scenario of the target user.
[0144] In some embodiments, the scenario type determination unit 302 is configured to: Obtain the target user's current geographical location information; Get the current time and location of the target user at their current geographical location; The scenario type is determined based on geographical location information and the current time.
[0145] For example, the scenario type determination unit 302 is used to obtain the current geographical location information uploaded by the target user's terminal, and when the current geographical location is obtained, it generates the corresponding current time. The geographical location information carries the location name of the environment in which the target user is located. Based on the location name and the current time, the scenario type of the current scenario in which the target user is located is determined.
[0146] (iii) Mapping Relationship Acquisition Unit 303; The mapping relationship acquisition unit 303 is used to acquire mapping relationship information, which includes the mapping relationship between scene type and preset audio effects.
[0147] For example, the mapping relationship acquisition unit 303 is used to establish a connection between the scene type and the preset audio features, so as to determine which preset audio effects should be used to process the audio of other users according to the scene type.
[0148] (iv) Target sound effects and special effects determination unit 304; The target sound effect determination unit 304 is used to determine the scene audio effect corresponding to the scene type based on the mapping relationship information.
[0149] For example, the target sound effect determination unit 304 is used to receive the determined scene type and find the scene audio effect corresponding to the scene type in the mapping relationship information according to the scene type.
[0150] In some embodiments, the target sound effect determination unit 304 is further configured to: Determine the acquisition time of other users' pending audio; The desired contextual audio effects for the target user are determined based on the acquisition time and the context type of the current situation.
[0151] For example, the target sound effect determination unit 304 receives a determined scenario type, which includes the target user's current geographical location information and the current time at that location. Based on the target user's geographical location at the current time, the target user's scenario audio effect is determined. This helps to make the selection of the corresponding scenario audio effect more accurate when performing sound effect scenario processing on other users' audio. It also helps the target user to directly obtain clear and comfortable sound, thus improving the humanized service of sound effect processing.
[0152] (v) Sound processing unit 305; The audio processing unit 305 is used to perform audio scene processing on the audio to be processed of other users using contextual audio effects to obtain the contextual audio of other users.
[0153] In some embodiments, the contextual audio effects have corresponding first-channel pitch operators and second-channel pitch operators, which are used to calculate the pitch difference between the left and right channels. Other users' contextual audio includes first-channel audio and second-channel audio. The audio effects processing unit is used for: The first channel audio is obtained by convolving the audio from other users with the first channel pitch operator. The second-channel audio is obtained by convolving the audio from other users using the second-channel pitch operator.
[0154] For example, the audio processing unit 305 can be used to convolve the audio to be processed (mono audio) of other users with the first channel pitch operator and the second channel pitch operator respectively to obtain the first channel audio and the second channel audio, so that the background audio of other users becomes stereo audio, and the target user can perceive a stereo sound when the background audio of other users is played.
[0155] In some embodiments, the sound processing unit 305 is further configured to: In some embodiments, after obtaining the contextual audio from other users, the contextual audio effect also has corresponding first channel audio intensity and second channel audio intensity, and the device is further used for: The amplitude of the first channel audio is adjusted according to the first channel audio intensity to obtain the processed first channel intensity audio. The amplitude of the second channel audio is adjusted according to the intensity of the second channel audio to obtain the processed second channel intensity audio.
[0156] For example, the sound processing unit 305 can also obtain the first channel audio intensity and the second channel audio intensity from the contextual audio effects. The audio intensity can make the listener perceive the distance of the sound source. The amplitude of the first channel audio is adjusted according to the first channel audio intensity, and the amplitude of the second channel audio is adjusted according to the second channel audio intensity, so that the sound after the audio intensity adjustment is more three-dimensional, which is beneficial for the target user to obtain clear and comfortable sound.
[0157] In some embodiments, the contextual audio effect also includes reverberation information, and after obtaining contextual audio from other users, the device is further used to: The background audio of other users is reverb-processed based on the reverb information to obtain the background audio of other users with reverb.
[0158] In some embodiments, the reverberation information includes corresponding direct phonon information, early reflected phonon information, and late reflected phonon information. The device performs reverberation processing on the ambient audio of other users based on the reverberation information to obtain reverberated ambient audio of other users. The device also includes: Based on the direct phonon information, the contextual audio of other users is processed to obtain the direct phonon audio; Early reflection phonon information is used to process the contextual audio of other users into early reflection phonons to obtain early reflection audio. Based on the post-reflection phonon information, the contextual audio of other users is processed by post-reflection sound to obtain post-reflection sound audio. By superimposing the direct sound audio, early reflected audio, and late reflected audio, the reverberated background audio for other users is obtained.
[0159] In some embodiments, the device performs direct sound processing on the contextual audio of other users based on direct sound phonon information to obtain direct sound audio. The direct sound phonon information includes a direct sound operator. The device is also used to: The direct sound operator is used to process the contextual audio of other users to obtain direct sound audio.
[0160] In some embodiments, the device performs early reflection processing on the contextual audio of other users based on early reflection phonon information to obtain early reflection audio. The early reflection phonon information includes a frequency delay filter type and an early reflection operator. The device is also used to: The contextual audio of other users is delayed based on the type of frequency delay filter to obtain frequency-delayed audio. Early reflection audio is obtained by using an early reflection operator to process the frequency-delayed audio.
[0161] In some embodiments, post-reflection phonon processing is performed on the context audio of other users based on post-reflection phonon information to obtain post-reflection audio. The post-reflection phonon information includes frequency delay filter type, frequency filter type, phase delay filter type, and post-reflection operator, including: Based on the frequency delay filter type, the contextual audio of other users is subjected to audio delay processing to obtain frequency-delayed audio; The frequency-delayed audio is filtered according to the type of frequency filter to obtain the frequency-filtered audio. The frequency-filtered audio is phase-filtered according to the phase delay filter type to obtain the phase-delayed audio. The post-reflection sound operator is used to process the phase-delayed audio with post-reflection sound to obtain the post-reflection sound audio.
[0162] For example, the sound effects processing unit 305 can also retrieve reverb information from the contextual audio effects. The sound effects processing unit can perform reverb processing on the contextual audio of other users based on the reverb information. It can also perform reverb processing on the first channel intensity audio and the second channel intensity audio to make the reverb processed audio have a sense of space.
[0163] (vi) Sound effect playback unit 306; For example, the sound effect playback unit 306 is used to play contextual audio from other users so that the contextual audio listened to by the target user matches the current context of the target user.
[0164] In some embodiments, the target user's terminal plays contextual audio from other users. The contextual audio played from other users is determined based on the target user's current context. Therefore, the contextual audio listened to by the target user can match the target user's current context, which is beneficial for the target user to directly obtain clear and comfortable sound.
[0165] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0166] As can be seen from the above, the sound effect processing device in this embodiment acquires the audio to be processed from other users by the audio acquisition unit, where other users are users other than the target user; the scene type determination unit determines the scene type of the current scene of the target user; the mapping relationship acquisition unit acquires mapping relationship information, which includes the mapping relationship between the scene type and preset audio effects; the target sound effect determination unit determines the scene audio effect corresponding to the scene type based on the mapping relationship information; the sound effect processing unit performs sound effect scene processing on the audio to be processed from other users using scene audio effects to obtain the scene audio of other users; and the sound effect playback unit plays the scene audio of other users so that the scene audio listened to by the target user matches the scene of the target user.
[0167] Therefore, the embodiments of this application improve the user-friendliness of sound effect processing. The embodiments of this application also provide an electronic device, which can be a terminal, a server, or other similar devices. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0168] In some embodiments, the sound processing device may also be integrated into multiple electronic devices, such as multiple servers, with multiple servers implementing the sound processing method of this application.
[0169] In this embodiment, a mobile terminal will be used as an example for detailed description. For example, ... Figure 4 As shown, it illustrates a structural diagram of the mobile terminal involved in an embodiment of this application. Specifically: The mobile terminal may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will understand that... Figure 4 The SSS structure shown does not constitute a limitation on the mobile terminal and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Wherein: The processor 401 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data of the mobile terminal, thereby performing overall detection of the mobile terminal. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 401.
[0170] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the mobile terminal, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0171] The mobile terminal also includes a power supply 403 that supplies power to the various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0172] The mobile terminal may also include an input module 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, microphone, optical or trackball signal inputs related to user settings and function control.
[0173] The mobile terminal may also include a communication module 405. In some embodiments, the communication module 405 may include a wireless module, through which the mobile terminal can perform short-range wireless transmission, thereby providing users with wireless broadband internet access. For example, the communication module 405 can be used to help users send and receive emails, browse web pages, and access streaming media.
[0174] Although not shown, the mobile terminal may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the mobile terminal loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows: Get audio files to be processed from other users, where other users are users other than the target user; Determine the scenario type of the target user's current situation; Obtain mapping relationship information, which includes the mapping relationship between scene type and preset audio effects; Based on the mapping relationship information, determine the scene audio effects corresponding to the scene type; Use contextual audio effects to process the audio of other users, and obtain the contextual audio of other users. Play contextual audio from other users so that the contextual audio listened to by the target user matches the current context of the target user.
[0175] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0176] As can be seen from the above, the embodiments of this application can reduce the impact of the user's environment on the audio to be played. Therefore, this solution can meet the user's requirements for audio effects in different environments, making the audio to be played adaptable to the user's current environment.
[0177] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0178] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the sound effect processing methods provided in embodiments of this application. For example, the instructions can execute the following steps: Get audio files to be processed from other users, where other users are users other than the target user; Determine the scenario type of the target user's current situation; Obtain mapping relationship information, which includes the mapping relationship between scene type and preset audio effects; Based on the mapping relationship information, determine the scene audio effects corresponding to the scene type; Use contextual audio effects to process the audio of other users, and obtain the contextual audio of other users; Play contextual audio from other users so that the contextual audio listened to by the target user matches the current context of the target user.
[0179] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0180] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the above embodiments for performing audio effects processing on audio from other users based on the user's environment.
[0181] Since the instructions stored in the storage medium can execute the steps of any of the sound effect processing methods provided in the embodiments of this application, the beneficial effects that any of the sound effect processing methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0182] The above provides a detailed description of a sound effect processing method, sound effect processing device, terminal, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An audio processing method, characterized by, include: During voice calls or voice message interactions on social media software, the system can acquire audio from users other than the target user listening to the audio in real time. Based on the target user's current geographical location information and the current time at which the target user is located at the current geographical location, the scenario type of the target user's current situation is determined; the scenario type is classified according to the geographical location and the level of environmental noise within a specific time period; Obtain mapping information, including the mapping relationship between scene type and preset audio effects; Based on the mapping relationship information, the context audio effect corresponding to the context type is determined. The context audio effect is a dialogue sound effect that simulates an offline conversation between a speaker and a listener. The audio to be processed is processed using a first channel pitch operator to obtain a first channel audio, and the audio to be processed using a second channel pitch operator to obtain a second channel audio. The first channel pitch operator and the second channel pitch operator are obtained based on the pitch parameters of the left channel and the right channel carried by the scene audio effect. The first channel pitch operator and the second channel pitch operator are used to calculate the pitch difference between the left and right channels. The audio to be processed is a mono audio, and there is a pitch difference between the first channel audio and the second channel audio. The sound pressure is adjusted by adjusting the amplitude of the first channel audio according to the first channel audio intensity to obtain the processed first channel intensity audio. The sound pressure is adjusted by adjusting the amplitude of the second channel audio according to the second channel audio intensity to obtain the processed second channel intensity audio. The level of sound pressure simulates the distance between the sound source and the listener. The amplitude of the audio affects the sound intensity, and the sound intensity is reflected in the volume of the sound. For the audio after intensity adjustment, the audio is processed by direct sound based on direct phonon information to obtain direct sound audio, the audio is processed by early reflection phonon information to obtain early reflection audio, and the audio is processed by late reflection phonon information to obtain late reflection audio. The direct sound audio, the early reflected sound audio, and the late reflected sound audio are superimposed to obtain reverberant audio, which serves as the contextual audio for the other users; the first channel audio intensity, the second channel audio intensity, the direct phonon information, the early reflected phonon information, and the late reflected phonon information are information carried in the contextual audio effects. Play the contextual audio so that the contextual audio listened to by the target user matches the current context.
2. The sound effect processing method of claim 1, wherein, The process of processing the audio to be processed using a first channel pitch operator to obtain a first channel audio, and processing the audio to be processed using a second channel pitch operator to obtain a second channel audio, includes: The first channel pitch operator is used to perform convolution processing on the audio of the other users to obtain the first channel audio. The second channel audio is obtained by convolving the audio of the other users with the second channel pitch operator.
3. The sound effect processing method of claim 1, wherein, The direct phonon information includes a direct sound operator, and the step of performing direct sound processing on the audio based on the direct phonon information to obtain direct sound audio includes: The direct sound operator is used to perform direct sound processing on the audio to obtain the direct sound audio.
4. The sound effect processing method of claim 1, wherein, The early reflection phonon information includes the frequency delay filter type and the early reflection operator. The step of performing early reflection processing on the audio based on the early reflection phonon information to obtain early reflection audio includes: The audio is delayed according to the frequency delay filter type to obtain frequency-delayed audio. The early reflection operator is used to process the frequency-delayed audio with early reflections to obtain the early reflection audio.
5. The sound effect processing method of claim 1, wherein, The post-reflection phonon information includes frequency delay filter type, frequency filter type, phase delay filter type, and post-reflection operator. The step of performing post-reflection processing on the audio based on the post-reflection phonon information to obtain post-reflection audio includes: The audio is subjected to audio delay processing according to the frequency delay filter type to obtain frequency-delayed audio. The frequency-delayed audio is frequency-filtered according to the frequency filter type to obtain the frequency-filtered audio. The frequency-filtered audio is phase-filtered according to the phase delay filter type to obtain phase-delayed audio. The phase-delayed audio is processed by the post-reflection operator to obtain the post-reflection audio.
6. The sound effect processing method as described in claim 1, characterized in that, Before determining the scenario type of the target user's current situation based on the target user's current geographical location information and the current time at which the target user is located at the current geographical location, the method further includes: Obtain the current geographical location information of the target user; Obtain the current time at which the target user is located in the current geographical location.
7. A sound effect processing device, characterized in that, include: The audio acquisition unit is used to acquire, in real time, the audio to be processed from users other than the target user listening to the audio during voice calls or voice message interactions on social software. The scenario type determination unit is used to determine the scenario type of the current scenario of the target user based on the target user's current geographical location information and the current time when the target user is at the current geographical location; the scenario type is classified according to the geographical location and the level of environmental noise within a specific time period; the mapping relationship acquisition unit is used to acquire mapping relationship information including the mapping relationship between scenario type and preset audio effects; The target sound effect determination unit is used to determine the scene audio effect corresponding to the scene type based on the mapping relationship information; the scene audio effect is a dialogue sound effect that simulates an offline conversation between a speaker and a listener. The audio processing unit is used for: The audio to be processed is processed using a first channel pitch operator to obtain a first channel audio, and the audio to be processed using a second channel pitch operator to obtain a second channel audio. The first channel pitch operator and the second channel pitch operator are obtained based on the pitch parameters of the left channel and the right channel carried by the scene audio effect. The first channel pitch operator and the second channel pitch operator are used to calculate the pitch difference between the left and right channels. The audio to be processed is a mono audio, and there is a pitch difference between the first channel audio and the second channel audio. The sound pressure is adjusted by adjusting the amplitude of the first channel audio according to the first channel audio intensity to obtain the processed first channel intensity audio. The sound pressure is adjusted by adjusting the amplitude of the second channel audio according to the second channel audio intensity to obtain the processed second channel intensity audio. The level of sound pressure simulates the distance between the sound source and the listener. The amplitude of the audio affects the sound intensity, and the sound intensity is reflected in the volume of the sound. For the audio after intensity adjustment, the audio is processed by direct sound based on direct phonon information to obtain direct sound audio, the audio is processed by early reflection phonon information to obtain early reflection audio, and the audio is processed by late reflection phonon information to obtain late reflection audio. The direct sound audio, the early reflected sound audio, and the late reflected sound audio are superimposed to obtain reverberant audio, which serves as the contextual audio for the other users; the first channel audio intensity, the second channel audio intensity, the direct phonon information, the early reflected phonon information, and the late reflected phonon information are information carried in the contextual audio effects. The sound effect playback unit is used to play the contextual audio so that the contextual audio listened to by the target user matches the current context.
8. A terminal, characterized in that, The system includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps of the sound processing method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the sound processing method according to any one of claims 1 to 6.