Voice playing method and device, computer device and storage medium
By identifying the acoustic scene of the playback terminal and matching the target spatial sound effect template to process the voice, the problem of poor listening quality in traditional voice playback methods is solved, and a better listening experience is achieved.
Patent Information
- Application Number
- CN202110818669.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-20
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-07-20
AI Technical Summary
In traditional voice playback methods, the way of adding spatial sound effects to voice is fixed and cannot meet the auditory needs of the sound receiver, resulting in poor listening quality.
By acquiring the ambient sound of the playback terminal, identifying the current acoustic scene, matching the target spatial sound effect template, and processing the original speech according to the sound effect parameters to generate the target speech, it ensures that the voice interaction method meets the expectations of the sound receiver.
The listening quality of voice playback is improved, so that the target voice meets the auditory needs of the sound receiver.
Smart Images

Figure CN115705839B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing technology, and in particular to a voice playback method, apparatus, computer equipment, and storage medium. Background Art
[0002] With the development of signal processing technology, computer equipment can perform various forms of processing on voice signals. For example, it can add spatial sound effects to the voice being played, allowing users to hear more three-dimensional and spatially layered sounds.
[0003] In traditional technology, when adding spatial sound effects to the played voice, the played voice is usually processed in a fixed manner to obtain the target voice with the added spatial sound effects, resulting in the played voice not meeting the auditory needs of the sound receiver and poor listening quality. Summary of the Invention
[0004] Based on this, it is necessary to provide a voice playback method, device, computer equipment and storage medium that can improve the listening quality of the played voice in order to address the above technical problems.
[0005] A voice playback method, the method comprising: obtaining an original voice to be played, obtaining the ambient sound in the playback environment of a playback terminal of the original voice; performing scene recognition on the ambient sound to obtain the current acoustic scene of the playback terminal; obtaining a target spatial sound effect template that matches the current acoustic scene, the target spatial sound effect template matches a target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of a sound receiving object in the current acoustic scene; processing the original voice according to target sound effect parameters in the target spatial sound effect template to obtain a target voice, so as to play the target voice in the playback terminal.
[0006] A voice playback device, the device comprising: a voice acquisition module, used to acquire the original voice to be played, and acquire the ambient sound in the playback environment where the playback terminal of the original voice is currently located; a scene recognition module, used to perform scene recognition on the ambient sound, and obtain the current acoustic scene where the playback terminal is located; a template acquisition module, used to acquire a target space sound effect template that matches the current acoustic scene, the target space sound effect template matches a target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene; a voice processing module, used to process the original voice according to the target sound effect parameters in the target space sound effect template to obtain the target voice, so as to play the target voice in the playback terminal.
[0007] In some embodiments, the target spatial sound effect template includes a target sound orientation parameter sequence corresponding to the dynamic interactive position relationship; the speech processing module is also used to divide the original speech into speech segments with the number of parameters in the target sound orientation parameter sequence; according to the order of the speech segments in the original speech, the target sound orientation parameters corresponding to the speech segments in the target sound orientation parameter sequence are determined; the speech segments are processed according to the target sound orientation parameters corresponding to the speech segments to obtain processed speech segments, and each of the processed speech segments forms the target speech in the speech order.
[0008] In some embodiments, the candidate spatial sound effect template set includes candidate spatial sound effect templates with fixed voice interaction position relationships; the template acquisition module is also used to select, from the candidate spatial sound effect template set, a candidate spatial sound effect template with fixed voice interaction position relationship when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, as a target spatial sound effect template matching the current acoustic scene.
[0009] In some embodiments, the expected voice interaction position relationship includes an expected voice interaction distance; the template acquisition module is also used to, when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, select from the candidate spatial sound effect template set a candidate spatial sound effect template whose sound distance corresponding to the sound effect parameters is fixed and whose sound distance matches the expected voice interaction distance, as a target spatial sound effect template matching the current acoustic scene.
[0010] In some embodiments, the expected voice interaction mode includes an expected voice interaction position relationship between a sound receiving object and a sound emitting object; the target sound effect parameters include position relationship sound effect parameters that match the expected voice interaction position relationship; the voice processing module is also used to use the position relationship sound effect parameters in the target spatial sound effect template to perform voice processing on the original voice to obtain a target voice, so that the target voice matches the expected voice interaction position relationship.
[0011] In some embodiments, the expected voice interaction position relationship includes an expected voice interaction distance and an expected interaction direction; the position relationship sound effect parameters include direction-related sound effect parameters and distance-related sound effect parameters; the voice processing module is also used to use the direction-related sound effect parameters to process the direction of the original voice, and use the distance-related sound effect parameters to process the sound pressure of the original voice to obtain the target voice; so that the direction of the target voice matches the expected interaction direction, and the sound pressure of the target voice matches the expected voice interaction distance.
[0012] In some embodiments, the target spatial sound effect template also matches the target voice interaction effect, and the target voice interaction effect is the expected voice interaction effect of the sound receiving object in the current acoustic scene. The target sound effect parameters include voice effect adjustment parameters that match the target voice interaction effect; the voice processing module is also used to use the positional relationship sound effect parameters in the target spatial sound effect template and the voice effect adjustment parameters to process the original voice to obtain the target voice.
[0013] In some embodiments, the scene recognition module is also used to obtain multiple sound sub-segments corresponding to the ambient sound, perform feature extraction on the sound sub-segments, and obtain sub-segment features; obtain the segment acoustic scenes corresponding to the sound sub-segments based on the sub-segment feature recognition; perform statistics on the segment acoustic scenes corresponding to the sound sub-segments to obtain the number of scenes corresponding to each segment acoustic scene; and select the segment acoustic scene with the largest number of scenes as the current acoustic scene of the playback terminal.
[0014] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program: obtaining an original voice to be played, and obtaining the ambient sound in the playback environment of the playback terminal of the original voice; performing scene recognition on the ambient sound to obtain the current acoustic scene of the playback terminal; obtaining a target spatial sound effect template that matches the current acoustic scene, wherein the target spatial sound effect template matches a target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene; processing the original voice according to the target sound effect parameters in the target spatial sound effect template to obtain the target voice, so as to play the target voice in the playback terminal.
[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps: obtaining an original voice to be played, and obtaining the ambient sound in the playback environment of a playback terminal of the original voice; performing scene recognition on the ambient sound to obtain the current acoustic scene of the playback terminal; obtaining a target spatial sound effect template that matches the current acoustic scene, wherein the target spatial sound effect template matches a target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of a sound receiving object in the current acoustic scene; processing the original voice according to target sound effect parameters in the target spatial sound effect template to obtain a target voice, so as to play the target voice in the playback terminal.
[0016] The above-mentioned voice playback method, device, computer equipment and storage medium obtain the original voice to be played, obtain the ambient sound in the playback environment where the playback terminal of the original voice is currently located, perform scene recognition on the ambient sound, identify the current acoustic scene where the playback terminal is located, obtain the target space sound effect template that matches the current acoustic scene, and process the original voice according to the target sound effect parameters in the target space sound effect template to obtain the target voice, so as to play the target voice in the playback terminal. Since the target voice is obtained by processing the original voice according to the target sound effect parameters in the target space sound effect template, the target space sound effect template matches the target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene, the obtained target voice can meet the auditory needs of the sound receiving object and improve the listening quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A diagram showing an application environment of a voice playback method in some embodiments;
[0018] Figure 2 Schematic diagram of the flow of the voice playback method in some embodiments;
[0019] Figure 3 is a flow chart of processing the original speech to obtain the target speech in some embodiments;
[0020] Figure 4 A schematic diagram of generating two-channel stereo speech in some embodiments;
[0021] Figure 5 is a schematic structural diagram of a reverberator in some embodiments;
[0022] Figure 6 is a schematic diagram of the structure of an acoustic scene recognition model in some embodiments;
[0023] Figure 7 Schematic diagram of the flow of the voice playback method in some specific embodiments;
[0024] Figure 8 A schematic diagram of displaying a prompt box for selecting a spatial sound effect template in a conversation interface in some embodiments;
[0025] Figure 9 is a structural block diagram of a voice playback device in some embodiments;
[0026] Figure 10 1 is a diagram of the internal structure of a computer device in some embodiments. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0028] The voice playback method provided in this application can be applied to Figure 1 In the application environment shown. Among them, the first terminal 102 and the second terminal 106 communicate with the server 104 through the network. The first terminal 102 can be a terminal corresponding to the sound emitting object, and the second terminal 106 can be a terminal corresponding to the sound receiving object. The second terminal 106 is used for voice playback, so it can be called a playback terminal. The first terminal 102 collects the sound of the sound emitting object to obtain the original voice, and sends the original voice to the second terminal 106 through the server 104. The second terminal 106 further collects the ambient sound in the current playback environment. Combined with the ambient sound, the playback terminal or server can process the original voice to obtain the target voice.
[0029] Taking the playback terminal's processing of the original voice as an example, the ambient sound is subjected to scene recognition to obtain the current acoustic scene in which the playback terminal is located, and the target spatial sound effect template matching the current acoustic scene is further obtained. According to the target sound effect parameters in the target spatial sound effect template, the original voice is processed to obtain the target voice. Among them, the target spatial sound effect template matches the target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene. The target voice processed by the playback terminal can be a stereo voice with spatial sound effects added, and the playback terminal can further play the target voice through headphones or a combination of two or more speakers.
[0030] The first terminal 102 and the second terminal 106 may be, but are not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.
[0031] In some embodiments, as Figure 2 As shown, a voice playing method is provided, which is applied to Figure 1 Taking the second terminal 106, i.e., the playback terminal, as an example, the following steps are included:
[0032] Step 202: Acquire the original voice to be played, and acquire the ambient sound in the current playback environment of the playback terminal of the original voice.
[0033] Among them, the original voice to be played refers to the voice that the playback terminal needs to play. The original voice to be played can be various types of voices, including voices in audio and video, voices in online games, or human voices during calls, etc. The playback environment in which the playback terminal of the original voice is currently located refers to the environment corresponding to the physical place where the playback terminal is currently located. For example, if the playback terminal is currently in the woods, then the environment corresponding to the woods is the playback environment in which the playback terminal is currently located. For example, if the playback terminal is currently in a bar, then the environment corresponding to the bar is the playback environment in which the playback terminal is currently located. Ambient sounds can be all or most of the sounds in the environment that can represent the environment. For example, the ambient sounds corresponding to the woods can be human voices, bird sounds, flowing water sounds, etc. in the woods. The ambient sounds corresponding to the bar can be human voices, music sounds, object collision sounds, etc. in the bar.
[0034] Specifically, the playback terminal may obtain the original voice to be played locally or from a server, and when the original voice needs to be played, the ambient sound in the playback environment currently located by the playback terminal is collected.
[0035] In some embodiments, the playback terminal is installed with an application that can play voice, which can be a real-time call application, so that the playback terminal obtains the real-time voice of the call partner during the real-time call to obtain the original voice to be played, and collects the ambient sound in the current playback environment, and then processes the original voice in combination with the ambient sound to obtain the target voice.
[0036] In some embodiments, the application program installed on the playback terminal that can play voice can also be an audio and video entertainment application, so that when the playback terminal plays audio or video, it determines the audio in the played audio or video as the original voice to be played, and collects the ambient sound in the current playback environment, and then processes the original voice in combination with the ambient sound to obtain the target voice.
[0037] In some embodiments, the application installed on the playback terminal that can play voice can also be an instant messaging application. This application can receive voice messages. After receiving the voice message, it can display the playback controls of the message, and when the user triggers the playback control, it determines the voice in the message as the original voice to be played, and collects the ambient sound in the current playback environment, and then processes the original voice in combination with the ambient sound to obtain the target voice, and plays the processed target voice.
[0038] Step 204: Perform scene recognition on the ambient sound to obtain the current acoustic scene where the playback terminal is located.
[0039] Among them, the sound scene refers to a scene divided according to sound, and the acoustic scene is divided according to the environment in which the playback terminal is located. In some embodiments, since the sounds in different places are usually different, the acoustic scene can be divided according to at least one of the type of place, the size of the sound, or the atmosphere of the environment. For example, the acoustic scene can be divided into at least one of a subway scene, a supermarket scene, a bar scene, or a seaside scene. In some other embodiments, considering that the sounds in different places usually have certain commonalities, the acoustic scene can be delineated according to the actual characteristics of the sound. For example, considering that the sounds in places such as bars and vegetable markets are often noisy and the sounds in the woods are quiet, the acoustic scene can be divided into a noisy scene, a quiet scene, and an ordinary scene between the noisy and quiet scenes. The ordinary scene is neither a noisy scene nor a quiet scene.
[0040] Specifically, the playback terminal can use a machine learning-based method to perform scene recognition on the ambient sound. After extracting audio features from the ambient sound on the other end, the audio features are input into the trained acoustic scene recognition model. The ambient sound is then recognized through the acoustic scene recognition model. The obtained scene recognition result represents the current acoustic scene, so that the playback terminal can obtain the specific current acoustic scene it is in.
[0041] The audio feature can be a power spectrum of the environmental sound or can be a mel-frequency cepstral coefficient of the environmental sound. The acoustic scene recognition model refers to a machine learning model that can be used for acoustic scene recognition, i.e., a process of classifying an acoustic scene. The scene recognition model can specifically be a classification model based on convolution operation, such as a recurrent neural network (RNN), a convolutional neural network (CNN), a long short-term memory (LSTM), a BiLSTM, a gate recurrent unit (GRU), a BiGRU, and the like. Among them, the CNN is a kind of feedforward neural network containing convolution calculation and having a deep structure; the RNN is a kind of recursive neural network taking sequence data as input, performing recursion in the evolution direction of the sequence, and connecting all nodes (recurrent units) in a chain; the LSTM is a kind of time recurrent neural network, which is specially designed to solve the long-term dependence problem of general RNN (recurrent neural network), wherein all RNNs have a chain form of repeated neural network modules, and the forward LSTM and the backward LSTM are combined into a BiLSTM; the GRU is a kind of RNN, and like the LSTM, it is also proposed to solve the problems of long-term memory and gradient in back propagation, and the forward GRU and the backward GRU are combined into a BiGRU.
[0042] In some embodiments, the acoustic scene recognition model can be deployed locally on the playback terminal, and the playback terminal can obtain the acoustic scene recognition model from a local memory, input the extracted audio feature into the acoustic scene recognition model, so as to quickly obtain the acoustic scene recognition result. In other embodiments, in order to save the storage space of the playback terminal, the acoustic scene recognition model can also be deployed on a server, and the playback terminal can send a scene recognition request carrying the audio feature to the server after obtaining the audio feature, the server inputs the audio feature in the scene recognition request into the acoustic scene recognition model, and returns the obtained scene recognition result to the playback terminal.
[0043] In step 206, a target spatial sound effect template matching the current acoustic scene is obtained, the target spatial sound effect template matches a target voice interaction mode, and the target voice interaction mode is an expected voice interaction mode of a sound receiving object in the current acoustic scene.
[0044] Among them, spatial sound effect refers to the spatial sound effect to be achieved. Spatial sound effect is a sound that has been processed through certain audio technology to allow users to hear more three-dimensional and spatially layered sounds. For example, through headphones or a combination of two or more speakers, the playback restores the actual auditory scene, allowing the listener (that is, the sound receiving object) to clearly identify the direction, distance and movement trajectory of different acoustic objects. It can also make the listener feel surrounded by the sound in all directions, giving the listener an immersive auditory experience as if they are in the actual environment.
[0045] The spatial sound effect template is a template used to perform spatial sound effect processing on speech, and includes various sound effect parameters for performing spatial sound effect processing. The spatial sound effect template can be configured by professionals. Different spatial sound effect templates match different speech interaction methods. The speech interaction method refers to the way in which the sound receiving object receives the sound of the sound emitting object. The speech interaction method includes at least one of the positional relationship between the interacting parties and the size of the sound. Matching the spatial sound effect template with the speech interaction method means that the matched spatial sound effect template can make the played speech consistent with the interaction method of the matched speech interaction method. For example, assuming that the speech interaction method is that the spatial distance between the sound emitting object and the sound receiving object is sometimes far and sometimes near, then after the sound is processed by the spatial sound effect template matched with it, the effect of the played sound being sometimes far and sometimes near can be achieved.
[0046] The sound receiving object refers to the object that receives the sound, and the sound receiving object can be the user corresponding to the playback terminal. The sound emitting object can be the object that emits the sound. In voice communication, the sound emitting object is the communication object during the voice communication process, while in audio and video entertainment, the sound emitting object can be the computer device that emits the sound. For example, when playing music through a playback terminal, the sound emitting object can be the playback terminal. The expected voice interaction mode of the sound receiving object refers to the voice interaction mode expected by the sound receiving object, including at least one of the expected positional relationship between the interacting parties and the expected sound volume.
[0047] Considering that the recipients of sound have different auditory experience requirements in different environments, the recipients' desired voice interaction methods vary in different acoustic scenarios. For example, in a noisy environment, such as a bar, the recipient prefers close-to-ear or very close-range voice interaction to avoid interference from ambient noise. In a quiet environment, such as a grove, the recipient prefers a more natural and casual voice interaction, hoping that the other party they are talking to will also move around freely and naturally. In an open environment, such as a stadium, the recipient prefers to hear sound with a slight reverberation effect to better match the actual environment.
[0048] Specifically, for each type of acoustic scene, one or more spatial sound effect templates can be configured, that is, for each type of acoustic scene, a correspondence between the acoustic scene and the spatial sound effect model is established. When the playback terminal identifies the current acoustic scene in which the playback terminal is located, it can select a spatial sound effect template that matches the current acoustic scene from multiple candidate spatial sound effect templates according to the correspondence as the target spatial sound effect template. Since the target spatial sound effect template matches the current acoustic scene, the target spatial sound effect template matches the target voice interaction mode. The target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene, that is, the voice interaction mode expected by the sound receiving object in the current acoustic scene.
[0049] In some embodiments, the playback terminal stores the correspondence between the acoustic scene identifier and the spatial sound effect model identifier. After identifying the current acoustic scene, the playback terminal obtains the target spatial sound effect template that matches the current acoustic scene based on the matching relationship between the identifiers. In other embodiments, the correspondence between the acoustic scene identifier and the spatial sound effect model identifier can also be stored in a server, so that after identifying the current acoustic scene, the playback terminal sends the acoustic scene identifier to the server, and the server obtains the target spatial sound effect template that matches the current acoustic scene based on the matching relationship between the identifiers and returns it to the playback terminal.
[0050] In some embodiments, when there are multiple spatial sound effect templates that match the current acoustic scene, the playback terminal can randomly select one of the matching spatial sound effect templates to determine as the target spatial sound effect model. In some other embodiments, when there are multiple spatial sound effect templates that match the current acoustic scene, the playback terminal can also display a list of matching spatial sound effect templates, which displays the names and introductions of each matching spatial sound effect model, and provides a selection control for the user to select the spatial sound effect template, and determines the spatial sound effect template selected by the user as the target spatial sound effect template.
[0051] In some embodiments, a plurality of candidate spatial sound effect templates can be pre-configured, including at least one of an ear-to-ear communication template, a strolling communication template, a lecture-style communication template, a surprise communication template, a surround template, or a fly-in-and-out template, wherein the ear-to-ear communication template refers to the sound emitting object being close to the ear of the sound receiving object for communication; the strolling communication template refers to the sound emitting object being within a certain distance range and communicating with the sound receiving object in slow motion according to a random or set motion trajectory; the lecture-style template refers to the sound emitting object being at a medium or long distance, with a loud sound and accompanied by a certain reverberation effect; the surprise communication template refers to the sound emitting object being at a medium or long distance, with a loud sound and accompanied by a certain reverberation effect; the surprise communication template refers to the sound emitting object being at a medium or long distance and communicating with the sound receiving object in slow motion. The communication template refers to the position of the sound emitting object is not fixed and the movement trajectory is random. For example: the first sentence appears in front of the left of the sound receiving object, the second sentence appears behind the sound receiving object, and the next sentence is close to the ear of the sound receiving object, giving people a surprising auditory experience; the surround template refers to the sound emitting object maintaining a certain distance from the sound receiving object, and communicating by rotating 360 degrees around the sound receiving object in the horizontal direction; the fly-in and fly-out template refers to the sound emitting object approaching the position of the sound receiving object from a distance at a high speed, or moving away from a position close to the sound receiving object to a distance at a high speed.
[0052] In some embodiments, when the candidate spatial sound effect templates include ear-to-ear communication template, strolling communication template, lecture-style template, surprise communication template, surround template, and fly-in-fly-out template, the acoustic scenes can be divided into three categories: noisy scenes, quiet scenes, and ordinary scenes, among which the noisy scenes correspond to the ear-to-ear communication template, the quiet scenes correspond to the surprise communication template, surround template, and fly-in-fly-out template, and the ordinary scenes correspond to the strolling communication template and lecture-style template.
[0053] In some embodiments, the target spatial sound effect template is also matched with a target voice interaction effect, where the target voice interaction effect is the expected voice interaction effect of the sound receiving object in the current acoustic scene.
[0054] Step 208 : Process the original speech according to the target sound effect parameters in the target spatial sound effect template to obtain the target speech, so as to play the target speech in the playback terminal.
[0055] Among them, sound effect parameters refer to parameters used to process sound effects for speech. Sound effect parameters include parameters corresponding to the speech interaction mode, specifically position-related sound effect parameters. Position-related sound effect parameters include direction-related sound effect parameters and distance-related sound effect parameters. Sound direction-related sound effect parameters can be, for example, the direction of the sound, and distance-related sound effect parameters can be, for example, the distance. Sound effect parameters also include parameters corresponding to speech interaction effects, specifically parameters for adjusting speech effects, such as reverberation parameters.
[0056] It is understandable that the position relationship sound effect parameters included in the target sound effect parameters match the expected voice interaction mode, and the voice effect adjustment parameters included in the target sound effect parameters match the expected voice interaction effect.
[0057] Specifically, after the playback terminal obtains the target space sound effect template that matches the current acoustic scene, it processes the original voice according to the target sound effect parameters in the target space sound effect template to obtain the target voice, and plays the target voice.
[0058] In some embodiments, processing the original speech to obtain the target speech according to the target sound effect parameters in the target spatial sound effect template can specifically be: performing speech processing on the original speech according to the positional relationship sound effect parameters in the target sound effect parameters to obtain the target speech, so that the target speech matches the expected speech interaction method.
[0059] In some embodiments, processing the original speech to obtain the target speech according to the target sound effect parameters in the target space sound effect template can specifically be: performing speech processing on the original speech according to the speech effect adjustment parameters in the target sound effect parameters to obtain the target speech, so that the target speech matches the expected speech interaction effect.
[0060] In the above-mentioned voice playback method, the original voice to be played is obtained, the ambient sound in the playback environment where the playback terminal of the original voice is currently located is obtained, scene recognition is performed on the ambient sound, the current acoustic scene in which the playback terminal is located is identified, and the target space sound effect template matching the current acoustic scene is obtained. According to the target sound effect parameters in the target space sound effect template, the original voice is processed to obtain the target voice, so as to play the target voice in the playback terminal. Since the target voice is obtained by processing the original voice according to the target sound effect parameters in the target space sound effect template, the target space sound effect template matches the target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene, the obtained target voice can meet the auditory needs of the sound receiving object and improve the quality of listening.
[0061] In some embodiments, the expected voice interaction mode includes an expected voice interaction position relationship between a sound receiving object and a sound emitting object; the step of obtaining a target spatial sound effect template that matches the current acoustic scene includes: obtaining a set of candidate spatial sound effect templates; the set of candidate spatial sound effect templates includes multiple candidate spatial sound effect templates corresponding to different voice interaction position relationships; from the set of candidate spatial sound effect templates, a target spatial sound effect template that matches the current acoustic scene is selected, wherein the voice interaction position relationship corresponding to the target spatial sound effect template matches the expected voice interaction position relationship corresponding to the current acoustic scene.
[0062] The expected voice interaction position relationship between the sound receiving object and the sound emitting object refers to the interaction position relationship between the sound emitting object and the sound receiving object that the sound receiving object expects. There are two types of interaction position relationships: no change in position relationship and change in position relationship. No change in position relationship means that the distance between the sound receiving object and the sound emitting object is fixed. No change in position relationship includes close distance and long distance. Close distance means that the sound emitting object is within a certain range close to the sound receiving object, and long distance means that the sound emitting object is within a certain range away from the sound receiving object. Change in position relationship means that the position relationship between the sound receiving object and the sound emitting object changes over time, including at least one of the following: the speed of the sound emitting object changes over time, and the direction of movement changes over time.
[0063] In this embodiment, multiple candidate spatial sound effect templates can be pre-configured. These candidate spatial sound effect templates constitute a candidate spatial sound effect template set. Each candidate spatial sound effect template corresponds to a different voice interaction position relationship. After obtaining the current acoustic scene, the playback terminal selects a target spatial sound effect template that matches the current acoustic scene from the candidate spatial sound effect template set. Since the target spatial sound effect template matches the current acoustic scene, the voice interaction position relationship corresponding to the target spatial sound effect template matches the expected voice interaction position relationship corresponding to the current acoustic scene.
[0064] For example, assuming that the current acoustic scene is a noisy scene, the expected voice interaction position relationship of the sound receiving object in this scene is close-range interaction, so as to avoid interference of ambient noise on the listening process. Therefore, the playback terminal can select a close-range communication template from the candidate spatial sound effect template set. The voice interaction position relationship corresponding to the close-range communication template is close-range interaction, which matches the expected voice interaction position relationship.
[0065] In this embodiment, by selecting a candidate spatial sound effect template whose voice interaction position relationship matches the expected voice interaction position relationship corresponding to the current acoustic scene from a set of candidate spatial sound effect templates including candidate spatial sound effect templates corresponding to multiple different voice interaction position relationships as the target spatial sound effect template, a spatial sound effect template suitable for the current acoustic scene can be selected, thereby obtaining a played voice with high listening quality. At the same time, since multiple candidate spatial sound effect templates corresponding to different voice interaction position relationships are pre-configured, it can be suitable for multiple different acoustic scenes, thereby expanding the scope of application.
[0066] In some embodiments, the candidate spatial sound effect template set includes candidate spatial sound effect templates with changes in voice interaction position relationships; selecting a target spatial sound effect template that matches the current acoustic scene from the candidate spatial sound effect template set includes: when the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, selecting a candidate spatial sound effect template with changes in voice interaction position relationships from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene.
[0067] The candidate spatial sound effect template set includes candidate spatial sound effect templates for changes in voice interaction position relationships. When the candidate spatial sound effect templates for changes in voice interaction position relationships are used to process the original speech, the voice interaction position relationships corresponding to the target speech obtained by the speech interaction method change over time. For example, the candidate spatial sound effect templates for changes in voice interaction position relationships can be any of the following: a surprise communication template, a surround template, and a fly-in / fly-out template.
[0068] The current voice interaction position relationship corresponding to the current acoustic scene refers to the expected voice interaction method of the sound receiving object in the current acoustic scene. The expected voice interaction method of the sound receiving object is different in different acoustic scenes. A dynamic interaction position relationship means that the position relationship changes dynamically over time. In a dynamic interaction position relationship, the sound emitting object moves according to a set or random motion trajectory. For example, the sound emitting object rotates 360 degrees, or appears in front of the left of the listener for the previous sentence, behind the listener for the next sentence, and close to the listener's ear for the next sentence. Or it approaches the position of the sound receiving object from a distance, or moves away from the position close to the sound receiving object.
[0069] In this embodiment, since the expected voice interaction methods of the sound receiving objects in different acoustic scenes are different, after the playback terminal identifies the current acoustic scene, it is equivalent to knowing the current voice interaction position relationship corresponding to the current acoustic scene. When the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, the candidate spatial sound effect template with a changed voice interaction position relationship is selected from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene.
[0070] In some embodiments, when the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, the current acoustic scene may be a quiet scene.
[0071] In some embodiments, when there are multiple candidate spatial sound effect templates for changes in voice interaction position relationships, the playback terminal may display a list of candidate spatial sound effect templates for changes in each voice interaction position relationship for the user to select, or randomly select a candidate spatial sound effect template for changes in voice interaction position relationships as the target spatial sound effect template.
[0072] In the above embodiment, when the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, the playback terminal can select a candidate spatial sound effect template with a changing voice interaction position relationship from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene, thereby selecting a spatial sound effect template suitable for the current acoustic scene.
[0073] In some embodiments, the target spatial sound template includes a target sound orientation parameter sequence corresponding to the dynamic interaction position relationship; Figure 3 As shown, according to the target sound effect parameters in the target space sound effect template, the original speech is processed to obtain the target speech, including:
[0074] Step 302: Segment the original speech into speech segments having the same number of parameters as in the target sound position parameter sequence.
[0075] Among them, the sound orientation parameter refers to a parameter related to the sound orientation, and the sound orientation parameter can be a specific sound orientation. The sound orientation parameter corresponding to the dynamic interactive position relationship changes dynamically over time, and multiple sound orientation parameters that change over time form a target sound orientation parameter sequence in chronological order. The number of parameters in the target sound orientation parameter sequence refers to the number of sound orientation parameters corresponding to different times in the target sound orientation parameter sequence. For example, if the target sound orientation parameter sequence includes sound orientation parameters corresponding to 6 different times, then the number of parameters in the target sound orientation parameter sequence is 6.
[0076] Specifically, the playback terminal segments the original speech according to the number of parameters in the target sound orientation parameter sequence to obtain speech segments with the same number of parameters as in the target sound orientation parameter sequence. For example, assuming that the number of parameters in the target sound orientation parameter sequence is 6, the original speech is segmented to obtain 6 speech segments. The segmentation method can be random segmentation, or the original speech can be segmented into equal lengths according to the number of parameters in the target sound orientation parameter sequence. For example, if the original speech is 6 minutes and needs to be segmented into 6 speech segments, the original speech can be segmented into 6 speech segments with a duration of 1 minute.
[0077] Step 304 : determining the target sound position parameters corresponding to the speech segments in the target sound position parameter sequence according to the order of the speech segments in the original speech.
[0078] The order of the speech segments in the original speech refers to the order in which the multiple speech segments are sorted according to the time sequence in the original speech.
[0079] Specifically, since the sound orientation parameters in the target sound orientation parameter sequence are arranged in chronological order, each sound orientation parameter has a corresponding order, and the number of speech segments is the same as the number of parameters in the target sound orientation parameter sequence, each speech segment can correspond to a sound orientation parameter, so that the playback terminal can determine the sound orientation parameters in the target sound orientation parameter sequence whose order is the same as the order of the speech segments in the original speech as the target sound orientation parameters corresponding to the speech segment.
[0080] Step 306 : Process the speech segments according to the target sound orientation parameters corresponding to the speech segments to obtain processed speech segments. The processed speech segments form the target speech in the speech sequence.
[0081] Specifically, the playback terminal processes the orientation of the voice segment according to the target sound orientation parameters corresponding to the voice segment, generates stereo audio corresponding to the voice segment as the processed voice segment, and each processed voice segment forms the target audio in the order of the audio.
[0082] In some embodiments, the playback terminal can process the orientation of the speech segment based on an HRTF (Head-Related Transfer Function) to generate stereo audio corresponding to the speech segment. HRTF (Head-Related Transfer Function) is short for Head-Related Transfer Function. HRTF is a sound source position function that integrates ITD (Time Delay Difference), IID (Instrument Inertia Difference), and the body acoustic reflection spectrum characteristics, that is, the response of the sound transmission path. The time-domain stimulus response data corresponding to the HRTF transfer function is HRIR (Head Related Impulse Response). The most commonly used HRIR data sets include the CIPIC (Center for Image Processing and Integrated Computing) dataset and the MIT (Massachusetts Institute of Technology) dataset. For example, the CIPIC dataset collects time-domain measurement data of binaural listening signals from 45 measurement subjects, each at 25 different horizontal and 50 different vertical positions, for a total of 1,250 positions.
[0083] HRTF-based stereo generation is to convolve the original mono input signal u(n) with the target HRIR data h(n), and the output is a two-channel stereo signal y(n), refer to the following formula (1):
[0084]
[0085] When determining the HRIR data, the target orientation parameter may be matched with the orientation in the HRIR data set, and the HRIR data corresponding to the matching orientation may be determined as the target HRIR data.
[0086] It should be noted that h(n) is divided into HRIR data of left channel and right channel, so the generated y(n) also corresponds to the left channel and right channel signal results, such as Figure 4 Reference Figure 4 , the original speech and the left channel HRIR data are convolved to obtain the left channel speech signal, and the original speech and the right channel HRIR data are convolved to obtain the right channel speech signal.
[0087] In the above embodiment, by dividing the original speech into multiple speech segments, each speech segment is processed by different target sound orientation parameters in the target sound orientation parameter sequence, so that stereo speech of different orientations can be generated for each different speech segment, thereby obtaining a target speech that matches the target spatial sound effect template and whose interactive position relationship dynamically changes.
[0088] In some embodiments, the candidate spatial sound effect template set includes candidate spatial sound effect templates with fixed voice interaction position relationships; selecting a target spatial sound effect template that matches the current acoustic scene from the candidate spatial sound effect template set includes: when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, selecting a candidate spatial sound effect template with a fixed voice interaction position relationship from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene.
[0089] Among them, the candidate spatial sound effect template set includes candidate spatial sound effect templates with fixed voice interaction position relationships. When the candidate spatial sound effect templates with fixed voice interaction position relationships are used to process the original voice, the voice interaction position relationship corresponding to the voice interaction mode of the target voice obtained is fixed. The voice interaction position relationship can be, for example, the interaction distance. The interaction distance corresponding to the candidate spatial sound effect templates with fixed voice interaction position relationships can be close or long distance. The candidate spatial sound effect templates with fixed voice interaction position relationships can be any one of the ear-touching communication templates, the strolling mobile communication template, and the lecture template. The interaction distance corresponding to the ear-touching communication template is close distance, while the strolling mobile communication template and the lecture template can be long distance.
[0090] The current voice interaction position relationship corresponding to the current acoustic scene refers to the expected voice interaction method for the sound receiving object in the current acoustic scene. The expected voice interaction method for the sound receiving object varies in different acoustic scenes. A fixed interaction position relationship refers to a fixed position relationship. In a fixed interaction position relationship, the interaction distance between the sound emitting object and the sound receiving object is fixed.
[0091] In this embodiment, since the expected voice interaction methods of the sound receiving objects in different acoustic scenes are different, after the playback terminal identifies the current acoustic scene, it is equivalent to knowing the current voice interaction position relationship corresponding to the current acoustic scene. When the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, a candidate spatial sound effect template with a fixed voice interaction position relationship is selected from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene.
[0092] In the above embodiment, when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, the playback terminal can select a candidate spatial sound effect template with a fixed voice interaction position relationship from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene, thereby selecting a spatial sound effect template suitable for the current acoustic scene.
[0093] In some embodiments, the expected voice interaction position relationship includes an expected voice interaction distance; when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, selecting a candidate spatial sound effect template with a fixed voice interaction position relationship from the candidate spatial sound effect template set as a target spatial sound effect template that matches the current acoustic scene includes: when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, selecting a candidate spatial sound effect template with a fixed sound distance corresponding to the sound effect parameters and a sound distance that matches the expected voice interaction distance from the candidate spatial sound effect template set as a target spatial sound effect template that matches the current acoustic scene.
[0094] The expected voice interaction distance refers to the voice interaction distance expected by the sound receiving object in the current acoustic scene. In some embodiments, when the current acoustic scene is a noisy scene, the sound receiving object expects close communication to avoid interference from ambient noise on the listening process. That is, the expected voice interaction distance in the current acoustic scene is a short distance. The short distance can be the distance between the sound receiving object and the sound emitting object within a preset distance range. The preset distance range can be set as needed.
[0095] The fixed sound distance corresponding to the sound effect parameter means that the sound distance corresponding to each time in the sound effect parameter is the same. The matching of the sound distance corresponding to the sound effect parameter with the expected voice interaction distance means that the sound distance value corresponding to the sound effect parameter is consistent with the expected voice interaction distance.
[0096] Specifically, in different acoustic scenes, the expected voice interaction distance of the sound receiving object is different. In some acoustic scenes, the expected voice interaction distance of the sound receiving object may be a short distance, while in other acoustic scenes, the expected voice interaction distance of the sound receiving object may be a long distance or a medium-long distance. The long distance here may be that the distance between the sound receiving object and the sound emitting object is greater than a preset distance threshold. The preset distance threshold can be set as needed. The medium-long distance may be a distance between a short distance and a long distance. In this embodiment, since the expected voice interaction distance of the sound receiving object is different in different acoustic scenes, when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, the playback terminal can select a candidate spatial sound effect template from the candidate spatial sound effect template set, whose sound distance corresponding to the sound effect parameters is fixed and whose sound distance matches the expected voice interaction distance, as the target spatial sound effect template that matches the current acoustic scene.
[0097] In some embodiments, when the current acoustic scene is a noisy scene and the sound receiving object expects close communication, the playback terminal can select an ear-to-ear communication template from the candidate spatial sound effect template set as the target spatial sound effect template.
[0098] In the above embodiment, a candidate spatial sound effect template whose sound distance corresponding to the sound effect parameters is fixed and whose sound distance matches the expected voice interaction distance is selected from the set of candidate spatial sound effect templates as the target spatial sound effect template that matches the current acoustic scene. Since the selected target spatial sound effect template not only has a fixed sound distance but also matches the expected voice interaction distance, it is better suitable for the current acoustic scene.
[0099] In some embodiments, the expected voice interaction mode includes the expected voice interaction position relationship between the sound receiving object and the sound emitting object; the target sound effect parameters include position relationship sound effect parameters that match the expected voice interaction position relationship; according to the target sound effect parameters in the target space sound effect template, the original voice is processed to obtain the target voice, including: using the position relationship sound effect parameters in the target space sound effect template to perform voice processing on the original voice to obtain the target voice, so that the target voice matches the expected voice interaction position relationship.
[0100] The position-related sound effect parameters refer to sound effect parameters related to position, including direction-related sound effect parameters and distance-related sound effect parameters. Matching the position-related sound effect parameters with the expected voice interaction position relationship means that the position indicated by the position-related sound effect parameters is consistent with the position indicated by the expected voice interaction position relationship.
[0101] In this embodiment, when the playback terminal processes the original speech according to the target sound effect parameters in the target space sound effect template to obtain the target speech, it can specifically use the position relationship sound effect parameters in the target space sound effect template to perform speech processing on the original speech to obtain the target speech. Since the position relationship sound effect parameters in the target space sound effect template match the expected voice interaction position relationship, the obtained target speech matches the expected voice interaction position relationship.
[0102] In the above embodiment, since a target speech that matches the expected speech interaction position relationship can be obtained, the obtained target speech is adapted to the current acoustic scene and the speech listening quality is high.
[0103] In some embodiments, the expected voice interaction position relationship includes the expected voice interaction distance and the expected interaction direction; the position relationship sound effect parameters include direction-related sound effect parameters and distance-related sound effect parameters; using the position relationship sound effect parameters in the target space sound effect template to perform voice processing on the original voice to obtain the target voice includes: using the direction-related sound effect parameters to process the direction of the original voice, and using the distance-related sound effect parameters to process the sound pressure of the original voice to obtain the target voice; so that the direction of the target voice matches the expected interaction direction, and the sound pressure of the target voice matches the expected voice interaction distance.
[0104] Among them, the expected voice interaction distance refers to the voice interaction position expected by the sound receiving object in the current acoustic scene, and the expected interaction direction refers to the voice interaction direction expected by the sound receiving object in the current acoustic scene. Direction-related sound effect parameters refer to sound effect parameters related to direction, and the direction-related sound effect parameters can specifically be direction data. Distance-related sound effect parameters refer to sound effect parameters related to distance, and the direction-related sound effect parameters can specifically be one of distance data or sound pressure data. Sound pressure is used to characterize the size of a sound. The greater the sound pressure, the louder the sound. Conversely, the smaller the sound pressure, the quieter the sound.
[0105] In this embodiment, the target sound effect parameters include positional relationship sound effect parameters that match the expected voice interaction positional relationship, which may specifically include: distance-related sound effect parameters that match the expected voice interaction distance and orientation-related sound effect parameters that match the expected interaction orientation.
[0106] Specifically, the playing terminal processes the direction of the original speech by using the direction-related sound effect parameter to generate a stereo speech, and processes the sound pressure of the stereo speech by using the distance-related sound effect parameter to obtain the target speech. Since the direction-related sound effect parameter is matched with the expected interaction direction, and the distance-related sound effect parameter is matched with the expected speech interaction distance, the direction of the obtained target speech is matched with the expected interaction direction, and the sound pressure of the target speech is matched with the expected speech interaction distance.
[0107] In some embodiments, the playing terminal processes the direction of the original speech by using the direction-related sound effect parameter to generate a stereo speech can be specifically: matching the direction indicated by the direction-related sound effect parameter with the direction in the HRIR database, and convolving the matched HRIR data with the direction-related sound effect parameter to generate a stereo speech.
[0108] In some embodiments, since the closer the sound source is to the listener, the higher the sound pressure of the heard sound is, and conversely, the farther the sound source is from the listener, the lower the sound pressure of the heard sound is. The sound pressure level difference value formula of the distance r1 and the distance r2 is referred to the following formula (2), where lp2 is the sound pressure corresponding to the distance r2, and lp1 is the sound pressure corresponding to the distance r1:
[0109] lp2 = lp1 - 20lg(r2 / r1) (2)
[0110] In actual application, the sound pressure value lp1 at a specific distance r1 can be measured, and the sound pressure lp2 of the target distance r2 is generated by the above formula relationship. And it is mapped to the corresponding sound signal, thereby realizing the auditory perception effect of different distances. For example, r2 / r1 is equal to 2, and the sound pressure attenuation is 6db.
[0111] In other embodiments, the relationship between the sound pressure lp and the sound signal x(n) is as follows formula (3), where LP0 is a bias value. According to the corresponding relationship, the amplitude of the input signal can be adjusted to adjust the sound pressure, and then the effect of the change of the distance between the listener and the sound source is realized:
[0112]
[0113] In the above embodiments, the direction of the original speech is processed by using the direction-related sound effect parameter, and the sound pressure of the original speech is processed by using the distance-related sound effect parameter to obtain the target speech. The direction of the obtained target speech is matched with the expected interaction direction, and the sound pressure of the target speech is matched with the expected speech interaction distance, thereby obtaining the target speech with high listening quality.
[0114] In some embodiments, the target spatial sound effect template also matches the target voice interaction effect, which is the expected voice interaction effect of the sound receiving object in the current acoustic scene, and the target sound effect parameters include voice effect adjustment parameters that match the target voice interaction effect; using the positional relationship sound effect parameters in the target spatial sound effect template to perform voice processing on the original voice to obtain the target voice includes: using the positional relationship sound effect parameters and voice effect adjustment parameters in the target spatial sound effect template to process the original voice to obtain the target voice.
[0115] The expected voice interaction effect refers to the voice interaction effect expected by the sound receiver in the current acoustic scenario. The voice interaction effect may specifically be a reverberation effect. The voice effect adjustment parameter refers to a parameter used to adjust the sound effect. The voice effect adjustment parameter may specifically be an attenuation factor or a filter parameter.
[0116] Specifically, when the playback terminal uses the positional relationship sound effect parameters in the target spatial sound effect template to perform voice processing on the original voice to obtain the target voice, it can specifically use the positional relationship sound effect parameters in the target spatial sound effect template to process the original voice to obtain stereo voice, and use the voice effect adjustment parameters to process the voice effect of the original voice to obtain the target voice. The voice interaction effect of the obtained target voice matches the expected voice interaction effect.
[0117] like Figure 5 FIG. 1 is a schematic diagram of using voice effect adjustment parameters to adjust voice effects in some embodiments. Figure 5 The voice effect adjustment parameters are the parameters in the reverberator. The playback terminal performs reverberation processing on the voice through the reverberator to obtain the target voice with reverberation effect. Specifically, there are three branches inside the reverberator. These three branches can respectively obtain the direct voice signal, early reflected voice signal, and late reflected voice signal. The voice signals output by the three branches are superimposed to obtain the final target signal with reverberation signal. Among them:
[0118] Branch 1: The original speech x(n) is multiplied by the attenuation factor to obtain the direct speech signal;
[0119] Branch 2: The original speech x(n) passes through an 18-point filter, and the filtering result is multiplied by the early reflection attenuation factor to obtain the early reflection speech signal;
[0120] Branch 3: The original speech x(n) passes through an 18-point filter, then a weighted sum of six low-pass comb filters, and finally an all-pass filter. Finally, the filtered result is multiplied by the late reflection attenuation factor to obtain the late reflection speech signal.
[0121] In the above embodiment, the target sound effect parameters include voice effect adjustment parameters that match the target voice interaction effect. When the playback terminal processes the original voice, it can use the position relationship sound effect parameters and voice effect adjustment parameters in the target spatial sound effect template to process the original voice to obtain the target voice, so that the obtained target voice can match the expected voice interaction effect in the current acoustic scene, and the target voice has a higher listening quality, which can better meet the needs of the sound receiving object.
[0122] In some embodiments, scene recognition is performed on the ambient sound to obtain the current acoustic scene in which the playback terminal is located, including: obtaining multiple sound sub-segments corresponding to the ambient sound, performing feature extraction on the sound sub-segments, and obtaining sub-segment features; obtaining the segment acoustic scenes corresponding to the sound sub-segments based on the sub-segment feature recognition; counting the segment acoustic scenes corresponding to the sound sub-segments to obtain the number of scenes corresponding to each segment acoustic scene; and selecting the segment acoustic scene with the largest number of scenes as the current acoustic scene in which the playback terminal is located.
[0123] The sound sub-segments are obtained by segmenting the ambient sound. The ambient sound segmentation can be performed according to target time intervals. The sub-segment features can be power spectra or Mel-frequency cepstral coefficients.
[0124] Specifically, the playback terminal can divide the ambient sound into multiple sound sub-segments. For each sound sub-segment, it can be input into the trained acoustic scene recognition model to identify the acoustic scene corresponding to the sound sub-segment, that is, the segment acoustic scene. Each sound sub-segment corresponds to a segment acoustic scene. Therefore, the ambient sound corresponds to multiple segment acoustic scenes. The number of scenes of these segment acoustic scenes is counted to obtain the number of scenes corresponding to each segment acoustic scene. Then, the segment acoustic scene with the largest number of scenes can be selected as the current acoustic scene of the playback terminal.
[0125] For example, assuming that the ambient sound is divided into 6 sound sub-segments, namely sound sub-segment 1, sound sub-segment 2, sound sub-segment 3, sound sub-segment 4, sound sub-segment 5 and sound sub-segment 6, the segment acoustic scene corresponding to sound sub-segment 1 is scene A, the segment acoustic scene corresponding to sound sub-segment 2 is scene A, the segment acoustic scene corresponding to sound sub-segment 3 is scene B, the segment acoustic scene corresponding to sound sub-segment 4 is scene A, and the segment acoustic scene corresponding to sound sub-segment 3 is scene C. The final statistics show that the number of scenes corresponding to scene A is 4, the number of scenes corresponding to scene B is 1, and the number of scenes corresponding to scene C is 1. Scene A is then used as the current acoustic scene corresponding to the ambient sound, that is, the current acoustic scene where the playback terminal is located.
[0126] In some embodiments, the scene recognition model can be trained in the following manner: collecting environmental sounds in different acoustic scenes, dividing the environmental sounds into sound sub-segments and determining corresponding training labels, inputting each sound sub-segment into the scene recognition model, and training the acoustic scene recognition model with the training labels corresponding to each sound sub-segment as the expected output until the training stop condition is met, thereby obtaining a trained acoustic scene recognition model. Among them, the model parameters of the acoustic scene recognition model can be adjusted using the stochastic gradient descent algorithm, Adagrad (Adaptive Gradient) algorithm, Adadelta (an improvement of the AdaGrad algorithm), RMSprop (an improvement of the AdaGrad algorithm), Adam (Adaptive Moment Estimation) algorithm, etc. When the training stop condition is met, the training is completed and a trained acoustic scene recognition model is obtained. The training stop condition can be that the model parameters no longer change, the loss reaches the minimum value, the number of training times reaches the maximum number of iterations, and so on.
[0127] In some specific embodiments, such as Figure 6 The following is a schematic diagram of the model structure of the acoustic scene recognition model. Figure 6 , where the acoustic scene recognition model includes a five-layer convolutional network. The first layer is a dense convolutional network (DenseConvolutional Network, referred to as DenseNet), and the second to fourth layers are all GRU networks (Gate Recurrent Unit). The network parameters of the second to fourth layers are different. The fifth layer is a softmax layer, and the softmax layer can also adopt a DenseNet network structure. The final output of the acoustic scene recognition model can be a preset scene identifier. For example, if there are a total of 5 categories of scenes, the scene output identifier of the training sample is 5 binary numbers, such as 00100, which means that the sample corresponds to the third category of scenes. The final output of the acoustic scene recognition model can also be the probability of identifying each category of scenes. The final result is the final recognition scene result with the highest output probability of all categories.
[0128] In the above embodiment, by obtaining multiple sound sub-segments corresponding to the ambient sound, scene recognition is performed on each sound sub-segment to obtain multiple segment acoustic scenes corresponding to the ambient sound, the number of scenes of each segment acoustic scene is counted, and finally the segment acoustic scene with the largest number of scenes is selected as the current acoustic scene of the playback terminal. By dividing the segments, the sounds of each time period in the ambient sound can be fully considered, and the scene recognition results obtained are more accurate. The ambient sound is sliced and input into the acoustic scene recognition model. Since the length of the input signal is reduced, the model recognition results are also more accurate.
[0129] In some specific embodiments, a voice playback method is provided. Specifically, the playback terminal collects environmental sounds, performs acoustic scene recognition based on the collected environmental sounds, selects a spatial sound effect template that matches the scene based on the current acoustic scene obtained by recognition, generates virtual stereo based on the sound effect parameters in the spatial sound effect template and the original voice, and finally plays the generated stereo. Among them, the original voice signal can be a mono voice signal of the collected sound emitting object. In some specific embodiments, the virtual stereo generation process includes: for the collected mono signal, based on the sound effect parameter target orientation, reverberation parameters and distance parameters in the matching spatial sound effect template, first generate stereo based on HRTF technology, then perform distance volume adjustment, and finally perform reverberation processing to obtain a two-channel stereo signal.
[0130] This application also provides an application scenario, which applies the above-mentioned voice playback method. In this application scenario, the sound emitting object and the sound receiving object conduct a real-time voice call through the network. Specifically, refer to Figure 7 The application of the voice playback method in this application scenario is as follows:
[0131] Step 702: The playback terminal obtains the real-time voice of the sound emitting object through the network as the original voice, and collects the ambient sound of the current playback environment.
[0132] Step 704: The playback terminal divides the ambient sound into sound sub-segments.
[0133] Step 706: extract the power spectrum or Mel-frequency cepstral coefficients of the sound sub-segment to obtain the sub-segment features corresponding to the sound sub-segment.
[0134] Step 708: Input the sub-segment features into the acoustic scene recognition model to obtain the segment acoustic scene corresponding to the sound sub-segment.
[0135] Step 710: Count the segment acoustic scenes corresponding to the sound sub-segments to obtain the number of scenes corresponding to each segment acoustic scene, and select the segment acoustic scene with the largest number of scenes as the current acoustic scene of the playback terminal.
[0136] Among them, acoustic scenes include three categories: noisy scenes, quiet scenes and ordinary scenes. Ordinary scenes are scenes between noisy scenes and quiet scenes.
[0137] Step 712: Obtain a target spatial sound effect template that matches the current acoustic scene. The target spatial sound effect template matches the target voice interaction mode. The target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene.
[0138] Six candidate spatial sound effect templates can be pre-configured, including ear-to-ear communication template, strolling mobile communication template, lecture template, surprise communication template, surround template, and fly-in and fly-out template. Among them, the ear-to-ear communication template refers to the sound emitting object being close to the sound receiving object's ear for communication; the strolling mobile communication template refers to the sound emitting object being within a certain distance range, communicating with the sound receiving object in slow motion according to a random or set motion trajectory; the lecture template refers to the sound emitting object being at a medium or long distance, with a loud sound and a certain reverberation effect; the surprise communication template refers to the sound emitting object being at a medium or long distance, with a loud sound and a certain reverberation effect; the surprise communication template refers to the sound emitting object being at a medium or long distance, with a loud sound and a certain reverberation effect. The position of the sound emitting object is not fixed and its movement trajectory is random. For example, the sound appears in front of the left of the sound receiving object in the first sentence, behind the sound receiving object in the second sentence, and close to the ear of the sound receiving object in the third sentence, giving people a surprising auditory experience; the surround template refers to the sound emitting object maintaining a certain distance from the sound receiving object, and rotating 360 degrees around the sound receiving object in the horizontal direction to communicate; the fly-in and fly-out template refers to the sound emitting object approaching the position of the sound receiving object from a distance at a high speed, or moving away from a position close to the sound receiving object to a distance at a high speed.
[0139] Each spatial sound template contains a series of sound and image orientation, distance, reverberation parameters, etc. Based on these parameters, virtual stereo generation is achieved through relevant technologies. The generated stereo is played through headphones or multiple speakers. The multi-speaker playback here involves upmix conversion from dual channels to multi-channels and crosstalk elimination technology.
[0140] When the current acoustic scene is a noisy scene, the playback terminal obtains the ear-to-ear communication template as the target spatial sound effect template; when the current acoustic scene is a quiet scene, the playback terminal can select three spatial sound effect templates, namely, the surprise communication template, the surround template, and the fly-in-fly-out template, and provide a selection interface, and the user can select one from these three spatial sound effect templates as the target spatial sound effect template; when the current acoustic scene is an ordinary scene, the playback terminal can select two spatial sound effect templates, namely, the lecture template and the surprise communication template, and provide a selection interface, and the user can select one from these two spatial sound effect templates as the target spatial sound effect template.
[0141] Alternatively, ordinary scenes can be subdivided into natural scenes, such as a grove. At this time, the playback terminal can select a walking mobile communication template as the target space sound effect template. Ordinary scenes can be subdivided into open scenes, such as a large stadium. At this time, the playback terminal can select a lecture-style template as the target space sound effect template.
[0142] Step 714: Based on the target orientation in the spatial sound effect template, the playback terminal may determine the orientation data that matches it from the HRIR database, and convolve the HRIR data corresponding to the orientation data with the original voice signal to generate virtual stereo.
[0143] Step 716: Determine the stereophonic adjustment of the original speech according to the distance parameter in the spatial sound effect template, and perform reverberation processing according to the reverberation parameter in the spatial sound effect template to obtain a stereo sound signal with a reverberation effect.
[0144] In this application scenario, the technology that combines the real acoustic environment with the acoustics of a virtual space can address the different hearing needs of users in different acoustic scenarios and provide a brand new listening experience. By presenting an actual live listening experience through spatial sound effects, the audio quality of speech is significantly improved.
[0145] This application also provides another application scenario, in which the above-mentioned voice playing method is applied. In this application scenario, a sound emitting object and a sound receiving object conduct an instant voice call via a network.
[0146] In this application scenario, the playback terminal determines the voice message sent by the sound-emitting object as the original voice. When the playback terminal receives the sound-receiving object to trigger the playback operation of the original voice, it collects the ambient sound in the current playback environment, performs scene recognition on the ambient sound, obtains the current acoustic scene in which the playback terminal is located, obtains the target spatial sound effect template that matches the current acoustic scene, processes the voice message according to the target sound effect parameters in the target spatial sound effect template to obtain the target voice, and plays the target voice. In this application scenario, when the playback terminal obtains multiple spatial sound effect templates that match the current acoustic scene, the playback terminal may pop up a prompt box for selecting a spatial sound effect template, prompting the user to select a spatial sound effect template as the target spatial sound effect template.
[0147] refer to Figure 8, is a schematic diagram of a prompt box for selecting a spatial sound effect template displayed by the playback terminal on the conversation interface in some embodiments. In this embodiment, when the user clicks on the voice message 802, the playback terminal receives the user's playback operation on the voice message. After the current acoustic scene is determined, the spatial sound effect templates matching the current acoustic scene are obtained, including a surprising communication template, a surround template, and a fly-in and fly-out template. A prompt box 804 for selecting a spatial sound effect template is displayed on the conversation interface 800. The terminal can click on the text information corresponding to any spatial sound effect model to select the spatial sound effect template. The playback terminal determines the spatial sound effect template selected by the user as the target spatial sound effect template. The prompt box 804 can also display an option to not use the spatial sound effect template. When the user clicks not to use the spatial sound effect template, the playback terminal plays the voice message without any spatial sound effect processing.
[0148] In some embodiments, when the conversation interface diagram has multiple voice messages to be played, upon receiving a play operation for the current voice message, acoustic scene recognition is performed. When the current acoustic scene is recognized, a prompt box for selecting a spatial sound effect template pops up. After the user selects the target spatial sound effect template, the playback terminal continues acoustic scene recognition for the next voice message. If the recognized acoustic scene remains unchanged (i.e., the same as the acoustic scene of the previous voice message), the spatial sound effect template selected by the user in the previous voice message is automatically adopted as the target spatial sound effect template corresponding to the voice message. If the recognized current acoustic scene changes, a prompt box for selecting the spatial sound effect template corresponding to the changed acoustic scene pops up.
[0149] It should be understood that although Figure 2-8 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 2-8 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0150] In some embodiments, as Figure 9 As shown, a voice playback device 900 is provided. The device can be a software module or a hardware module, or a combination of the two to form a part of a computer device. The device specifically includes:
[0151] The voice acquisition module 902 is used to acquire the original voice to be played and the ambient sound in the current playback environment of the playback terminal of the original voice;
[0152] The scene recognition module 904 is used to perform scene recognition on the ambient sound to obtain the current acoustic scene of the playback terminal;
[0153] The template acquisition module 906 is used to obtain a target spatial sound effect template that matches the current acoustic scene. The target spatial sound effect template matches the target voice interaction mode. The target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene.
[0154] The speech processing module 908 is used to process the original speech according to the target sound effect parameters in the target space sound effect template to obtain the target speech, so as to play the target speech in the playback terminal.
[0155] The above-mentioned voice playback device obtains the original voice to be played, obtains the ambient sound in the playback environment where the playback terminal of the original voice is currently located, performs scene recognition on the ambient sound, identifies the current acoustic scene in which the playback terminal is located, obtains the target space sound effect template that matches the current acoustic scene, and processes the original voice according to the target sound effect parameters in the target space sound effect template to obtain the target voice, so as to play the target voice in the playback terminal. Since the target voice is obtained by processing the original voice according to the target sound effect parameters in the target space sound effect template, the target space sound effect template matches the target voice interaction mode, and the target voice interaction mode is the expected voice interaction mode of the sound receiving object in the current acoustic scene, the obtained target voice can meet the auditory needs of the sound receiving object and improve the quality of listening.
[0156] In some embodiments, the expected voice interaction mode includes an expected voice interaction position relationship between a sound receiving object and a sound emitting object; the template acquisition module is also used to obtain a set of candidate spatial sound effect templates; the candidate spatial sound effect template set includes multiple candidate spatial sound effect templates corresponding to different voice interaction position relationships; from the candidate spatial sound effect template set, a target spatial sound effect template that matches the current acoustic scene is selected, wherein the voice interaction position relationship corresponding to the target spatial sound effect template matches the expected voice interaction position relationship corresponding to the current acoustic scene.
[0157] In some embodiments, the candidate spatial sound effect template set includes candidate spatial sound effect templates with changes in voice interaction position relationships; the template acquisition module is also used to select candidate spatial sound effect templates with changes in voice interaction position relationships from the candidate spatial sound effect template set when the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, as the target spatial sound effect template that matches the current acoustic scene.
[0158] In some embodiments, the target spatial sound effect template includes a target sound orientation parameter sequence corresponding to the dynamic interaction position relationship; the speech processing module is also used to divide the original speech into speech segments with the number of parameters in the target sound orientation parameter sequence; according to the order of the speech segments in the original speech, the target sound orientation parameters corresponding to the speech segments in the target sound orientation parameter sequence are determined; the speech segments are processed according to the target sound orientation parameters corresponding to the speech segments to obtain processed speech segments, and each processed speech segment forms the target speech in the speech order.
[0159] In some embodiments, the candidate spatial sound effect template set includes candidate spatial sound effect templates with fixed voice interaction position relationships; the template acquisition module is also used to select candidate spatial sound effect templates with fixed voice interaction position relationships from the candidate spatial sound effect template set as the target spatial sound effect template matching the current acoustic scene when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship.
[0160] In some embodiments, the expected voice interaction position relationship includes an expected voice interaction distance; the template acquisition module is also used to select, from the candidate spatial sound effect template set, a candidate spatial sound effect template whose sound distance corresponding to the sound effect parameters is fixed and whose sound distance matches the expected voice interaction distance when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, as a target spatial sound effect template that matches the current acoustic scene.
[0161] In some embodiments, the expected voice interaction mode includes the expected voice interaction position relationship between the sound receiving object and the sound emitting object; the target sound effect parameters include position relationship sound effect parameters that match the expected voice interaction position relationship; the voice processing module is also used to use the position relationship sound effect parameters in the target space sound effect template to perform voice processing on the original voice to obtain the target voice, so that the target voice matches the expected voice interaction position relationship.
[0162] In some embodiments, the expected voice interaction position relationship includes the expected voice interaction distance and the expected interaction direction; the position relationship sound effect parameters include direction-related sound effect parameters and distance-related sound effect parameters; the voice processing module is also used to use the direction-related sound effect parameters to process the direction of the original voice, and use the distance-related sound effect parameters to process the sound pressure of the original voice to obtain the target voice; so that the direction of the target voice matches the expected interaction direction, and the sound pressure of the target voice matches the expected voice interaction distance.
[0163] In some embodiments, the target spatial sound effect template also matches the target voice interaction effect, which is the expected voice interaction effect of the sound receiving object in the current acoustic scene, and the target sound effect parameters include voice effect adjustment parameters that match the target voice interaction effect; the voice processing module is also used to use the positional relationship sound effect parameters and voice effect adjustment parameters in the target spatial sound effect template to process the original voice to obtain the target voice.
[0164] In some embodiments, the scene recognition module is also used to obtain multiple sound sub-segments corresponding to the ambient sound, perform feature extraction on the sound sub-segments, and obtain sub-segment features; obtain the segment acoustic scenes corresponding to the sound sub-segments based on the sub-segment feature recognition; count the segment acoustic scenes corresponding to the sound sub-segments to obtain the number of scenes corresponding to each segment acoustic scene; and select the segment acoustic scene with the largest number of scenes as the current acoustic scene of the playback terminal.
[0165] For the specific definition of the voice playback device, please refer to the definition of the voice playback method above, which will not be repeated here. The various modules in the above-mentioned voice playback device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0166] In some embodiments, a computer device is provided. The computer device may be a playback terminal, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a voice playback method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0167] Those skilled in the art will understand that Figure 10The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0168] In some embodiments, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0169] In some embodiments, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0170] In some embodiments, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of each of the above method embodiments.
[0171] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0172] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0173] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A voice playing method, characterized in that: Executed by a playback terminal, the method includes: During a real-time voice call between the playback terminal and the first terminal, obtaining an original voice to be played obtained by the first terminal collecting the voice of the sound emitting object, and obtaining ambient sound in a playback environment currently located in the playback terminal of the original voice; Using the trained acoustic scene recognition model, identify the acoustic scenes corresponding to the multiple sound sub-segments of the ambient sound, and use the acoustic scene with the largest number of scenes as the current acoustic scene of the playback terminal; Obtaining a target spatial sound effect template that matches the current acoustic scene, the target spatial sound effect template matching a target voice interaction mode, the target voice interaction mode being an expected voice interaction mode for a sound receiving object in the current acoustic scene; the expected voice interaction mode including an expected voice interaction positional relationship between the sound receiving object and the sound emitting object; According to the target sound effect parameters in the target spatial sound effect template, virtual stereo generation is performed on the original voice to obtain the target voice, so that the target voice matches the expected voice interaction position relationship, so as to play the target voice in the playback terminal; the target sound effect parameters include positional relationship sound effect parameters that match the expected voice interaction position relationship.
2. The method according to claim 1, characterized in that The step of obtaining a target spatial sound effect template that matches the current acoustic scene includes: Obtaining a set of candidate spatial sound effect templates; the set of candidate spatial sound effect templates includes a plurality of candidate spatial sound effect templates corresponding to different voice interaction position relationships; From the candidate spatial sound effect template set, a target spatial sound effect template that matches the current acoustic scene is selected, wherein the voice interaction position relationship corresponding to the target spatial sound effect template matches the expected voice interaction position relationship corresponding to the current acoustic scene.
3. The method according to claim 2, characterized in that The candidate spatial sound effect template set includes candidate spatial sound effect templates with changes in voice interaction position relationships; The selecting a target spatial sound effect template that matches the current acoustic scene from the candidate spatial sound effect template set includes: When the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, a candidate spatial sound effect template with a changing voice interaction position relationship is selected from the candidate spatial sound effect template set as the target spatial sound effect template that matches the current acoustic scene.
4. The method according to claim 3, characterized in that The target spatial sound effect template includes a target sound orientation parameter sequence corresponding to the dynamic interactive position relationship; The generating of the target speech by performing virtual stereo on the original speech according to the target sound effect parameters in the target spatial sound effect template to obtain the target speech comprises: Cutting the original speech into speech segments of the number of parameters in the target sound orientation parameter sequence; determining, according to the order of the speech segments in the original speech, target sound orientation parameters corresponding to the speech segments in the target sound orientation parameter sequence; The speech segments are processed according to the target sound orientation parameters corresponding to the speech segments to obtain processed speech segments, and each of the processed speech segments forms the target speech in a speech order.
5. The method according to claim 2, characterized in that The candidate spatial sound effect template set includes candidate spatial sound effect templates with fixed voice interaction position relationships; The selecting a target spatial sound effect template that matches the current acoustic scene from the candidate spatial sound effect template set includes: When the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, a candidate spatial sound effect template with a fixed voice interaction position relationship is selected from the candidate spatial sound effect template set as the target spatial sound effect template matching the current acoustic scene.
6. The method according to claim 5, characterized in that The expected voice interaction position relationship includes an expected voice interaction distance; When the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, selecting a candidate spatial sound effect template with a fixed voice interaction position relationship from the candidate spatial sound effect template set as a target spatial sound effect template matching the current acoustic scene includes: When the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, a candidate spatial sound effect template whose sound distance corresponding to the sound effect parameters is fixed and whose sound distance matches the expected voice interaction distance is selected from the candidate spatial sound effect template set as the target spatial sound effect template matching the current acoustic scene.
7. The method according to claim 1, characterized in that The expected voice interaction position relationship includes an expected voice interaction distance and an expected interaction direction; the position relationship sound effect parameters include direction-related sound effect parameters and distance-related sound effect parameters; The step of generating a virtual stereo image of the original speech according to the target sound effect parameters in the target spatial sound effect template to obtain the target speech includes: The orientation of the original speech is processed using the orientation-related sound effect parameters, and the sound pressure of the original speech is processed using the distance-related sound effect parameters to obtain a target speech; so that the orientation of the target speech matches the expected interaction orientation, and the sound pressure of the target speech matches the expected speech interaction distance.
8. The method according to claim 7, characterized in that The target spatial sound effect template also matches a target voice interaction effect, where the target voice interaction effect is the expected voice interaction effect of the sound receiving object in the current acoustic scene, and the target sound effect parameters include voice effect adjustment parameters that match the target voice interaction effect; The step of generating a virtual stereo image of the original speech according to the target sound effect parameters in the target spatial sound effect template to obtain the target speech includes: The original speech is processed using the positional relationship sound effect parameters in the target spatial sound effect template and the speech effect adjustment parameters to obtain the target speech.
9. The method according to claim 1, characterized in that The step of identifying the acoustic scenes corresponding to the plurality of sound sub-segments of the ambient sound by using the trained acoustic scene recognition model, and using the acoustic scene with the largest number of scenes as the current acoustic scene of the playback terminal includes: Acquire multiple sound sub-segments corresponding to the environmental sound, perform feature extraction on the sound sub-segments, and obtain sub-segment features; Obtaining, by means of a trained acoustic scene recognition model, a segment acoustic scene corresponding to the sound sub-segment based on the sub-segment feature recognition; Counting the segment acoustic scenes corresponding to the sound sub-segments to obtain the number of scenes corresponding to each segment acoustic scene; The segment acoustic scene with the largest number of scenes is selected as the current acoustic scene of the playback terminal.
10. A voice playback device, characterized in that: The device comprises: A voice acquisition module is used to acquire, during a real-time voice call between a playback terminal and a first terminal, an original voice to be played obtained by the first terminal collecting the voice of a sound-emitting object, and to acquire the ambient sound of the playback environment in which the playback terminal of the original voice is currently located; A scene recognition module is used to identify the acoustic scenes corresponding to the multiple sound sub-segments of the ambient sound using a trained acoustic scene recognition model, and use the acoustic scene with the largest number of scenes as the current acoustic scene of the playback terminal; A template acquisition module is configured to acquire a target spatial sound effect template that matches the current acoustic scene, wherein the target spatial sound effect template matches a target voice interaction mode, wherein the target voice interaction mode is an expected voice interaction mode for a sound receiving object in the current acoustic scene; the expected voice interaction mode includes an expected voice interaction positional relationship between the sound receiving object and the sound emitting object; A speech processing module is used to generate virtual stereo for the original speech according to the target sound effect parameters in the target spatial sound effect template to obtain the target speech, so that the target speech matches the expected speech interaction position relationship, so as to play the target speech in the playback terminal; the target sound effect parameters include positional relationship sound effect parameters that match the expected speech interaction position relationship.
11. The device according to claim 10, characterized in that The template acquisition module is also used to obtain a set of candidate spatial sound effect templates; the candidate spatial sound effect template set includes multiple candidate spatial sound effect templates corresponding to different voice interaction position relationships; from the candidate spatial sound effect template set, a target spatial sound effect template that matches the current acoustic scene is selected, wherein the voice interaction position relationship corresponding to the target spatial sound effect template matches the expected voice interaction position relationship corresponding to the current acoustic scene.
12. The device according to claim 11, characterized in that The candidate spatial sound effect template set includes candidate spatial sound effect templates with changes in voice interaction position relationships; the template acquisition module is also used to select, from the candidate spatial sound effect template set, candidate spatial sound effect templates with changes in voice interaction position relationships when the current voice interaction position relationship corresponding to the current acoustic scene is a dynamic interaction position relationship, as the target spatial sound effect template matching the current acoustic scene.
13. The device according to claim 12, characterized in that The target spatial sound effect template includes a target sound orientation parameter sequence corresponding to the dynamic interactive position relationship; the speech processing module is also used to divide the original speech into speech segments with the number of parameters in the target sound orientation parameter sequence; determine the target sound orientation parameters corresponding to the speech segments in the target sound orientation parameter sequence according to the order of the speech segments in the original speech; process the speech segments according to the target sound orientation parameters corresponding to the speech segments to obtain processed speech segments, and each of the processed speech segments forms the target speech in the speech order.
14. The device according to claim 11, characterized in that The candidate spatial sound effect template set includes candidate spatial sound effect templates with fixed voice interaction position relationships; the template acquisition module is also used to select, from the candidate spatial sound effect template set, a candidate spatial sound effect template with fixed voice interaction position relationship as a target spatial sound effect template matching the current acoustic scene when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship.
15. The device according to claim 14, characterized in that The expected voice interaction position relationship includes an expected voice interaction distance; the template acquisition module is also used to, when the current voice interaction position relationship corresponding to the current acoustic scene is a fixed interaction position relationship, select from the candidate spatial sound effect template set, a candidate spatial sound effect template whose sound distance corresponding to the sound effect parameters is fixed and whose sound distance matches the expected voice interaction distance, as a target spatial sound effect template matching the current acoustic scene.
16. The device according to claim 10, characterized in that The expected voice interaction position relationship includes the expected voice interaction distance and the expected interaction direction; the position relationship sound effect parameters include direction-related sound effect parameters and distance-related sound effect parameters; the voice processing module is also used to use the direction-related sound effect parameters to process the direction of the original voice, and use the distance-related sound effect parameters to process the sound pressure of the original voice to obtain the target voice; so that the direction of the target voice matches the expected interaction direction, and the sound pressure of the target voice matches the expected voice interaction distance.
17. The device according to claim 16, characterized in that The target spatial sound effect template also matches the target voice interaction effect, and the target voice interaction effect is the expected voice interaction effect of the sound receiving object in the current acoustic scene. The target sound effect parameters include voice effect adjustment parameters that match the target voice interaction effect; the voice processing module is also used to use the position relationship sound effect parameters in the target spatial sound effect template and the voice effect adjustment parameters to process the original voice to obtain the target voice.
18. The device according to claim 10, characterized in that The scene recognition module is further configured to obtain a plurality of sound sub-segments corresponding to the ambient sound, perform feature extraction on the sound sub-segments, and obtain sub-segment features; Through the trained acoustic scene recognition model, the segment acoustic scene corresponding to the sound sub-segment is obtained based on the sub-segment feature recognition; the segment acoustic scenes corresponding to the sound sub-segment are counted to obtain the number of scenes corresponding to each segment acoustic scene; the segment acoustic scene with the largest number of scenes is selected as the current acoustic scene of the playback terminal.
19. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.
20. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
21. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method for adding background sound effect according to position of mobile terminal, and mobile terminal
CN104410748A
Audio playing method and device, computer equipment and storage medium
CN111142838A