Accompaniment generation method, device and storage medium

By merging and processing dry sound signals and background music in a virtual three-dimensional space, a full-dimensional surround accompaniment is generated, which solves the problem of non-stereoscopic sound effects in multi-person chorus scenes and improves the user experience.

CN114242025BActive Publication Date: 2025-09-09TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111527995.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-09-09
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The existing virtual 3D audio technology does not provide three-dimensional sound effects in multi-person chorus scenarios, resulting in a poor user experience.

Method used

By obtaining a set of dry sound signals, a virtual sound signal is generated based on the sound and image position in the virtual three-dimensional space, and after merging and processing, it is synthesized with the background music to generate a full-dimensional surround accompaniment.

Benefits of technology

It achieves a stereo audio surround effect and enhances the user's immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242025B_ABST
    Figure CN114242025B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a method, device and storage medium for generating an accompaniment, and the method for generating the accompaniment includes: obtaining a dry sound signal set, wherein the dry sound signal set includes x dry sound signals corresponding to a target song; generating a virtual sound signal based on the dry sound signal corresponding to each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions, wherein the x dry sound signals correspond to N virtual three-dimensional sound image positions, the N virtual three-dimensional sound image positions are different, and each virtual three-dimensional sound image position is allowed to correspond to one or more dry sound signals in the x dry sound signals; merging and processing each virtual sound signal in the virtual sound signal set to obtain a chorus dry sound; performing sound effect synthesis processing on the chorus dry sound and the background music of the target song according to the sound effect optimization rules to obtain the accompaniment of the target song. By adopting the present application, the effect of stereo surround of the accompaniment can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer application technology, and in particular to a method, device and storage medium for generating an accompaniment. Background Art

[0002] With the development of virtual reality (VR) technology, virtual three-dimensional (3D) audio technology is also gradually being optimized. Virtual 3D audio technology can create a three-dimensional dynamic effect. When applied to singing software, it can provide users with an immersive experience. Currently, when applying virtual 3D audio technology to multi-person chorus scenarios, the existing technical solution is to directly weight and superimpose multiple voices. However, this processing method makes the sound effect less three-dimensional, resulting in a poor user experience. Summary of the Invention

[0003] The embodiments of the present application provide a method, device, and storage medium for generating an accompaniment, which can achieve an all-round audio stereo surround effect and improve the user experience.

[0004] On the one hand, an embodiment of the present application provides a method for generating an accompaniment, comprising:

[0005] Obtain a dry sound signal set, where the dry sound signal set includes x dry sound signals corresponding to the target song, where x is an integer greater than 1;

[0006] generating a virtual sound signal based on a dry sound signal corresponding to each virtual three-dimensional spatial sound image position in N virtual three-dimensional spatial sound image positions, wherein the x dry sound signals correspond to the N virtual three-dimensional spatial sound image positions, N is an integer greater than 1, the N virtual three-dimensional spatial sound image positions are different, and each virtual three-dimensional spatial sound image position is allowed to correspond to one or more dry sound signals in the x dry sound signals;

[0007] Merging each virtual sound signal in a virtual sound signal set to obtain a dry chorus sound, the virtual sound signal set including: a virtual sound signal at each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions;

[0008] According to the sound effect optimization rules, the dry chorus voice and the background music of the target song are synthesized to obtain the accompaniment of the target song.

[0009] On the one hand, an embodiment of the present application provides a method for playing an accompaniment, comprising:

[0010] displaying a user interface, the user interface being used to receive a selection instruction for a target song;

[0011] If the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode, then obtaining the accompaniment corresponding to the target song;

[0012] Play the accompaniment corresponding to the target song;

[0013] The accompaniment is generated based on the chorus dry voice and background music. The chorus dry voice is generated based on multiple dry voice signals in a dry voice signal set. The multiple dry voice signals in the dry voice signal set correspond to multiple different virtual three-dimensional sound and image positions. The dry voice signal set is obtained based on the dry voice signals recorded by multiple users for the target song.

[0014] On the other hand, an embodiment of the present application provides an accompaniment generation device, comprising:

[0015] An acquisition unit is used to acquire a dry sound signal set, where the dry sound signal set includes x dry sound signals corresponding to a target song, where x is an integer greater than 1; and a virtual sound signal is generated based on the dry sound signal corresponding to each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions, where the x dry sound signals correspond to N virtual three-dimensional sound image positions, where N is an integer greater than 1, the N virtual three-dimensional sound image positions are different, and each virtual three-dimensional sound image position is allowed to correspond to one or more dry sound signals in the x dry sound signals.

[0016] The processing unit is used to merge the various virtual sound signals in the virtual sound signal set to obtain a dry chorus sound, where the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions; and perform sound effect synthesis processing on the dry chorus sound and the background music of the target song according to the sound effect optimization rules to obtain the accompaniment of the target song.

[0017] On the other hand, an embodiment of the present application provides an accompaniment playback processing device, comprising:

[0018] The acquisition unit is used to display a user interface, which is used to receive a selection instruction for a target song; if the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode, the accompaniment corresponding to the target song is acquired.

[0019] A processing unit is used to play the accompaniment corresponding to the target song; the accompaniment is generated based on the chorus dry sound and background music, the chorus dry sound is generated based on multiple dry sound signals in a dry sound signal set, the multiple dry sound signals in the dry sound signal set correspond to multiple different virtual three-dimensional space sound and image positions, and the dry sound signal set is obtained based on the dry sound signals recorded by multiple users for the target song.

[0020] Accordingly, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a network interface, wherein the processor is connected to the memory and the network interface, wherein the network interface is used to provide network communication functions, the memory is used to store program code, and the processor is used to call the program code and execute the method in the embodiment of the present application.

[0021] Accordingly, an embodiment of the present application provides a computer-readable storage medium, including: a computer program stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method in the embodiment of the present application is implemented.

[0022] Accordingly, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, which are stored in a computer-readable storage medium. The processor of a computer device reads and executes the computer instructions from the computer-readable storage medium, so that the computer device performs the method in the embodiment of the present application.

[0023] By implementing the embodiments of the present application, on the one hand, the virtual sound signals of each dry sound signal corresponding to the target song in the dry sound signal set in different virtual three-dimensional sound and image positions can be obtained, and then the virtual sound signals corresponding to each dry sound signal can be merged and processed to obtain the chorus dry sound, and finally the chorus dry sound and the background music of the target song can be synthesized according to the sound effect optimization rules to obtain the accompaniment of the target song; on the other hand, the user's selection instruction for the target song can be received, and when the received selection instruction indicates that the accompaniment mode of the target song is the chorus accompaniment mode, the accompaniment corresponding to the target song can be obtained and played. In this way, the sound and image positions of the dry sound signals in the virtual three-dimensional space can be simulated in all directions, and the effect of audio stereo surround can be achieved, so that the user can have an immersive feeling in listening when obtaining the corresponding accompaniment, and obtain an immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 This is a schematic diagram of an application scenario of a method for generating an accompaniment provided in an embodiment of the present application;

[0026] Figure 2 This is a flow chart of a method for generating an accompaniment provided in an embodiment of the present application;

[0027] Figure 3a Schematic diagram of a horizontal plane, an upper plane, and a lower plane in a method for generating an accompaniment provided in an embodiment of the present application;

[0028] Figure 3b Schematic diagram of the virtual three-dimensional sound image position in a method for generating an accompaniment provided in an embodiment of the present application;

[0029] Figure 3c This is a schematic diagram of dividing each plane at preset angle intervals in a method for generating an accompaniment provided in an embodiment of the present application;

[0030] Figure 4 1 is a flow chart of another method for generating an accompaniment provided in an embodiment of the present application;

[0031] Figure 5 This is a flow chart of obtaining a dual-channel signal corresponding to a dry sound signal in a dry sound signal set in a method for generating an accompaniment provided in an embodiment of the present application;

[0032] Figure 6 1 is a flow chart of a method for playing an accompaniment provided in an embodiment of the present application;

[0033] Figure 7a This is a flow chart of obtaining the accompaniment corresponding to the target song in an accompaniment playback processing method provided in an embodiment of the present application;

[0034] Figure 7b This is a schematic diagram showing a first single sentence interface in an accompaniment playback processing method provided by an embodiment of the present application;

[0035] Figure 7c This is a schematic diagram showing a second single sentence interface in an accompaniment playback processing method provided by an embodiment of the present application;

[0036] Figure 8a 1 is a schematic structural diagram of an accompaniment generation device provided in an embodiment of the present application;

[0037] Figure 8b 1 is a schematic structural diagram of an accompaniment playback processing device provided in an embodiment of the present application;

[0038] Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0040] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0041] 1) Dry sound signal: The dry sound signal in the embodiment of the present application refers to a pure human voice signal without accompaniment music. The dry sound signal is a monophonic sound signal, that is, it does not include any directional information.

[0042] 2) Dual-channel signal: Dual-channel means having two sound channels. The principle is that when people hear a sound, they can determine the specific location of the sound source based on the phase difference between the left and right ears. In the embodiments of the present application, the dual-channel signal refers to the left channel sound signal and the right channel sound signal.

[0043] 3) Head-Related Transfer Functions (HRTFs): HRTFs, also known as binaural transfer functions, describe the transmission of sound waves from a sound source to both ears. HRTFs are a set of filters that employ the principle that time-domain convolution is equivalent to frequency-domain convolution. Based on the HRTF data corresponding to the sound source's location, they calculate the virtual sound signals transmitted to both ears.

[0044] The embodiment of the present application provides a method, device and storage medium for generating an accompaniment. By implementing the embodiment of the present application, a dry sound signal set consisting of multiple dry sound signals corresponding to the same target song can be obtained, and the virtual sound signals corresponding to the dry sound signals included in the dry sound signal set in different virtual three-dimensional sound and image positions can be obtained. Then, the virtual sound signals corresponding to the dry sound signals are merged to obtain the chorus dry sound, and finally, the chorus dry sound and the background music of the target song are synthesized according to the sound effect optimization rules to obtain the accompaniment of the target song. In this way, on the one hand, the sound and image positions of the dry sound signals included in the dry sound signal set in the virtual three-dimensional space can be simulated in all directions, and then the chorus dry sound obtained by merging the virtual sound signals corresponding to the dry sound signals in different virtual three-dimensional sound and image positions can be obtained, thereby achieving an audio stereo surround effect. On the other hand, the chorus dry sound and the background music can be synthesized according to the sound effect optimization rules to obtain the accompaniment, thereby enhancing the immersiveness of the audio effect. In general, compared with the processing method of directly superimposing the dry sound signals, the present application can obtain richer audio processing effects and improve the user experience.

[0045] See Figure 1 , Figure 1 This is a schematic diagram of an application scenario of a method for generating an accompaniment provided in an embodiment of the present application. Figure 1 As shown, the application scenario may include a smart device 100 , which communicates with a server 110 via a wired or wireless manner, and the server 110 is connected to a database 120 .

[0046] The method for generating accompaniment provided in the embodiment of the present application can be implemented by an electronic device such as the smart device 100. For example, when the smart device 100 receives a selection instruction indicating that the accompaniment mode for the target song is a chorus accompaniment mode, it can obtain a set of dry sound signals corresponding to the target song, and generate a virtual sound signal based on the dry sound signal corresponding to each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions. For example, the virtual sound signal can be a two-channel signal, and then the virtual sound signals corresponding to the various dry sound signals are merged to obtain a chorus dry sound, and then the chorus dry sound and the background music of the target song are synthesized according to the sound effect optimization rules to obtain the accompaniment. As an example, in Figure 1 The smart device 100 displays a "Chorus Accompaniment" option. The user can generate a chorus accompaniment mode selection instruction through voice control or by triggering a selection control displayed on the user interface. The dry sound signal set can be pre-stored locally by the smart device 100 or obtained by the smart device 100 from the server 110 or database 120.

[0047] The method for generating the accompaniment provided in the embodiment of the present application can also be implemented by an electronic device such as the server 110. For example, when the server 110 receives a selection instruction indicating that the accompaniment mode for the target song is a chorus accompaniment mode, it can obtain a set of dry sound signals corresponding to the target song, and generate a virtual sound signal based on the dry sound signal corresponding to each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions. The virtual sound signal can be, for example, a two-channel signal, and then the virtual sound signals corresponding to the various dry sound signals are merged to obtain a chorus dry sound, and then the chorus dry sound and the background music of the target song are synthesized according to the sound effect optimization rules to obtain an accompaniment. The dry sound signal set can be pre-stored locally by the server 110, or it can be obtained by the server 110 from the database 120. The final accompaniment can be stored locally or stored in the database 120 for calling when needed. Of course, the server 110 does not need to start generating the accompaniment when the received selection instruction indicates that the accompaniment mode for the target song is the chorus accompaniment mode. The server 110 can start executing the relevant steps of the accompaniment generation method of the present application to generate the accompaniment at an appropriate time, such as when the server 110 load is low, or when the server 110 receives a new dry sound signal of the target song, or when receiving a management operation for generating the accompaniment. Preferably, the chorus version of the accompaniment can be generated in advance and then stored in the server. After generating the accompaniment of a large number of songs, the user can use the smart device 100 to select "chorus accompaniment" on the user interface to issue a selection instruction for the target song. In this way, the server 110 can respond to the selection instruction, find the chorus accompaniment of the target song from the large number of generated accompaniments, and send the chorus accompaniment to the smart device 100.

[0048] The accompaniment generation method provided in the embodiment of the present application can also be implemented collaboratively by an electronic device such as the smart device 100 and an electronic device such as the server 110. For example, the server 110 can generate a virtual sound signal based on the dry sound signal corresponding to each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions. The virtual sound signal can be, for example, a two-channel signal. The virtual sound signals corresponding to the respective dry sound signals are then merged to obtain a chorus dry sound. The chorus dry sound is then synthesized with the background music of the target song according to the sound effect optimization rules to obtain an accompaniment, and the obtained accompaniment is sent to the smart device 100.

[0049] The accompaniment generation method provided in the embodiments of the present application can also be implemented by running a computer program on an electronic device such as the smart device 100 and an electronic device such as the server 110. For example, the computer program can be a native program or software module in an operating system, a local application (APP), or a mini-program. In short, the computer program can be any form of application, module, or plug-in, and the embodiments of the present application do not specifically limit this.

[0050] The smart devices involved in the embodiments of the present application may be personal computers, laptops, smart phones, tablet computers, smart watches, smart voice interaction devices, smart home appliances, vehicle-mounted terminals and smart wearable devices, etc., but are not limited thereto. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Smart devices and servers may be connected directly or indirectly via wired or wireless communications, which is not specifically limited in the embodiments of the present application.

[0051] It should be understood that Figure 1 The numbers of dry sound signals and virtual three-dimensional sound image positions shown in the figure are merely illustrative. Depending on implementation requirements, the dry sound signal set may include any number of dry sound signals, and the virtual three-dimensional space may include any number of virtual three-dimensional sound image positions.

[0052] For further information, see Figure 2 , Figure 2 This is a flow chart of a method for generating an accompaniment provided in an embodiment of the present application. The method in the embodiment of the present application can be applied to electronic devices, such as smart devices such as smartphones, tablet computers, smart wearable devices, personal computers, and servers. The method may include but is not limited to the following steps:

[0053] S201: Acquire a dry sound signal set.

[0054] In the implementation of the present application, the electronic device can obtain a dry sound signal set, which includes several dry sound signals corresponding to the target song.

[0055] In one embodiment, the dry sound signal set may be obtained from an audio database that includes initial dry sound signals recorded by multiple users when performing the same song. It should be noted that the initial dry sound signals in the audio database are recorded with the authorization of the users. The electronic device may filter out dry sound signals that meet the requirements based on the sound parameters of the initial dry sound signals to form the dry sound signal set.

[0056] In one embodiment, the electronic device can filter out dry sound signals that meet the conditions from the initial dry sound signal set based on the pitch characteristic parameters and the sound quality characteristic parameters. The pitch characteristic parameters may include any one or more of the pitch parameters, rhythm parameters and rhythmic parameters. The dry sound signals that meet the conditions filtered out based on the pitch characteristic parameters have the characteristics of high consistency between the song pitch, rhythm and accompaniment melody; the sound quality characteristic parameters may include any one or more of the noise parameters, energy parameters and speed parameters. The dry sound signals that meet the conditions filtered out based on the sound quality characteristic parameters have the characteristics of clear audio, appropriate audio energy and uniform audio speed. The embodiment of the present application does not limit the order of filtering dry sound signals that meet the conditions. For example, the electronic device can first filter out dry sound signals that meet the conditions based on the pitch characteristic parameters, and then filter out dry sound signals that meet the preset sound quality characteristic parameter conditions from the dry sound signals that meet the preset pitch characteristic parameter conditions. It can also first filter out dry sound signals that meet the conditions based on the sound quality characteristic parameters, and then filter out dry sound signals that meet the preset audio characteristic parameter conditions from the dry sound signals that meet the preset sound quality characteristic parameter conditions. The dry sound signal set formed by the dry sound signals obtained by screening the initial dry sound signal set in this way has excellent pitch and sound quality.

[0057] S202: Generate a virtual sound signal based on a dry sound signal corresponding to each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions.

[0058] In an embodiment of the present application, the electronic device may simulate different sound image positions of each dry sound signal in a virtual three-dimensional space, and then generate a virtual sound signal based on the dry sound signal corresponding to each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions. The virtual sound signal may be, for example, a two-channel signal. The N virtual three-dimensional sound image positions are different, and each virtual three-dimensional sound image position may correspond to one or more dry sound signals.

[0059] In one embodiment, N virtual three-dimensional spatial sound image positions can be simulated in the virtual three-dimensional space in the following manner: Figure 3aAs shown, the positive directions of the x, y, and z axes in the virtual three-dimensional space correspond to the front, left side, and top of the head, respectively, dividing the virtual three-dimensional space into three planes: a horizontal plane 301, an upper plane 302 whose angle with the horizontal plane is a first angle threshold, and a lower plane 303 whose angle with the horizontal plane is a second angle threshold. Figure 3b As shown, each virtual three-dimensional space sound image position in the virtual three-dimensional space includes an azimuth angle and an elevation angle. Assume that θ represents the azimuth angle of the virtual three-dimensional space sound image position, and θ represents the elevation angle. Represents the elevation angle of the virtual three-dimensional space sound image position, then each virtual three-dimensional space sound image position can be Represents. Accordingly, the horizontal plane 301 is the plane corresponding to the 0° elevation angle, the upper plane is the plane corresponding to the first angle threshold elevation angle, and the first angle threshold can be any angle value above the horizontal plane, and the lower plane is the plane corresponding to the second angle threshold elevation angle, and the second angle threshold can be any angle value below the horizontal plane. For example, the upper plane can be the plane corresponding to the 40° elevation angle, and the lower plane can be the plane corresponding to the -40° elevation angle. Among them, the azimuth angle θ can be used to describe the angle between the virtual three-dimensional sound and image position on the plane in a clockwise direction to the target direction line. Further, as Figure 3c As shown, by dividing the planes corresponding to different elevation angles at intervals of corresponding preset angles, multiple virtual three-dimensional sound image positions can be obtained. Specifically, by dividing the horizontal plane at intervals of a first preset angle, n1 virtual three-dimensional sound image positions can be obtained on the horizontal plane; by dividing the upper plane at intervals of a second preset angle, n2 virtual three-dimensional sound image positions can be obtained on the upper plane; and by dividing the lower plane at intervals of a third preset angle, n3 virtual three-dimensional sound image positions can be obtained on the lower plane. For example, assuming the first preset angle is 10°, and the second and third preset angles are both 15°, then dividing the horizontal plane at intervals of 10° will result in 36 virtual three-dimensional sound image positions, dividing the upper plane at intervals of 15° will result in 24 virtual three-dimensional sound image positions, and dividing the lower plane at intervals of 15° will also result in 24 virtual three-dimensional sound image positions, for a total of 84 different virtual three-dimensional sound image positions. It should be noted that the first, second, and third preset angles in the embodiments of the present application can be any preset angle values. The specific values ​​of the three preset angles are provided for example purposes only and do not limit the embodiments of the present application. In this way, multiple virtual three-dimensional sound and image positions can be created at varying azimuth angles within three different planes in the virtual three-dimensional space, achieving a fully immersive simulation of the sound source.

[0060] In one embodiment, each virtual three-dimensional space sound image position may correspond to one dry sound signal or multiple dry sound signals. The electronic device may obtain a virtual sound signal corresponding to one or more dry sound signals at each virtual three-dimensional space sound image position. Specifically, the electronic device may adopt the following scheme to obtain the virtual sound signal corresponding to each dry sound signal in the virtual three-dimensional space: obtain the azimuth and elevation of the virtual three-dimensional space sound image position corresponding to the dry sound signal, determine the head-related transfer function HRTF corresponding to the virtual three-dimensional space sound image position according to the azimuth and elevation of the virtual three-dimensional space sound image position, and calculate the virtual sound signal corresponding to the dry sound signal at the virtual space sound image position according to the azimuth and elevation of the virtual three-dimensional space sound image position and the corresponding HRTF data. For example, the azimuth and elevation of the virtual three-dimensional space sound image position corresponding to the dry sound signal X are The HRTF data expression corresponding to the virtual three-dimensional sound image position is:

[0061]

[0062] The virtual sound signal corresponding to the dry sound signal X at the virtual three-dimensional sound image position is calculated to be the two-channel signal Y L and Y R , where Y L is the left channel signal, Y R For the right channel signal.

[0063] In one embodiment, the electronic device can obtain part of the dry sound signal from the dry sound signal set, for example, it can randomly obtain or filter out dry sound signals with better pitch and sound quality according to new screening rules, and perform delay processing operations on the filtered dry sound signals to obtain delayed dual-channel signals corresponding to each dry sound signal in the dry sound signal. Specifically, when performing delay processing operations on a dry sound signal, 8 pairs of different time parameters can be selected. It should be noted that the 8 pairs of time parameters represent 8 time parameters for obtaining delayed left channel signals and 8 time parameters for obtaining delayed right channel signals, and a total of 16 time parameters will be selected. For example, 16 parameters of varying lengths can be selected from the range of 21ms to 79ms (80ms is selected as the reverberation time based on the general room impulse response) as time parameters, or 16 (or other values) of varying lengths can be randomly selected from other reasonable ranges as time parameters according to actual needs. In this way, the dry sound signal located at the left ear or the right ear of the human head can be simulated to make the audio effect richer. In one embodiment, the selection method and the setting of the time parameters (delay duration parameters) during the delay processing operation can be adjusted through a single interface to facilitate flexible configuration by users producing chorus audio. It should be noted that the above-mentioned steps of obtaining the virtual sound signal, obtaining the delayed left channel signal, and obtaining the delayed right channel signal can be performed simultaneously or sequentially, and this application does not limit this.

[0064] S203: merging the virtual sound signals in the virtual sound signal set to obtain a dry chorus sound.

[0065] In an embodiment of the present application, the electronic device may combine and process the individual virtual sound signals in the virtual sound signal set to obtain a dry chorus sound.

[0066] In one embodiment, the merging of the virtual sound signals corresponding to the dry sound signals can be achieved through normalization to adjust the loudness of the merged virtual sound signals to [-1dB, 1dB]. The virtual sound signals targeted during the merging include: a virtual sound signal corresponding to the dry sound signal at each of the N acquired virtual three-dimensional spatial sound image positions, and delayed two-channel signals obtained by delaying some of the dry sound signals in the dry sound signal set.

[0067] S203: Performing sound effect synthesis processing on the dry chorus voice and the background music of the target song according to the sound effect optimization rule to obtain an accompaniment.

[0068] In an embodiment of the present application, the electronic device synthesizes the chorus dry vocals and the background music of the target song according to a sound effect optimization rule to obtain a final accompaniment. The sound effect optimization rule may, for example, adjust the sound parameters of the virtual sound signal corresponding to the background music of the target song and the multiple dry vocal signals obtained above. The sound parameters may be common adjustable parameters such as loudness and timbre.

[0069] In one embodiment, after obtaining the dry chorus sound, the electronic device can obtain the background music of the target song. If the energy relationship between the obtained dry chorus sound and the background music of the target song does not meet the energy ratio condition, the electronic device can adjust the energy relationship between the dry chorus sound and the background music of the target song. Here, the energy ratio condition can be set to the ratio between the energy value of the dry chorus sound and the energy value of the background music of the target song being less than a ratio threshold, or it can be set to the loudness of the dry chorus sound being 3dB lower than the loudness of the background music of the target song. In this way, it can avoid the energy of the dry chorus sound being greater than the background music of the target song, making the final accompaniment more harmonious.

[0070] By implementing the embodiments of the present application, virtual sound signals corresponding to various dry sound signals at different virtual three-dimensional sound and image positions in the virtual three-dimensional space can be obtained, and then the various virtual sound signals are merged and processed to obtain the chorus dry sound. The chorus dry sound and the background music of the target song are then synthesized according to the sound effect optimization rules to obtain the accompaniment of the target song, thereby achieving a three-dimensional surround effect in the audio listening sense, and enhancing the immersiveness of the audio effect, so that the user experience is excellent.

[0071] For further information, see Figure 4 , Figure 4 This is a flow chart of another method for generating an accompaniment provided in an embodiment of the present application. The method in the embodiment of the present application can be applied to electronic devices, such as smartphones, tablet computers, smart wearable devices, personal computers, servers, etc. The method may include but is not limited to the following steps:

[0072] S401: Acquire an initial dry sound signal set from an audio database.

[0073] In an embodiment of the present application, the electronic device may obtain an initial dry sound signal set from an audio database. It should be noted that the initial dry sound signal set in the audio database is recorded with the authorization and consent of the user.

[0074] In one embodiment, the audio database can be an independent database or integrated with the electronic device, that is, the audio database can be considered to be stored within the electronic device. Here, the initial dry sound signal set refers to the set of original dry sound signals in the audio database that are recorded with authorization when the user sings the same song.

[0075] S402: Filter out dry sound signals from the initial dry sound signal set according to sound parameters of the initial dry sound signals, and the filtered dry sound signals constitute the dry sound signal set.

[0076] In an embodiment of the present application, the electronic device may filter out dry sound signals that meet the conditions from the initial dry sound signal set according to the sound parameters of each initial dry sound signal, so as to narrow down the initial dry sound signal set to form a dry sound signal set.

[0077] In one embodiment, the sound parameters of the initial dry sound signal may include pitch characteristic parameters and sound quality characteristic parameters of the initial dry sound signal. The pitch characteristic parameters may include any one or more of pitch parameters, rhythm parameters, and rhythmic parameters. The sound quality characteristic parameters may include any one or more of noise parameters, energy parameters, and speed parameters. In this way, the initial dry sound signals with poor audio effects such as noisy, out-of-tune, short audio time, low audio energy, and popping sounds can be removed from the initial dry sound signal set to obtain a dry sound signal set with excellent pitch and sound quality.

[0078] S403: Obtaining a head-related transfer function corresponding to each of the N virtual three-dimensional sound image positions.

[0079] In an embodiment of the present application, the electronic device can obtain N virtual three-dimensional spatial sound image positions in the virtual three-dimensional space, and then obtain the head-related transfer function corresponding to each virtual three-dimensional spatial sound image position based on the N virtual three-dimensional spatial sound image positions.

[0080] In one embodiment, the head-related transfer functions corresponding to each virtual three-dimensional sound image position in the virtual three-dimensional space can be pre-saved in a head-related transfer function database, so that the electronic device can call the corresponding head-related transfer function from the head-related transfer function database according to the virtual three-dimensional sound image position.

[0081] S404: Processing the dry sound signal corresponding to the target virtual three-dimensional space sound image position by using the head-related transfer function corresponding to the target virtual three-dimensional space sound image position to obtain a virtual sound signal at the target virtual three-dimensional space sound image position.

[0082] In an embodiment of the present application, the electronic device can process the target dry sound signal according to the head-related transfer function corresponding to the target virtual three-dimensional space sound and image position to obtain a virtual sound signal corresponding to the target dry sound signal at the target virtual three-dimensional space sound and image position. The target virtual three-dimensional space sound and image position can be any one of the N virtual three-dimensional space sound and image positions, and the target dry sound signal can be any one of the dry sound signals in the dry sound signal set.

[0083] In one embodiment, the head-related transfer function corresponding to the target virtual three-dimensional spatial sound image position is HRTF data corresponding to the target virtual three-dimensional spatial sound image position. The HRTF data corresponding to the target virtual three-dimensional spatial sound image position can be determined from known HRTF data based on the azimuth and elevation angles of the target virtual three-dimensional spatial sound image position. The electronic device can then convolve the target dry sound signal with the HRTF data corresponding to the target virtual three-dimensional spatial position to obtain a virtual sound signal corresponding to the target dry sound signal at the target virtual three-dimensional spatial sound image position.

[0084] S405: Obtain p dry sound signals from the x dry sound signals included in the dry sound signal set.

[0085] In the embodiment of the present application, the electronic device may randomly obtain p dry sound signals from the x dry sound signals included in the dry sound signal set. It should be noted that S404 and S405 may be performed simultaneously or sequentially, and this application does not limit this.

[0086] S406: Perform a delay processing operation on each of the p dry sound signals to obtain a delayed left channel signal and a delayed right channel signal corresponding to each of the p dry sound signals.

[0087] In an embodiment of the present application, the electronic device can perform a delay processing operation of m1 time parameters on each of the p dry sound signals to obtain m1 delayed dry sound signals corresponding to each of the p dry sound signals, and obtain a delayed left channel signal corresponding to each of the p dry sound signals by superimposing the m1 delayed dry sound signals corresponding to each dry sound signal, where m1 is a positive integer; and then perform a delay processing operation of m2 time parameters on each of the p dry sound signals to obtain m2 delayed dry sound signals corresponding to each of the p dry sound signals, and obtain a delayed right channel signal corresponding to each of the p dry sound signals by superimposing the m2 delayed dry sound signals corresponding to each dry sound signal, where m2 is a positive integer.

[0088] In one embodiment, the electronic device can process a dry sound signal through 16 delay devices with different time parameters to obtain 16 dry sound signals with different delays and attenuation degrees, and then divide the 16 dry sound signals with different delays and attenuation degrees into two groups on average, and superimpose the dry sound signals with different delays and attenuation degrees in each group respectively, and finally obtain the delayed left channel signal and delayed right channel signal corresponding to the dry sound signal.

[0089] In one embodiment, before obtaining the delayed two-channel signal corresponding to each of the p dry sound signals, the sound field of the dry sound signal can be widened by adding a bass enhancement and reverberation simulation module to reduce the correlation between the delayed left channel signal and the delayed right channel signal in the two-channel signal obtained by delay processing. It should be noted that the steps of S403 and S404 for obtaining the virtual sound signal and the steps of S405 and S406 for obtaining the delayed left channel signal and the delayed right channel signal can be performed simultaneously or successively, and this application does not limit this. Among them, S405 and S406 are two optional steps.

[0090] S407: merging the virtual sound signals in the virtual sound signal set to obtain a dry chorus sound.

[0091] In an embodiment of the present application, the electronic device may combine and process each virtual sound signal in a virtual sound signal set to obtain a chorus dry sound. Here, the virtual sound signal set includes virtual sound signals corresponding to each dry sound signal obtained by the electronic device by simulating N virtual three-dimensional spatial positions, and a delayed two-channel signal obtained by the electronic device by performing a delay processing operation on p dry sound signals in the dry sound signal set.

[0092] In one embodiment, each virtual sound signal in the virtual sound signal set is a two-channel signal, and the two-channel signal includes a left channel signal and a right channel signal. Merging the virtual sound signals allows the left channel signal and the right channel signal to be processed separately, and the left channel signal and the right channel signal are subject to the same processing rule. Here, the merging process can be implemented by normalization processing so that the loudness of the two-channel signal after merging is [-1dB, 1dB]. For example, assuming there are 1000 two-channel signals, including 1000 left channel signals and 1000 right channel signals, each left channel signal is normalized separately, and the sum of the 1000 normalized left channel signals is divided by 1000 to obtain the merged left channel signal. Similarly, each right channel signal is normalized, and the sum of the 1000 normalized right channel signals is divided by 1000 to obtain the merged right channel signal. In this way, you can get a dry chorus sound.

[0093] In one embodiment, the obtained energy relationship between the dry chorus and the background music of the target song may or may not satisfy the energy ratio condition. If the obtained energy relationship between the dry chorus and the background music satisfies the energy ratio condition, step S408 may be omitted; correspondingly, if the obtained energy relationship between the dry chorus and the background music does not satisfy the energy ratio condition, step S408 is executed.

[0094] S408: Obtain the background music of the target song, and adjust the energy relationship between the dry chorus sound and the background music.

[0095] In an embodiment of the present application, the electronic device can target the background music of the song and adjust the energy relationship between the dry chorus sound and the corresponding background music, and the energy relationship between the adjusted dry chorus sound and the adjusted background music meets the energy ratio condition.

[0096] In one embodiment, the energy of the dry chorus may be too high, causing it to overwhelm the energy of the background music. By adjusting the dry chorus and the background music, the energy relationship between the adjusted dry chorus and the adjusted background music can be made to meet an energy ratio condition. In this way, the situation where the energy of the dry chorus is too high can be addressed. The energy ratio condition can be set to the ratio between the energy value of the dry chorus and the energy value of the background music being less than a ratio threshold, or it can be set to the loudness of the dry chorus being 3dB lower than the loudness of the background music.

[0097] In one embodiment, after obtaining the background music of the target song, a detailed description of the virtual sound signal can be generated based on the dry sound signal corresponding to each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions according to the above S202, and the background music can also be processed in the same way to obtain chorus dry sound and background music with similar effects, thereby achieving a more harmonious and unified auditory experience.

[0098] S409: Perform spectrum equalization processing on the chorus dry sound at a preset frequency band.

[0099] In an embodiment of the present application, the electronic device may perform spectrum equalization processing on the chorus dry sound in a preset frequency band.

[0100] In one embodiment, the electronic device can achieve spectrum equalization by adding a spectrum notch in a preset frequency band. For example, the electronic device can add a spectrum notch of approximately 6dB around 4kHz. This makes the dry chorus sound more natural and prevents high-frequency current noise caused by spectral disharmony.

[0101] S410: Obtain the loudness of background music.

[0102] In an embodiment of the present application, the electronic device can obtain the loudness of the background music.

[0103] S411: If the loudness is less than the loudness threshold, the loudness of the adjusted background music is increased to the loudness threshold.

[0104] In an embodiment of the present application, if the loudness of the background music is less than the loudness threshold, the electronic device may increase the loudness of the background music to the loudness threshold. For example, the loudness threshold may be set to -14dB. If the loudness of the background music is less than -14dB, the electronic device may increase it to -14dB.

[0105] S412: Get the accompaniment.

[0106] In an embodiment of the present application, the electronic device superimposes the dry chorus sound and the background music to obtain the final accompaniment. It should be noted that the accompaniment can be obtained according to any one step or a combination of multiple steps in S408 to S411, and, in one embodiment, S408 to S411 can be selectively executed according to actual needs. For example, there may be a situation where the energy relationship between the dry chorus sound and the background music does not need to be adjusted, so S408 is not executed. Similarly, it is also optional to perform spectral equalization processing on the dry chorus sound at a preset frequency band. For another example, steps S410 and S411 may also not be executed. In Figure 4 It only points out the technical solutions adopted to make the accompaniment more harmonious and natural and the sound quality better. In addition, the energy relationship adjustment in S408, the spectrum equalization adjustment in S409, and the loudness adjustment embodied in S410 and S411 are performed. The order of these three aspects is not limited in this application.

[0107] In one implementation, after the final accompaniment is obtained, it can be stored in a database so that the electronic device can directly obtain the corresponding accompaniment from the database when receiving a chorus request for the same song.

[0108] By implementing the embodiments of the present application, virtual sound signals corresponding to each dry sound signal can be obtained by simulating N virtual three-dimensional spatial positions. Delayed two-channel signals corresponding to each dry sound signal can also be obtained by delaying each dry sound signal, thereby enriching the chorus dry sound. Furthermore, by adjusting the energy relationship between the chorus dry sound and the background music, the resulting accompaniment sounds more harmonious and natural, allowing users to clearly experience the sense of space and immersion during the chorus.

[0109] For further information, see Figure 5 , Figure 5This is a flow chart of obtaining a virtual sound signal in a method for generating an accompaniment provided in an embodiment of the present application. Obtaining the virtual sound signal includes: obtaining a virtual sound signal corresponding to a dry sound signal at each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions, and performing a delay processing operation on each of the p dry sound signals to obtain a delayed two-channel signal corresponding to each of the p dry sound signals.

[0110] In an embodiment of the present application, after obtaining the dry sound signal set, a virtual sound signal corresponding to the dry sound signal at each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions can be obtained, and a delay processing operation can be performed on each of the p dry sound signals in the dry sound signal set to obtain a delayed dual-channel signal corresponding to each of the p dry sound signals.

[0111] As an example, Figure 5 As shown, the dry sound signal X and the dry sound signal W included in the dry sound signal set are respectively used to obtain corresponding virtual sound signals through the above two methods, wherein the dry sound signal X and the dry sound signal W can be any dry sound signal in the dry sound signal set. Specifically, after obtaining the dry sound signal X in the dry sound signal set, the electronic device can describe the position information of the virtual three-dimensional space sound image position according to the azimuth and elevation angle of the virtual three-dimensional space sound image position, that is, Then, according to the position information of the virtual three-dimensional sound image position The head-related transfer function HRTF corresponding to the virtual three-dimensional sound image position can be found The head-related transfer function HRTF corresponding to the dry sound signal X and the virtual three-dimensional sound image position is calculated. After the convolution operation, the virtual sound signal corresponding to the dry sound signal at the virtual three-dimensional sound image position can be obtained. The virtual sound signal is a two-channel signal, including the left channel signal.

[0112] Y L and right channel signal Y R The virtual sound signal corresponding to the dry sound signal obtained in this way can enhance the user's three-dimensional immersion.

[0113] In addition, after obtaining the dry sound signal W in the dry sound signal set, the electronic device can perform a delay processing operation on the dry sound signal W. For example, the electronic device can delay the dry sound signal W through d L (1) d L (2), ..., d L (8) and d R (1) d R (2), ..., d R(8) A total of 16 delay devices with different time parameters are used for delay processing, and then the delay time is calculated by d L (1) d L (2), ..., d L (8) The eight dry sound signals obtained by the eight delay devices are superimposed to obtain the delayed left channel signal W corresponding to the dry sound signal W. L , will be passed d R (1) d R (2), ..., d R (8) The eight dry sound signals obtained by the eight delay devices are superimposed to obtain the delayed right channel signal W corresponding to the dry sound signal W. R The delayed two-channel signal corresponding to the dry sound signal obtained in this way can simulate the two-channel signal at the left or right ear of the human head, enriching the user's listening experience.

[0114] In one embodiment, the final virtual sound signal set includes the above two cases, that is, the final virtual sound signal set is: Z = {Z L , Z R}, Z L =Y L +W L ; Z R =Y R +W L . It should be noted that the above steps of obtaining the virtual sound signal, obtaining the delayed left channel signal and the delayed right channel signal can be performed simultaneously or successively, and this application does not limit this. The dry sound signal in the dry sound signal set is used in the above two different ways to obtain the corresponding virtual sound signal, which can fully display the scene experience during chorus and make the audio effect richer.

[0115] For further information, see Figure 6 , Figure 6 This is a method for playing an accompaniment provided in an embodiment of the present application. The method of the embodiment of the present application can be applied to an electronic device, such as a smart phone, tablet computer, smart wearable device, personal computer, or other smart device, or a server, etc. The method may include but is not limited to the following steps:

[0116] S601: Displaying a user interface.

[0117] In an embodiment of the present application, the electronic device may display a user interface for receiving a user's selection instruction for a target song.

[0118] In one embodiment, the selection instruction may include a selection instruction for the accompaniment mode of the target song. The accompaniment mode of the target song may be a chorus accompaniment mode, an acoustic accompaniment mode, and an artificial intelligence (AI) accompaniment mode, but is not limited thereto.

[0119] In one embodiment, the selection instruction may be an instruction generated by the user by triggering a selection control displayed on a user interface, or may be a selection instruction generated by the user by controlling the electronic device through voice. For example, the user's voice control of the electronic device may be "Please play in chorus accompaniment mode." In this way, the electronic device may generate a selection instruction indicating that the accompaniment mode for the target song is a chorus accompaniment mode.

[0120] S602: If the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode, then the accompaniment corresponding to the target song is obtained.

[0121] In an embodiment of the present application, if the electronic device receives a selection instruction in the user interface indicating that the accompaniment mode for the target song is a chorus accompaniment mode, it can obtain the corresponding accompaniment in the chorus accompaniment mode of the target song.

[0122] In one embodiment, a user interface may display a selection control for an accompaniment mode of a target song, and the selection control may include a chorus accompaniment mode selection control and an acoustic accompaniment mode selection control. Before obtaining the accompaniment corresponding to the target song, the electronic device may detect whether a selection operation for the chorus accompaniment mode selection control has been received, and if so, confirm that the selection instruction received on the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode.

[0123] In one embodiment, the corresponding accompaniment in the chorus accompaniment mode is generated based on the chorus dry sound and background music. Among them, the chorus dry sound can be generated based on a virtual sound signal set, and the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional space sound image position in N virtual three-dimensional space sound image positions generated based on the acquired dry sound signal set, and the multiple dry sound signals in the dry sound signal set can correspond to multiple different virtual three-dimensional space sound image positions, and each virtual three-dimensional space sound image position can correspond to one or more dry sound signals. The dry sound signal set is obtained based on the dry sound signals recorded by multiple users for the target song. It should be noted that the user's dry sound signal for the target song is recorded after the user's authorization and consent. Specifically, the method for generating the corresponding accompaniment in the chorus accompaniment mode can be seen above Figure 2-Figure 5 The embodiments shown will not be described in detail here.

[0124] S603: Play the accompaniment corresponding to the target song.

[0125] In an embodiment of the present application, after the electronic device obtains the corresponding accompaniment in the chorus accompaniment mode of the target song, it can play the accompaniment to the user.

[0126] In one implementation, the accompaniment corresponding to the target song can be applied to a karaoke scene, and the user can sing while the accompaniment is playing. With the user's authorization and consent, the electronic device can collect the user's singing voice and integrate it with the accompaniment corresponding to the target song before playing it, giving the user a unique experience as if they were at a concert.

[0127] In one embodiment, Figure 7a As shown, the electronic device may obtain the accompaniment corresponding to the target song, including but not limited to the following steps:

[0128] S701: Send an accompaniment request to the server.

[0129] In an embodiment of the present application, the electronic device may send an accompaniment request to the server, and the accompaniment request may include identification information of the target song.

[0130] In one embodiment, the identification information of the target song is information used to uniquely identify the target song. For example, the identification information may be the song title of the target song.

[0131] S702: Receive the chorus dry voice and background music returned by the server in response to the accompaniment request.

[0132] In an embodiment of the present application, the electronic device can receive the dry chorus and background music returned by the server in response to the accompaniment request for the target song.

[0133] In one embodiment, the server may return the dry chorus and the background music separately, or may combine the dry chorus and the background music before returning them. The specific return method may be selected according to the user's settings.

[0134] S703: Determine a target chorus dry voice segment from the chorus dry voice.

[0135] In an embodiment of the present application, the electronic device may determine a target chorus dry voice segment based on the returned chorus dry voice.

[0136] In one embodiment, the electronic device may display a first single sentence interface, such as Figure 7b As shown, the first single sentence interface displays each single sentence in the text data corresponding to the chorus dry voice according to the time playback node order of the chorus dry voice. The user can select the target chorus dry voice segment based on each single sentence displayed in the first single sentence interface.

[0137] In one embodiment, the target dry chorus segment may be composed of a portion of the dry chorus sentences, or may be composed of all the dry chorus sentences, which may be determined by a user's selection operation.

[0138] S704: Obtain an accompaniment corresponding to the target song according to the chorus dry voice and background music corresponding to the target chorus dry voice segment.

[0139] In an embodiment of the present application, the electronic device can obtain the accompaniment corresponding to the target song based on the chorus dry voice and background music corresponding to the target chorus dry voice segment selected by the user.

[0140] In one embodiment, the electronic device may display a second single sentence interface, such as Figure 7c As shown, the second single sentence interface can be displayed during the playing of the accompaniment corresponding to the target song, and each single sentence in the text data corresponding to the accompaniment can be displayed in the order of the time playing nodes of the accompaniment.

[0141] In one embodiment, the electronic device can also detect whether a mute selection operation for the dry chorus in the accompaniment is obtained during the playback process. If a mute selection operation for the dry chorus in the accompaniment is received from the user, the playback of the dry chorus in the accompaniment can be canceled at the current time playback node, and only the playback of the background music in the accompaniment can be retained.

[0142] By implementing the embodiments of the present application, on the one hand, a user's selection instruction for a target song can be received. When the user's selection instruction indicates that the accompaniment mode for the target song is a chorus accompaniment mode, the accompaniment corresponding to the target song can be obtained and played. On the other hand, the accompaniment of the target song in the chorus accompaniment mode is generated based on the chorus dry voice and background music. The target chorus segment can be determined from the chorus dry voice, and the accompaniment corresponding to the target song can be generated based on the chorus dry voice and background music corresponding to the target chorus dry voice segment. In this way, when playing the accompaniment corresponding to the target song, the user can have the experience of being at a concert and have an immersive sense of listening. In addition, the user can also flexibly select the chorus dry voice in the accompaniment, which enhances the fun of the accompaniment and improves the user experience.

[0143] For further information, see Figure 8a , Figure 8a is a schematic diagram of the structure of an accompaniment generation device provided in an embodiment of the present application. The device in the embodiment of the present application can be applied to an electronic device, such as a smart phone, a tablet computer, a smart wearable device, a personal computer, a server, etc. In one embodiment, Figure 8a As shown, the accompaniment generating device 80 may include:

[0144] An acquisition unit 801 is used to acquire a dry sound signal set, where the dry sound signal set includes x dry sound signals corresponding to a target song, where x is an integer greater than 1; and a virtual sound signal is generated based on the dry sound signal corresponding to each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions, where the x dry sound signals correspond to N virtual three-dimensional sound image positions, where N is an integer greater than 1, the N virtual three-dimensional sound image positions are different, and each virtual three-dimensional sound image position is allowed to correspond to one or more dry sound signals in the x dry sound signals.

[0145] The processing unit 802 is used to merge the various virtual sound signals in the virtual sound signal set to obtain the dry chorus sound, where the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions; and perform sound effect synthesis processing on the dry chorus sound and the background music of the target song according to the sound effect optimization rules to obtain the accompaniment of the target song.

[0146] In one embodiment, the acquisition unit 801 can also be used to obtain an initial dry sound signal set from an audio database, which includes initial dry sound signals recorded when multiple users sing the same song; the processing unit 802 can also be used to filter out dry sound signals from the initial dry sound signal set based on the sound parameters of each initial dry sound signal, and the filtered dry sound signals constitute the dry sound signal set.

[0147] In one embodiment, the dry sound signal set includes: dry sound signals filtered out from the initial dry sound signal set based on pitch characteristic parameters and sound quality characteristic parameters; the pitch characteristic parameters include any one or more of pitch parameters, rhythm parameters and prosody parameters; the sound quality characteristic parameters include any one or more of noise parameters, energy parameters and speed parameters.

[0148] In one embodiment, the N virtual three-dimensional spatial sound and image positions include: n1 virtual three-dimensional spatial sound and image positions on the horizontal plane obtained by dividing the horizontal plane at intervals of a first preset angle; n2 virtual three-dimensional spatial sound and image positions on the upper plane obtained by dividing the upper plane at intervals of a second preset angle; the angle between the upper plane and the horizontal plane is a first angle threshold; n3 virtual three-dimensional spatial sound and image positions on the lower plane obtained by dividing the lower plane at intervals of a third preset angle; the angle between the lower plane and the horizontal plane is a second angle threshold; wherein n1, n2 and n3 are positive integers and the sum is equal to N.

[0149] In one embodiment, the acquisition unit 801 can also be used to obtain the head-related transfer function corresponding to each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions; the processing unit 802 can also be used to process the dry sound signal corresponding to the target virtual three-dimensional sound image position through the head-related transfer function corresponding to the target virtual three-dimensional sound image position to obtain the virtual sound signal at the target virtual three-dimensional sound image position; the virtual sound signal at the target virtual three-dimensional sound image position is a two-channel signal; the target virtual three-dimensional sound image position is any virtual three-dimensional sound image position among the N virtual three-dimensional sound image positions.

[0150] In one embodiment, the virtual sound signal set also includes: a delayed left channel signal and a delayed right channel signal corresponding to each of the p dry sound signals; the acquisition unit 801 can also be used to obtain p dry sound signals from the x dry sound signals included in the dry sound signal set, where p is a positive integer and is less than or equal to x; the processing unit 802 can also be used to perform a delay processing operation of m1 time parameters on each of the p dry sound signals to obtain m1 delayed dry sound signals corresponding to each of the p dry sound signals, and obtain a delayed left channel signal corresponding to each of the p dry sound signals by superimposing the m1 delayed dry sound signals corresponding to each dry sound signal, where m1 is a positive integer; perform a delay processing operation of m2 time parameters on each of the p dry sound signals to obtain m2 delayed dry sound signals corresponding to each of the p dry sound signals, and obtain a delayed right channel signal corresponding to each of the p dry sound signals by superimposing the m2 delayed dry sound signals corresponding to each dry sound signal, where m2 is a positive integer.

[0151] In one embodiment, the acquisition unit 801 can also be used to obtain the background music of the target song, and the processing unit 802 can also be used to adjust the energy relationship between the dry chorus sound and the background music, and the energy relationship between the adjusted dry chorus sound and the adjusted background music meets the energy ratio condition; the accompaniment is obtained based on the adjusted dry chorus sound and the background music.

[0152] In one embodiment, the processing unit 802 can also be used to perform spectral equalization processing on the dry chorus sound at a preset frequency band; the acquisition unit 801 can also be used to obtain the loudness of the background music; the processing unit 802 can also be used to increase the loudness of the background music to the loudness threshold if the loudness of the background music is less than the loudness threshold; the accompaniment is obtained based on the dry chorus sound after spectral equalization processing and the background music after loudness processing.

[0153] It should be noted that Figure 8a The details not mentioned in the corresponding embodiments and the specific implementation of each step can be found in Figure 2-Figure 5The illustrated embodiments and the aforementioned contents will not be described in detail here.

[0154] For further information, see Figure 8b , Figure 8b is a structural diagram of an accompaniment playback processing device provided in an embodiment of the present application. The device in the embodiment of the present application can be applied to an electronic device, such as a smart phone, a tablet computer, a smart wearable device, a personal computer, a server, etc. In one embodiment, Figure 8b As shown, the accompaniment playing processing device 81 may include:

[0155] The acquisition unit 811 is used to display a user interface, which is used to receive a selection instruction for a target song; if the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode, the accompaniment corresponding to the target song is acquired.

[0156] Processing unit 812 is used to play the accompaniment corresponding to the target song; the accompaniment is generated based on the chorus dry sound and background music, the chorus dry sound is generated based on multiple dry sound signals in a dry sound signal set, the multiple dry sound signals in the dry sound signal set correspond to multiple different virtual three-dimensional space sound and image positions, and the dry sound signal set is obtained based on the dry sound signals recorded by multiple users for the target song.

[0157] In one embodiment, the chorus dry sound is generated based on a virtual sound signal set, and the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional space sound image position in N virtual three-dimensional space sound image positions generated based on the acquired dry sound signal set; wherein, multiple dry sound signals in the dry sound signal set correspond to N virtual three-dimensional space sound image positions, N is an integer greater than 1, the N virtual three-dimensional space sound image positions are different, and each virtual three-dimensional space sound image position is allowed to correspond to one or more dry sound signals.

[0158] In one embodiment, a selection control for the accompaniment mode of the target song is displayed on the user interface, and the accompaniment mode selection control includes: a chorus accompaniment mode selection control and an acoustic accompaniment mode selection control; before obtaining the accompaniment corresponding to the target song, the processing unit 812 can also be used to detect whether a selection operation for the chorus accompaniment mode selection control is obtained; if so, it is confirmed that the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode.

[0159] In one embodiment, the processing unit 812 can also be used to send an accompaniment request to the server, which accompaniment request includes identification information of the target song; the acquisition unit 811 can also be used to receive the chorus dry voice and background music returned by the server in response to the accompaniment request; the processing unit 812 can also be used to determine the target chorus dry voice segment from the chorus dry voice; and obtain the accompaniment corresponding to the target song based on the chorus dry voice and background music corresponding to the target chorus dry voice segment.

[0160] In one embodiment, the processing unit 812 can also be used to display a first single sentence interface, displaying each single sentence in the text data corresponding to the chorus dry voice in the order of the time playback nodes of the chorus dry voice; the target chorus dry voice segment is determined based on the single sentence selection operation on the first single sentence interface.

[0161] In one embodiment, the processing unit 812 can also be used to display a second single sentence interface, displaying each single sentence in the text data corresponding to the accompaniment in the order of the time playback node of the accompaniment; detecting whether a mute selection operation for the dry chorus in the accompaniment is obtained during the playback process; if so, canceling the playback of the dry chorus at the current time playback node.

[0162] It should be noted that Figure 8b The details not mentioned in the corresponding embodiments and the specific implementation of each step can be found in Figure 2 -The embodiment shown in FIG7 and the aforementioned contents will not be described in detail here.

[0163] For further information, see Figure 9 , Figure 9: This is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may include: a network interface 901, a memory 902 and a processor 903. The network interface 901, the memory 902 and the processor 903 are connected via one or more communication buses, and the communication bus is used to realize the connection and communication between these components. The network interface 901 may include a standard wired interface and a wireless interface (such as a WIFI interface). The memory 902 may include a volatile memory (volatile memory), such as a random-access memory (RAM); the memory 902 may also include a non-volatile memory (non-volatile memory), such as a flash memory (flash memory), a solid-state drive (SSD), etc.; the memory 902 may also include a combination of the above types of memory. The processor 903 may be a central processing unit (CPU). The processor 903 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), etc. The PLD may be a field-programmable gate array (FPGA), a generic array logic (GAL), or the like.

[0164] Optionally, the memory 902 is further configured to store program instructions, and the processor 903 may further call the program instructions to implement:

[0165] Obtain a dry sound signal set, where the dry sound signal set includes x dry sound signals corresponding to the target song, where x is an integer greater than 1;

[0166] generating a virtual sound signal based on a dry sound signal corresponding to each virtual three-dimensional spatial sound image position in N virtual three-dimensional spatial sound image positions, wherein the x dry sound signals correspond to the N virtual three-dimensional spatial sound image positions, N is an integer greater than 1, the N virtual three-dimensional spatial sound image positions are different, and each virtual three-dimensional spatial sound image position is allowed to correspond to one or more dry sound signals in the x dry sound signals;

[0167] Merging each virtual sound signal in a virtual sound signal set to obtain a dry chorus sound, the virtual sound signal set including: a virtual sound signal at each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions;

[0168] According to the sound effect optimization rules, the dry chorus voice and the background music of the target song are synthesized to obtain the accompaniment of the target song.

[0169] In one embodiment, the processor 903 can also call the program instructions to achieve: obtaining an initial dry sound signal set from an audio database, the audio database including initial dry sound signals recorded when multiple users sing the same song; filtering out dry sound signals from the initial dry sound signal set according to the sound parameters of each initial dry sound signal, and the filtered dry sound signals constitute a dry sound signal set.

[0170] In one embodiment, the dry sound signal set includes: dry sound signals filtered out from the initial dry sound signal set based on pitch characteristic parameters and sound quality characteristic parameters; the pitch characteristic parameters include any one or more of pitch parameters, rhythm parameters and prosody parameters; the sound quality characteristic parameters include any one or more of noise parameters, energy parameters and speed parameters.

[0171] In one embodiment, the N virtual three-dimensional spatial sound and image positions include: n1 virtual three-dimensional spatial sound and image positions on the horizontal plane obtained by dividing the horizontal plane at intervals of a first preset angle; n2 virtual three-dimensional spatial sound and image positions on the upper plane obtained by dividing the upper plane at intervals of a second preset angle; the angle between the upper plane and the horizontal plane is a first angle threshold; n3 virtual three-dimensional spatial sound and image positions on the lower plane obtained by dividing the lower plane at intervals of a third preset angle; the angle between the lower plane and the horizontal plane is a second angle threshold; wherein n1, n2 and n3 are positive integers and the sum is equal to N.

[0172] In one embodiment, the processor 903 may also call the program instructions to implement: obtaining the head-related transfer function corresponding to each virtual three-dimensional spatial sound image position among N virtual three-dimensional spatial sound image positions; processing the dry sound signal corresponding to the target virtual three-dimensional spatial sound image position through the head-related transfer function corresponding to the target virtual three-dimensional spatial sound image position to obtain a virtual sound signal at the target virtual three-dimensional spatial sound image position; the virtual sound signal at the target virtual three-dimensional spatial sound image position is a two-channel signal; the target virtual three-dimensional spatial sound image position is any virtual three-dimensional spatial sound image position among the N virtual three-dimensional spatial sound image positions.

[0173] In one embodiment, the virtual sound signal set also includes: a delayed left channel signal and a delayed right channel signal corresponding to each of the p dry sound signals; the processor 903 can also call the program instructions to implement: obtaining p dry sound signals from the x dry sound signals included in the dry sound signal set, where p is a positive integer and is less than or equal to x; performing a delay processing operation of m1 time parameters on each of the p dry sound signals to obtain m1 delayed dry sound signals corresponding to each of the p dry sound signals, and obtaining a delayed left channel signal corresponding to each of the p dry sound signals by superimposing the m1 delayed dry sound signals corresponding to each dry sound signal, where m1 is a positive integer; performing a delay processing operation of m2 time parameters on each of the p dry sound signals to obtain m2 delayed dry sound signals corresponding to each of the p dry sound signals, and obtaining a delayed right channel signal corresponding to each of the p dry sound signals by superimposing the m2 delayed dry sound signals corresponding to each dry sound signal, where m2 is a positive integer.

[0174] In one embodiment, the processor 903 can also call the program instructions to achieve: obtaining the background music of the target song, and adjusting the energy relationship between the dry chorus sound and the background music, so that the energy relationship between the adjusted dry chorus sound and the adjusted background music meets the energy ratio condition; the accompaniment is obtained based on the adjusted dry chorus sound and the background music.

[0175] In one embodiment, the processor 903 may also call the program instructions to implement: performing spectrum equalization processing on the dry chorus sound at a preset frequency band; obtaining the loudness of the background music; if the loudness of the background music is less than the loudness threshold, increasing the loudness of the background music to the loudness threshold; and the accompaniment is obtained based on the dry chorus sound after spectrum equalization processing and the background music after loudness processing.

[0176] Optionally, the memory 902 is further configured to store program instructions, and the processor 903 may further call the program instructions to implement:

[0177] displaying a user interface, the user interface being used to receive a selection instruction for a target song;

[0178] If the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode, then obtaining the accompaniment corresponding to the target song;

[0179] Play the accompaniment corresponding to the target song; the accompaniment is generated based on the chorus dry voice and background music, the chorus dry voice is generated based on multiple dry voice signals in a dry voice signal set, the multiple dry voice signals in the dry voice signal set correspond to multiple different virtual three-dimensional sound and image positions, and the dry voice signal set is obtained based on the dry voice signals recorded by multiple users for the target song.

[0180] In one embodiment, the chorus dry sound is generated based on a virtual sound signal set, and the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional space sound image position in N virtual three-dimensional space sound image positions generated based on the acquired dry sound signal set; wherein, multiple dry sound signals in the dry sound signal set correspond to N virtual three-dimensional space sound image positions, N is an integer greater than 1, the N virtual three-dimensional space sound image positions are different, and each virtual three-dimensional space sound image position is allowed to correspond to one or more dry sound signals.

[0181] In one embodiment, a selection control for the accompaniment mode of the target song is displayed on the user interface, and the accompaniment mode selection control includes: a chorus accompaniment mode selection control and an acoustic accompaniment mode selection control; before obtaining the accompaniment corresponding to the target song, the processor 903 can also call the program instruction to achieve: detecting whether a selection operation for the chorus accompaniment mode selection control is obtained; if so, confirming that the selection instruction received on the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode.

[0182] In one embodiment, the processor 903 can also call the program instructions to achieve: sending an accompaniment request to the server, the accompaniment request including identification information of the target song; receiving the dry chorus and background music returned by the server in response to the accompaniment request; determining the target dry chorus segment from the dry chorus; and obtaining the accompaniment corresponding to the target song based on the dry chorus and background music corresponding to the target dry chorus segment.

[0183] In one embodiment, the processor 903 can also call the program instructions to achieve: displaying a first single sentence interface, displaying each single sentence in the text data corresponding to the chorus dry voice in the order of the time playback nodes of the chorus dry voice; the target chorus dry voice segment is determined based on the single sentence selection operation on the first single sentence interface.

[0184] In one embodiment, the processor 903 can also call the program instructions to achieve: displaying a second single sentence interface, displaying each single sentence in the text data corresponding to the accompaniment in the order of the time playback node of the accompaniment; detecting whether a mute selection operation for the dry chorus in the accompaniment is obtained during the playback process; if so, canceling the playback of the dry chorus at the current time playback node.

[0185] It should be understood that the principles and beneficial effects of the electronic device 90 described in the embodiment of the present application are similar to those of the present application. Figure 2 - The principles and beneficial effects of the embodiment shown in FIG7 and the aforementioned content in solving the problem are similar, and for the sake of brevity, they will not be repeated here.

[0186] In addition, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the aforementioned embodiment is implemented.

[0187] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the aforementioned embodiment.

[0188] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.

[0189] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.

[0190] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0191] The above disclosure is only part of the embodiments of the present application, and it is certainly not intended to limit the scope of the rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of the above embodiments and equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.

Claims

1. A method for generating an accompaniment, characterized in that: The method comprises: Obtain a dry sound signal set, where the dry sound signal set includes x dry sound signals corresponding to the target song, where x is an integer greater than 1; generating a virtual sound signal based on a dry sound signal corresponding to each virtual three-dimensional spatial sound image position in N virtual three-dimensional spatial sound image positions, wherein the x dry sound signals correspond to the N virtual three-dimensional spatial sound image positions, N is an integer greater than 1, the N virtual three-dimensional spatial sound image positions are different, and each virtual three-dimensional spatial sound image position is allowed to correspond to one or more dry sound signals in the x dry sound signals; Merging each virtual sound signal in a virtual sound signal set to obtain a dry chorus sound, wherein the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional sound image position in N virtual three-dimensional sound image positions; Performing sound effect synthesis processing on the dry chorus sound and the background music of the target song according to a sound effect optimization rule to obtain the accompaniment of the target song; The step of synthesizing the dry chorus sound and the background music of the target song according to the sound effect optimization rule to obtain the accompaniment of the target song comprises: obtaining the background music of the target song, and adjusting the energy relationship between the dry chorus sound and the background music, wherein the energy relationship between the adjusted dry chorus sound and the adjusted background music satisfies the energy ratio condition; the accompaniment is obtained based on the adjusted dry chorus sound and the background music; or The method comprises performing sound effect synthesis processing on the dry chorus sound and the background music of the target song according to the sound effect optimization rule to obtain the accompaniment of the target song, including: performing spectrum equalization processing on the dry chorus sound at a preset frequency band; obtaining the loudness of the background music; if the loudness of the background music is less than a loudness threshold, increasing the loudness of the background music to the loudness threshold; the accompaniment is obtained based on the dry chorus sound after spectrum equalization processing and the background music after loudness processing.

2. The method according to claim 1, wherein The obtaining of the dry sound signal set includes: Acquire an initial dry voice signal set from an audio database, wherein the audio database includes initial dry voice signals recorded by multiple users when singing a target song; According to the sound parameters of the respective initial dry sound signals, x dry sound signals are selected from the initial dry sound signal set to form the dry sound signal set.

3. The method according to claim 2, wherein The sound parameters include: pitch characteristic parameters and sound quality characteristic parameters; The pitch characteristic parameters include any one or more of pitch parameters, rhythm parameters and prosody parameters; and the sound quality characteristic parameters include any one or more of noise parameters, energy parameters and speed parameters.

4. The method according to claim 1, wherein: N virtual three-dimensional spatial audio and video positions include; After dividing the horizontal plane at intervals of a first preset angle, n1 virtual three-dimensional spatial sound image positions on the horizontal plane are obtained; After dividing the upper plane at intervals of a second preset angle on the upper plane, n2 virtual three-dimensional spatial sound image positions on the upper plane are obtained; the angle between the upper plane and the horizontal plane is a first angle threshold; After dividing the lower plane at intervals of a third preset angle, n3 virtual three-dimensional spatial sound image positions on the lower plane are obtained; the angle between the lower plane and the horizontal plane is a second angle threshold; Wherein, n1, n2 and n3 are positive integers and their sum is equal to N.

5. The method according to any one of claims 1 to 4, characterized in that The generating of the virtual sound signal based on the dry sound signal corresponding to each virtual three-dimensional sound image position in the N virtual three-dimensional sound image positions includes: Obtaining a head-related transfer function corresponding to each virtual three-dimensional spatial sound image position in N virtual three-dimensional spatial sound image positions; processing a dry sound signal corresponding to the target virtual three-dimensional spatial sound image position using a head-related transfer function corresponding to the target virtual three-dimensional spatial sound image position to obtain a virtual sound signal at the target virtual three-dimensional spatial sound image position; The virtual sound signal at the target virtual three-dimensional space sound image position is a two-channel signal; The target virtual three-dimensional spatial sound image position is any virtual three-dimensional spatial sound image position among the N virtual three-dimensional spatial sound image positions.

6. The method according to any one of claims 1 to 4, characterized in that The virtual sound signal set further includes: a delayed left channel signal and a delayed right channel signal corresponding to each of the p dry sound signals; Before merging the virtual sound signals in the virtual sound signal set to obtain the chorus dry sound, the method further includes: Obtaining p dry sound signals from the x dry sound signals included in the dry sound signal set, where p is a positive integer and is less than or equal to x; Performing a delay processing operation of m1 time parameters on each of the p dry sound signals to obtain m1 delayed dry sound signals corresponding to each of the p dry sound signals, and obtaining a delayed left channel signal corresponding to each of the p dry sound signals by superimposing the m1 delayed dry sound signals corresponding to each of the dry sound signals, where m1 is a positive integer; Each of the p dry sound signals is subjected to a delay processing operation of m2 time parameters to obtain m2 delayed dry sound signals corresponding to each of the p dry sound signals. The delayed right channel signal corresponding to each of the p dry sound signals is obtained by superimposing the m2 delayed dry sound signals corresponding to each dry sound signal, where m2 is a positive integer.

7. A method for playing an accompaniment, characterized in that: include: displaying a user interface, wherein the user interface is used to receive a selection instruction for a target song; If the selection instruction received on the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode, obtaining the accompaniment corresponding to the target song; Play the accompaniment corresponding to the target song; The accompaniment is generated based on the chorus dry sound and background music, the chorus dry sound is generated based on multiple dry sound signals in a dry sound signal set, the multiple dry sound signals in the dry sound signal set correspond to multiple different virtual three-dimensional sound image positions, and the dry sound signal set is obtained based on dry sound signals recorded by multiple users for the target song; The chorus dry sound is generated based on a virtual sound signal set, and the virtual sound signal set includes: a virtual sound signal at each virtual three-dimensional space sound image position in N virtual three-dimensional space sound image positions generated based on the acquired dry sound signal set; wherein, multiple dry sound signals in the dry sound signal set correspond to N virtual three-dimensional space sound image positions, N is an integer greater than 1, the N virtual three-dimensional space sound image positions are different, and each virtual three-dimensional space sound image position is allowed to correspond to one or more dry sound signals.

8. The method according to claim 7, wherein The user interface displays a selection control for the accompaniment mode of the target song, and the accompaniment mode selection control includes: a chorus accompaniment mode selection control and an acoustic accompaniment mode selection control; before obtaining the accompaniment corresponding to the target song, the method further includes: detecting whether a selection operation for the chorus accompaniment mode selection control is obtained; If so, it is confirmed that the selection instruction received in the user interface indicates that the accompaniment mode for the target song is a chorus accompaniment mode.

9. The method according to claim 7, wherein The step of obtaining the accompaniment corresponding to the target song includes: Sending an accompaniment request to a server, wherein the accompaniment request includes identification information of the target song; receiving the dry chorus and the background music returned by the server in response to the accompaniment request; determining a target chorus dry voice segment from the chorus dry voice; The accompaniment corresponding to the target song is obtained according to the chorus dry voice corresponding to the target chorus dry voice segment and the background music.

10. The method according to claim 9, wherein Before determining the target chorus dry voice segment from the chorus dry voice, the method further includes: Displaying a first single sentence interface, displaying each single sentence in the text data corresponding to the chorus dry voice in the order of the time playback nodes of the chorus dry voice; The target chorus dry voice segment is determined based on a single sentence selection operation on the first single sentence interface.

11. The method according to claim 7 or 9, characterized in that After playing the accompaniment corresponding to the target song, the method includes: Displaying a second single sentence interface, displaying each single sentence in the text data corresponding to the accompaniment in the order of the time playback nodes of the accompaniment; Detecting whether a mute selection operation for the dry chorus voice in the accompaniment is obtained during playback; If so, the playing of the dry chorus sound is canceled at the current time playing node.

12. An electronic device, characterized in that: include: A memory, a processor, and a network interface, wherein the processor is connected to the memory and the network interface, wherein the network interface is used to provide a network communication function, the memory is used to store program code, and the processor is used to call the program code to execute the method described in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Method, device and system for performing audio recording, equipment and storage medium

    CN110491358A