Audio processing methods, computer equipment and computer storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-13
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]但是,由于每个用户的网络和设备不同,且用户之间的唱歌水平存在差异,全部合流会导致合唱歌声数据的听感存在较大的不确定性,例如有人在噪音很大的环境下加入合唱,必然会对合唱的整体效果带来很大的干扰,导致合唱歌声的听感不佳
[0016]根据与演唱端的网络传输状况信息确定演唱端的候选干声音频的网络传输性能标签,使用目标音频听感评价模型获得候选干声音频的听感标签,使用目标音频音质评价模型获得候选干声音频的音质标签,从多路候选干声音频中确定网络传输性能标签、听感标签以及音质标签满足预设要求的多路目标干声音频,并将多路目标干声音频进行混合,得到合唱音频。基于各项评价指标对多路干声音频的质量进行衡量,从而筛选出优质的干声音频,进而由多路优质的干声音频合成的合唱音频的听感效果更佳,提升用户合唱的兴趣和体验。
Smart Images

Figure CN117373480B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing, specifically to an audio processing method, a computer device, and a computer storage medium. Background Technology
[0002] In a multi-person chorus scenario, each user's terminal uploads the member's singing data to the server, and the server combines the singing data of multiple users into chorus singing data.
[0003] However, due to differences in each user's network and device, as well as variations in singing ability, merging all voices can lead to significant uncertainty in the perceived sound of the chorus. For example, if someone joins the chorus in a noisy environment, it will inevitably cause considerable interference with the overall effect of the chorus, resulting in a poor listening experience. Summary of the Invention
[0004] This application provides an audio processing method, computer device, and computer storage medium for filtering high-quality dry audio from multiple dry audio streams to improve the listening experience of choral audio.
[0005] A first aspect of this application provides an audio processing method, the method being applied to a server, the server being connected to a singing end; the method includes:
[0006] Multiple candidate dry audio files are acquired. Each candidate dry audio file is a dry audio file of a user singing a chorus at a singing terminal. The chorus content sung by users at multiple singing terminals is the same.
[0007] When the number of dry audio channels exceeds a preset value, for each singing end, the network transmission performance label of the candidate dry audio of the singing end is determined based on the network transmission status information of the singing end;
[0008] Each candidate dry audio stream is input into a pre-trained target audio listening evaluation model to obtain listening labels output by the target audio listening evaluation model to describe the listening experience of the candidate dry audio stream.
[0009] Each candidate dry audio stream is input into a pre-trained target audio quality evaluation model to obtain a sound quality label output by the target audio quality evaluation model to describe the sound quality of the candidate dry audio stream.
[0010] From the multiple candidate dry audio files, determine the multiple target dry audio files whose network transmission performance label, auditory perception label, and sound quality label meet preset requirements;
[0011] The multi-channel target dry audio is mixed to obtain the chorus audio.
[0012] A second aspect of this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method of the first aspect described above.
[0013] A third aspect of this application provides a computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.
[0014] A fourth aspect of this application provides a computer program product, the computer program product including a computer program that, when executed by a processor, implements the method of the first aspect described above.
[0015] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0016] Based on network transmission status information from the singing end, network transmission performance labels are determined for candidate dry audio files. A target audio listening evaluation model is used to obtain listening quality labels for these candidate dry audio files, and a target audio sound quality evaluation model is used to obtain sound quality labels. From multiple candidate dry audio files, multiple target dry audio files that meet preset requirements for network transmission performance labels, listening quality labels, and sound quality labels are selected. These multiple target dry audio files are then mixed to obtain the choral audio. The quality of the multiple dry audio files is measured based on various evaluation indicators to select high-quality dry audio files. The resulting choral audio synthesized from multiple high-quality dry audio files has a better listening experience, enhancing the user's interest and experience in choral singing. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the network framework in an embodiment of this application;
[0018] Figure 2 This is a flowchart illustrating an audio processing method in an embodiment of this application;
[0019] Figure 3 This is another flowchart illustrating the audio processing method in this application embodiment;
[0020] Figure 4 This is a flowchart illustrating the steps of aligning multiple target dry audio signals in an embodiment of this application.
[0021] Figure 5 This is a schematic diagram illustrating the interaction between the server, the performing end, and the audience end in an embodiment of this application;
[0022] Figure 6This is a schematic diagram of a computer device in an embodiment of this application. Detailed Implementation
[0023] This application provides an audio processing method, computer device, and computer storage medium for filtering high-quality dry audio from multiple dry audio streams to improve the listening experience of choral audio.
[0024] Please see Figure 1 The network framework in this embodiment includes:
[0025] The system consists of a server and multiple vocal terminals. The server can communicate with each vocal terminal via a network. The server can have a corresponding data storage system that can store the data that the server needs to process, such as dry audio collected from the vocal terminals. The data storage system can be integrated into the server or deployed in the cloud or other network servers.
[0026] In this embodiment, a user can activate a singing client and participate in a chorus interaction in a virtual room. During the chorus, the singing client can collect the vocal signals generated when the user sings a song, obtaining the corresponding dry audio, which can then be uploaded to the server. Thus, the server can obtain dry audio from multiple singing clients in the virtual room, each collecting the same chorus content. Furthermore, the server can further determine the audio quality of each dry audio stream, and then evaluate the quality of each dry audio stream based on its audio quality, obtaining audio quality information for each dry audio stream. Finally, the server can identify multiple target dry audio streams from the multiple dry audio streams whose audio quality information meets preset conditions.
[0027] After obtaining multiple target dry audio streams, the server can mix them to produce a chorus audio stream for the audience in the virtual room. The singing and listening devices in the virtual room can include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, and smart in-vehicle systems, while portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. The server can be implemented using a dedicated server or a server cluster consisting of multiple servers.
[0028] The following is combined with Figure 1 The network framework described herein is used to illustrate the audio processing method in the embodiments of this application.
[0029] Please see Figure 2 One embodiment of the audio processing method in this application includes:
[0030] 201. Obtain multiple candidate dry audio files, where each candidate dry audio file is a dry audio file of a user singing a chorus at the singing terminal.
[0031] Users can use a vocal app to sing a choral piece. The app then collects the dry audio of the user's performance and uploads it to the server. Multiple users on different apps may be singing the same choral piece, such as a song or a traditional opera. The server receives the dry audio uploaded by each app in real time, generating multiple candidate dry audio streams. The server then selects the best dry audio stream to synthesize the choral audio.
[0032] 202. When the number of dry audio channels exceeds a preset value, for each singing end, determine the network transmission performance label of the candidate dry audio of the singing end based on the network transmission status information of the singing end;
[0033] The server can evaluate the quality of each dry audio stream, one evaluation metric being the network transmission performance between the singer and the server. Since choral projects emphasize the synchronization and harmony of multiple sound sources, and network transmission performance affects the sending and receiving of audio data packets, the harmony and synchronization among multiple candidate dry audio streams significantly impacts the overall choral quality. Therefore, for each singer, the server can determine a network transmission performance tag for the candidate dry audio streams based on the network transmission status information. This tag describes the network transmission performance of that candidate dry audio stream.
[0034] 203. Input each of the candidate dry audio streams into the pre-trained target audio listening evaluation model to obtain the listening labels output by the target audio listening evaluation model to describe the listening experience of the candidate dry audio streams.
[0035] The server can obtain a pre-trained target audio listening evaluation model. The target audio listening evaluation model is trained on multiple dry audio files based on machine learning algorithms. During the training process, it learns to determine the listening labels of the dry audio files based on their listening characteristics. Therefore, this target audio listening evaluation model can be used to determine the listening labels of each candidate dry audio file. That is, each candidate dry audio file is input into the target audio listening evaluation model to obtain the listening labels output by the target audio listening evaluation model to describe the listening experience of the candidate dry audio files.
[0036] The target audio listening evaluation model can be pre-trained on the server or trained on other devices and then deployed to the server.
[0037] 204. Input each of the candidate dry audio streams into the pre-trained target audio quality evaluation model to obtain the audio quality label output by the target audio quality evaluation model to describe the audio quality of the candidate dry audio stream.
[0038] The server can also obtain a pre-trained target audio quality evaluation model. The target audio quality evaluation model is trained on multiple dry audio files based on machine learning algorithms. During the training process, it learns to determine the audio quality labels of the dry audio files based on their audio quality characteristics. Therefore, this target audio quality evaluation model can be used to determine the audio quality labels of each candidate dry audio file. That is, each candidate dry audio file is input into the target audio quality evaluation model to obtain the audio quality labels output by the target audio quality evaluation model to describe the audio quality of the candidate dry audio files.
[0039] The target audio quality evaluation model can be pre-trained on the server or trained on other devices and then deployed to the server.
[0040] 205. Determine multiple target dry audio streams from the multiple candidate dry audio streams that meet preset requirements for network transmission performance label, auditory perception label, and sound quality label;
[0041] 206. Mix the multiple target dry audio streams to obtain the chorus audio;
[0042] After obtaining the evaluation indicators of each candidate dry audio, multiple target dry audios that meet the preset requirements for network transmission performance label, listening experience label, and sound quality label can be selected from the multiple candidate dry audios. That is, multiple target dry audios that meet the conditions for network transmission performance, listening experience, and sound quality are selected, and the multiple target dry audios are mixed to obtain the chorus audio.
[0043] In this embodiment, the network transmission performance labels of candidate dry audio files for the singing end are determined based on the network transmission status information of the singing end. A target audio listening evaluation model is used to obtain the listening perception labels of the candidate dry audio files, and a target audio sound quality evaluation model is used to obtain the sound quality labels of the candidate dry audio files. From multiple candidate dry audio files, multiple target dry audio files that meet preset requirements for network transmission performance labels, listening perception labels, and sound quality labels are selected. These multiple target dry audio files are then mixed to obtain the choral audio. The quality of the multiple dry audio files is measured based on various evaluation indicators, thereby selecting high-quality dry audio files. The choral audio synthesized from multiple high-quality dry audio files has a better listening effect, enhancing the user's interest and experience in choral singing.
[0044] The following will be discussed in the preceding text. Figure 2 Based on the illustrated embodiments, embodiments of this application will be described in further detail. Please refer to [link to relevant documentation]. Figure 3Another embodiment of the audio processing method in this application includes:
[0045] 301. Obtain multiple candidate dry audio files, each candidate dry audio file being a dry audio file of a user singing chorus content collected by the singing terminal, and the chorus content sung by multiple users of the singing terminals being the same.
[0046] In the specific implementation, users can launch a client and enter a virtual room, which can also be called a virtual karaoke room. Multiple clients can enter the same virtual room for real-time online interaction. In this embodiment, clients participating in the chorus interaction in the virtual room can be called singing clients, and those not participating in the chorus interaction can be called audience clients. For example, after receiving a chorus interaction initiation request from a client with the identity of a broadcaster (i.e., a broadcaster) or a client with the identity of an administrator (such as a room administrator) in the virtual room, the server can send chorus invitation information to each client in the virtual room. If a user confirms participation in the chorus interaction, it can trigger the client to send the corresponding confirmation information. The server can then identify the client that sent the confirmation information as a singing client. Clients that do not respond to the chorus invitation information (such as clients replying with a rejection command or failing to reply within a timeout) can be identified by the server as audience clients.
[0047] During the chorus, all participating users can sing the same content, meaning multiple users can sing in unison. While users are singing, the corresponding client can capture the current vocal signal to obtain the corresponding dry audio and upload it to the server. Therefore, during the chorus, the server can obtain the dry audio captured by each of the multiple singing devices in the virtual room, and the dry audio captured by each singing device can be for the same choral content.
[0048] 302. When the number of dry audio channels exceeds a preset value, for each singing end, determine the network transmission performance label of the candidate dry audio of the singing end based on the network transmission status information of the singing end;
[0049] In this embodiment, the network transmission status information can be information representing network transmission status such as packet loss rate and latency. When determining the packet loss rate, the server can determine the packet loss rate based on the number of all audio data packets sent by the singing end and the number of lost packets. When determining the latency, the latency of the dry audio can be determined based on the progress of mixing multiple dry audio streams and the progress of dry audio transmission.
[0050] For example, if the server mixes multiple dry audio frames at a rate of 20ms per frame, and the current frame of a certain dry audio stream is sent to the server 100ms after the server has mixed all the dry audio frames, then the delay time for this dry audio stream is 100ms. The server can preset an acceptable range for the delay time, such as 500ms. Dry audio streams with a delay time less than 500ms are acceptable, while those with a delay time greater than 500ms are not acceptable and can be discarded.
[0051] In this embodiment, when the number of candidate dry audio channels exceeds a preset value, that is, when the number of participants in the chorus is more than N (N is configurable), the operation process of selecting high-quality dry audio from multiple candidate dry audio channels can be executed in steps 302 to 305 of this embodiment; when the number of participants in the chorus is less than N (N is configurable), it is not necessary to select high-quality dry audio from multiple candidate dry audio channels.
[0052] For example, Figure 1 In the scenario shown, when there are more than two singing terminals and more than two candidate dry audio tracks are collected in total, such as 10 singing terminals collecting a total of 10 candidate dry audio tracks, the high-quality dry audio tracks can be selected from these 10 candidate dry audio tracks using the method of this embodiment. Assuming 6 relatively high-quality candidate dry audio tracks are selected, these 6 relatively high-quality candidate dry audio tracks can be mixed to obtain the chorus audio. The chorus audio then combines the excellent listening experience, sound quality, and smoothness of each high-quality dry audio track, resulting in a better listening experience. However, when there are only 2 candidate dry audio tracks, i.e., only two people are singing together, these 2 candidate dry audio tracks can be directly mixed to obtain the chorus audio without needing to select high-quality dry audio tracks. Alternatively, more users can join the chorus to obtain more candidate dry audio tracks before selecting high-quality dry audio tracks to mix and obtain the chorus audio.
[0053] 303. Input each of the candidate dry audio streams into the pre-trained target audio listening evaluation model to obtain the listening labels output by the target audio listening evaluation model to describe the listening experience of the candidate dry audio streams.
[0054] In this embodiment, the listening perception label includes evaluation information for describing the listening perception of the candidate dry audio. For example, the evaluation information can be information determined in a quantitative way, such as a score; or information determined in a non-quantitative way, such as classifying the listening perception of the candidate dry audio into levels, with each level being the evaluation information.
[0055] The training steps for the target audio listening experience evaluation model include:
[0056] Obtain an initial audio listening evaluation model and multiple sets of first training samples. Each set of first training samples includes dry audio and a first pre-labeled label used to describe the listening experience of the dry audio.
[0057] Multiple sets of first training samples are input into the initial audio listening evaluation model so that the initial audio listening evaluation model extracts the listening features of the dry audio in the first training samples and determines the first predicted label to describe the listening experience of the dry audio based on the listening features, and outputs the first predicted label. Training stops when the relationship between the first predicted label and the first pre-labeled label meets the convergence condition, and the target audio listening evaluation model is obtained.
[0058] The initial audio listening evaluation model can be a neural network model related to audio processing, such as the SANet lightweight parameterless network model. The listening features of dry audio can include features of multiple dimensions such as pitch features, rhythm features, emotional features, breath features, and skill features. The initial audio listening evaluation model extracts multiple listening features of dry audio and determines the first predicted label of dry audio based on these multiple listening features.
[0059] The first pre-label of the dry audio in the training samples can be the rating of the listener's perception of the dry audio. The initial audio listening evaluation model will also output the listening evaluation of the dry audio based on the various listening features of the dry audio. When the rating pre-labeled by the person and the rating output by the model meet the convergence condition, the model training will stop and the target audio listening evaluation model will be obtained.
[0060] The relationship between the first predicted label and the first pre-labeled label satisfies the convergence condition. This can be achieved by the loss function constructed based on the first predicted label and the first pre-labeled label becoming stable, or by the loss function being less than a preset value, or by the number of iterations meeting a preset number. This embodiment does not limit the conditions under which the convergence condition is met.
[0061] 304. Input each of the candidate dry audio streams into the pre-trained target audio quality evaluation model to obtain the audio quality label output by the target audio quality evaluation model to describe the audio quality of the candidate dry audio stream.
[0062] In this embodiment, the audio quality label includes a label that represents the audio quality classification result of the dry audio, and the audio quality classification result includes the dry audio with noise and the dry audio without noise; or, the audio quality classification result includes the type of noise mixed with the dry audio and the dry audio without noise.
[0063] Types of noise mixed in dry audio include current noise, reverberation, popping sounds, click sounds, background noise, back noise, and accompaniment. Click sounds can be sounds produced by actions such as knocking or closing a door.
[0064] The training steps for the target audio quality evaluation model include:
[0065] Obtain an initial audio quality evaluation model and multiple sets of second training samples. Each set of second training samples includes dry audio and a second pre-labeled label used to describe the quality of the dry audio.
[0066] Multiple sets of second training samples are input into the initial audio quality evaluation model to obtain the second predicted label output by the initial audio quality evaluation model to describe the sound quality of dry audio. Training stops when the relationship between the second predicted label and the second pre-labeled label meets the convergence condition, so as to obtain the target audio quality evaluation model.
[0067] For example, the sound quality classification result of dry audio indicates the type of noise mixed in the dry audio and the absence of noise in the dry audio. The second pre-label for the dry audio in the training samples can be a classification result representing the sound quality of the dry audio, such as the dry audio containing electrical noise, or the dry audio containing reverberation, or the dry audio containing popping sounds, etc. Furthermore, the initial audio sound quality evaluation model can extract the sound quality features of the dry audio in the training samples and identify these sound quality features to determine the sound quality classification result of the dry audio. For example, identifying sound quality features can determine whether the dry audio contains noise and, when determining that it contains noise, determine the type of noise based on the sound quality features, and output a second predicted label for the dry audio.
[0068] The relationship between the second predicted label and the second pre-labeled label satisfies the convergence condition, which can be achieved by the loss function constructed based on the second predicted label and the second pre-labeled label becoming stable, or by the loss function being less than a preset value, or by the number of iterations meeting a preset number. This embodiment does not limit the conditions under which the convergence condition is met.
[0069] The initial audio quality evaluation model can be a deep residual network, ResNet. ResNet is a deep convolutional neural network constructed from residual blocks, whose outputs are obtained through ReLU activation layers. The training of the initial audio quality evaluation model can be divided into two stages. The first stage uses the officially released dry audio dataset for training, and the second stage uses manually pre-labeled dry audio samples for training, in order to further adjust and optimize the model's parameters.
[0070] 305. Determine multiple target dry audio streams from the multiple candidate dry audio streams that meet preset requirements for network transmission performance label, auditory perception label, and sound quality label;
[0071] In this embodiment, multiple target dry audio files can be determined from multiple candidate dry audio files, based on the network transmission performance label meeting the first preset condition, the listening perception label meeting the second preset condition, and the sound quality label meeting the third preset requirement.
[0072] For example, one way to determine the target dry audio is if the candidate dry audio meets the corresponding preset conditions for each of the three indicators mentioned above; otherwise, the candidate dry audio will be discarded and not determined as the target dry audio. For example, if the delay time of a candidate dry audio is less than a preset duration and there is no noise, but its listening score is less than a preset threshold, then this candidate dry audio will be discarded and will not be used to produce choral audio.
[0073] In a preferred embodiment of this example, another way to determine the target dry audio is to preset a priority order among the network transmission performance label, listening perception label, and sound quality label of the candidate dry audio, and the server sorts the multiple candidate dry audio according to this priority order. For example, if the priority order is network transmission performance label > listening perception label > sound quality label, and there are 5 candidate dry audios A, B, C, D, and E, they are first sorted according to their network transmission performance labels. If there are multiple candidate dry audios with the same packet loss rate or latency, such as B, C, and D, they are sorted according to their listening perception labels, such as by listening perception score. If the listening perception scores of these 3 candidate dry audios are still the same, they are sorted according to their sound quality labels, with the one with no noise or the lowest noise level placed first, and the final sorting result is determined.
[0074] After obtaining the sorting results, the candidate dry audio files with the first preset number of positions are selected as the target dry audio files from the sorted candidate dry audio files. For example, in the above example, assuming the sorting results are A, B, C, D, and E from best to worst, and the number of dry audio files used to produce the chorus audio is 4, then the 4 dry audio files A, B, C, and D are selected to produce the chorus audio.
[0075] In another preferred embodiment, the sorting method for the multiple candidate dry audio files can be as follows: the network transmission performance label, listening experience label, and sound quality label of the multiple candidate dry audio files include a highest priority label, a second highest priority label, and a lowest priority label. The multiple candidate dry audio files can be sorted from best to worst or worst to best according to the highest priority label to obtain a first sorting result; the multiple candidate dry audio files can be sorted from best to worst or worst to best according to the second highest priority label to obtain a second sorting result; and the multiple candidate dry audio files can be sorted from best to worst or worst to best according to the lowest priority label to obtain a third sorting result. Based on the first, second, and third sorting results, a target sorting result for the multiple candidate dry audio files is determined. Furthermore, the candidate dry audio files with the first preset number of positions in the target sorting result can be identified as the target dry audio files.
[0076] In determining the target ranking of multiple candidate dry audio files based on the first, second, and third ranking results, the first ranking result is assigned the highest weight for the same position among the three results, followed by the second, and then the third. For example, if the priority order is network transmission performance label > listening experience label > sound quality label, and there are five candidate dry audio files A, B, C, D, and E, the files are first ranked from best to worst according to the network transmission performance label, then from best to worst according to the listening experience label, and finally from best to worst according to the sound quality label, resulting in three ranking results: the first, second, and third. Different weights are assigned to the same position among these three ranking results; the first ranking result has the highest weight, followed by the second, and then the third. For example, the weight of the first position in the first sorting result is 15, the weight of the first position in the second sorting result is 14, and the weight of the first position in the third sorting result is 13; the weight of the second position in the first sorting result is 12, the weight of the second position in the second sorting result is 11, the weight of the second position in the third sorting result is 10, and so on. Then, for each candidate dry audio signal, the weights of its corresponding positions in the three sorting results are accumulated to obtain the total weight of each candidate dry audio signal. The multiple candidate dry audio signals are then sorted according to the total weight to obtain the target sorting result.
[0077] Therefore, by selecting high-quality candidate dry audio based on various evaluation indicators of multiple candidate dry audio sources, it is possible to synthesize choral videos from these high-quality candidate dry audio sources, thereby improving the listening experience of the choral videos.
[0078] In addition to the aforementioned evaluation indicators such as network transmission performance, listening experience, and sound quality, the server can also determine the target dry audio based on other evaluation indicators. For example, the server can also determine the vocal position and timbre corresponding to the candidate dry audio based on its audio characteristics. The vocal position corresponding to the candidate dry audio will affect the stereo effect of the choral audio; while the vocal timbre will also affect the listening experience of the choral audio. For example, an extremely sharp timbre in a female chorus will affect the listening experience of the choral audio; the timbre of the dry audio needs to be compatible with the timbre of the instruments in the choral content. For example, if the instruments in the choral content have a cheerful style, then a bright and crisp dry audio is needed, and a deep dry audio is not suitable.
[0079] Therefore, when determining the target dry audio, another preferred approach is for the server to select multiple target dry audios from multiple candidate dry audios that meet preset requirements in terms of network transmission performance tags, listening experience tags, sound quality tags, and vocal position and timbre. This ensures that the target dry audios meet the preset requirements in multiple aspects such as network transmission, listening experience, sound quality, and vocal position and timbre, thereby further optimizing the listening experience of the synthesized choral audio.
[0080] 306. Mix the multiple target dry audio streams to obtain the chorus audio;
[0081] When mixing multiple target dry audio streams, the multiple target dry audio streams can be superimposed or weighted superimposed to obtain choral audio.
[0082] After mixing multiple target dry audio streams to obtain the choral audio, the server can send these multiple target dry audio streams to other singing terminals besides the one that collected the target dry audio, or send other target dry audio streams besides its own to the singing terminal that collected the target dry audio. This allows choral singers to experience the atmosphere of the chorus and the effect of others singing together. Furthermore, the choral audio can be sent to the audience so they can appreciate the choral work and provide feedback, enhancing the fun and interactivity of the choral performance.
[0083] For example, during the mixing process, such as Figure 4 As shown, taking singers A, B, C, D, and E as an example, the server can determine the corresponding dry audio A, dry audio B, dry audio C, dry audio D, and dry audio E as the target dry audio for synthesizing the chorus audio. Each target dry audio can carry accompaniment progress information and flow into different queues. Then, the queues are aligned according to the accompaniment progress information. The alignment standard can be that the offset between different queues is less than a certain number of milliseconds (such as 300 milliseconds). Then, 3-5 dry audio channels are designated as the main dry audio channels and the others as auxiliary dry audio channels, and then the mixing is performed.
[0084] After mixing, the choral audio with added accompaniment can be sent to various audience terminals in the virtual room. For example, the choral audio can be sent to the interface machine of the real-time audio and video communication server, and the audience terminals in the virtual room can retrieve the choral audio through the interface machine.
[0085] For example, a network transmission architecture can be as follows: Figure 5 As shown, the server can include a Real-Time Communication (RTC) server and other backend processing equipment servers. The RTC server can receive the dry audio uploaded by each chorus user's singing end and transmit it to the backend processing equipment server (such as a chorus mixing server) for mixing. The mixed chorus audio can be sent to the audience end to realize online real-time chorus with multiple singing ends in different locations, ensuring data transmission quality and improving the interactivity and fun of chorus.
[0086] The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 6 One embodiment of the computer device in this application includes:
[0087] The computer device 600 may include one or more central processing units (CPUs) 601 and a memory 605, in which one or more applications or data are stored.
[0088] The memory 605 can be volatile or persistent storage. The program stored in the memory 605 can include one or more modules, each module including a series of instruction operations on the computer device. Furthermore, the central processing unit 601 can be configured to communicate with the memory 605 and execute the series of instruction operations stored in the memory 605 on the computer device 600.
[0089] The computer device 600 may also include one or more power supplies 602, one or more wired or wireless network interfaces 603, one or more input / output interfaces 604, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0090] The central processing unit 601 can perform the aforementioned... Figures 2 to 3 The specific operations performed by the computer device in the illustrated embodiment will not be described in detail here.
[0091] This application also provides a computer storage medium, one embodiment of which includes: the computer storage medium storing instructions, which, when executed on a computer, cause the computer to perform the aforementioned... Figures 2 to 3 The operations performed by the computer device in the illustrated embodiment.
[0092] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0094] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0095] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0096] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. An audio processing method, characterized in that, The method is applied to a server, which is connected to a singing client; the method includes: Multiple candidate dry audio files are acquired. Each candidate dry audio file is a dry audio file of a user singing a chorus at a singing terminal. The chorus content sung by users at multiple singing terminals is the same. When the number of dry audio channels exceeds a preset value, for each singing end, the network transmission performance label of the candidate dry audio of the singing end is determined based on the network transmission status information of the singing end; Each candidate dry audio stream is input into a pre-trained target audio listening evaluation model to obtain listening labels output by the target audio listening evaluation model to describe the listening experience of the candidate dry audio stream. Each candidate dry audio stream is input into a pre-trained target audio quality evaluation model to obtain a sound quality label output by the target audio quality evaluation model to describe the sound quality of the candidate dry audio stream. From the multiple candidate dry audio files, determine the multiple target dry audio files whose network transmission performance label, auditory perception label, and sound quality label meet preset requirements; The multi-channel target dry audio is mixed to obtain the chorus audio.
2. The method according to claim 1, characterized in that, The step of determining the target dry audio streams from the multiple candidate dry audio streams that meet preset requirements for the network transmission performance label, the listening perception label, and the sound quality label includes: The multiple candidate dry audio files are sorted according to a preset priority order among the network transmission performance label, the auditory perception label, and the sound quality label; Among the sorted candidate dry audio streams, the candidate dry audio streams with the first preset number of positions are determined as the target dry audio stream.
3. The method according to claim 2, characterized in that, The network transmission performance label, the auditory perception label, and the sound quality label include the highest priority label, the second highest priority label, and the lowest priority label; The step of sorting the multiple candidate dry audio streams according to a preset priority order among the network transmission performance label, the auditory perception label, and the sound quality label includes: The candidate dry audio streams are sorted from best to worst or from worst to best according to the highest priority label to obtain the first sorting result; The candidate dry audio streams are sorted from best to worst or from worst to best according to the second highest priority label to obtain a second sorting result; The candidate dry audio streams are sorted from best to worst or from worst to best according to the lowest priority label to obtain a third sorting result; The target sorting result of the multi-channel candidate dry audio is determined based on the first sorting result, the second sorting result, and the third sorting result; The step of determining the target dry audio from the sorted multiple candidate dry audio streams by selecting the candidate dry audio streams with a predetermined order of the first preset number of positions includes: The candidate dry audio files with the first preset number of positions in the target sorting results are determined as the target dry audio files.
4. The method according to claim 1, characterized in that, The method further includes: The position and timbre of the human voice corresponding to the candidate dry audio are determined based on the audio characteristics of the candidate dry audio. The step of determining the target dry audio streams from the multiple candidate dry audio streams that meet preset requirements for the network transmission performance label, the listening perception label, and the sound quality label includes: The target dry audio is determined from the multiple candidate dry audio streams, based on the network transmission performance label, the auditory perception label, the sound quality label, the voice position, and the voice timbre, and the results satisfying the preset requirements.
5. The method according to claim 1, characterized in that, The network transmission performance label includes the packet loss rate of audio data packets transmitted from the singing end to the server, the latency of the candidate dry audio; and / or The auditory perception label includes evaluation information describing the auditory perception of the candidate dry audio; and / or The audio quality label includes a label for representing the audio quality classification result of the dry audio, wherein the audio quality classification result includes noise mixed in the dry audio and no noise mixed in the dry audio; or, the audio quality classification result includes the type of noise mixed in the dry audio and no noise mixed in the dry audio. The types of noise mixed in the dry audio include current noise, reverberation, popping sound, click sound, background noise, back noise, and accompaniment.
6. The method according to claim 1, characterized in that, The training steps of the target audio listening evaluation model include: Obtain an initial audio listening evaluation model and multiple sets of first training samples. Each set of first training samples includes dry audio and a first pre-labeled label used to describe the listening experience of the dry audio. Multiple sets of the first training samples are input into the initial audio listening evaluation model so that the initial audio listening evaluation model extracts the listening features of the dry audio in the first training samples and determines the first predicted label for describing the listening experience of the dry audio based on the listening features, and outputs the first predicted label. Training stops when the relationship between the first predicted label and the first pre-labeled label meets the convergence condition, and the target audio listening evaluation model is obtained. The auditory characteristics include pitch characteristics, rhythm characteristics, emotional characteristics, breath characteristics, and technique characteristics of dry audio.
7. The method according to claim 1, characterized in that, The training steps of the target audio quality evaluation model include: Obtain an initial audio quality evaluation model and multiple sets of second training samples. Each set of second training samples includes dry audio and a second pre-labeled label used to describe the quality of the dry audio. Multiple sets of the second training samples are input into the initial audio quality evaluation model to obtain the second predicted label output by the initial audio quality evaluation model to describe the sound quality of the dry audio. Training stops when the relationship between the second predicted label and the second pre-labeled label meets the convergence condition, so as to obtain the target audio quality evaluation model.
8. The method according to any one of claims 1 to 7, characterized in that, The server is also connected to the viewer; the method further includes: Send the multiple target dry audio signals to the singing terminals other than the singing terminal that collects the target dry audio signals; Send the target dry audio from the multiple target dry audio channels other than the target dry audio collected by itself to the singing end that collects the target dry audio; The chorus audio is sent to the audience.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.
10. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed on the computer, cause the computer to perform the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Vocal music performance scoring method and system based on neural network and audio-visual fusion
CN115579021A
Sound quality evaluation model determination method, sound quality evaluation method, equipment and medium
CN115631768A