Mixing method, device and electronic device
By identifying the venue logo and performing mixing processing, the problem of audio mixing accuracy in multi-channel audio interactive VR virtual scenes is solved, and efficient audio stream processing in large venues is realized, reducing costs.
Patent Information
- Application Number
- CN202210503586.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-09
AI Technical Summary
In the VR virtual scene of multi-channel audio interaction, how to accurately mix and push the audio from different venues back to the corresponding venue, as the number of venues and the number of people online increases, the existing technology is difficult to effectively solve.
By receiving the audio streams to be mixed corresponding to multiple venues, identifying the venue identifiers, mixing the audio streams to be mixed with the same venue, generating the target audio streams to be mixed, and sending them to multiple clients in the venue.
Accurate mixing of different venues is realized, ensuring that users of each venue can hear the sound in the entire venue, meeting the requirements of large-scale voice interactions in large venues, and reducing the mixing cost.
Smart Images

Figure CN114915750B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of artificial intelligence and cloud computing, and particularly to a mixing method, apparatus, and electronic device. Background Art
[0002] In a VR virtual scenario based on multi-channel audio interaction, intelligent mixing can mix multiple audio streams in a virtual venue and send them to each user in the venue, enabling users to vividly hear the voices of other users in the venue. With the increase in the number of venues and the number of concurrent online users, how to accurately mix the audio of different venues and then push it back to the corresponding venues has become an urgent problem to be solved. Summary of the Invention
[0003] A mixing method, apparatus, and electronic device are provided.
[0004] According to a first aspect, a mixing method is provided, including: receiving multiple to-be-mixed audio streams corresponding to multiple venues, where the to-be-mixed audio streams include the venue identifiers of the venues; performing mixing processing on the to-be-mixed audio streams including the same venue identifier to generate a target mixed audio stream corresponding to the venue; and sending the target mixed audio stream to multiple clients in the venue.
[0005] According to a second aspect, a mixing method apparatus is provided, including: a receiving module, configured to receive multiple to-be-mixed audio streams corresponding to multiple venues, where the to-be-mixed audio streams include the venue identifiers of the venues; a mixing module, configured to perform mixing processing on the to-be-mixed audio streams including the same venue identifier to generate a target mixed audio stream corresponding to the venue; and a sending module, configured to send the target mixed audio stream to multiple clients in the venue.
[0006] According to a third aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the mixing method according to the first aspect of the present disclosure.
[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the mixing method according to the first aspect of the present disclosure.
[0008] According to a fifth aspect, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the mixing method according to the first aspect of the present disclosure are implemented.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. Brief Description of the Drawings
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a schematic flowchart of a mixing method according to the first embodiment of the present disclosure;
[0012] Figure 2 is a schematic diagram of a scenario of a mixing method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic flowchart of a mixing method according to the second embodiment of the present disclosure;
[0014] Figure 4 is a schematic flowchart of a mixing method according to the third embodiment of the present disclosure;
[0015] Figure 5 is a schematic flowchart of a mixing method according to the fourth embodiment of the present disclosure;
[0016] Figure 6 is a schematic diagram of a scenario of a two-layer mixing service of a mixing method according to an embodiment of the present disclosure;
[0017] Figure 7 is a schematic flowchart of a mixing method according to the fifth embodiment of the present disclosure;
[0018] Figure 8 is a schematic flowchart of the mixing process of a mixing server included in the first-layer mixing service in the mixing method according to an embodiment of the present disclosure;
[0019] Figure 9 is a schematic flowchart of the mixing process of a mixing server included in other layer mixing services except the first-layer mixing service in the mixing method according to an embodiment of the present disclosure;
[0020] Figure 10 is a schematic diagram of the arrangement of audio streams to be mixed in the mixing method according to an embodiment of the present disclosure;
[0021] Figure 11 is a schematic diagram of generating a mixed audio stream after mixing the arranged audio streams to be mixed in the mixing method according to an embodiment of the present disclosure;
[0022] Figure 12 is a schematic flowchart of a mixing method according to the sixth embodiment of the present disclosure;
[0023] Figure 13 It is a schematic diagram of the mixing process of the mixing server included in the last - layer mixing service in the mixing method according to an embodiment of the present disclosure;
[0024] Figure 14 It is a block diagram of the mixing method device according to the first embodiment of the present disclosure;
[0025] Figure 15 It is a block diagram of the mixing method device according to the second embodiment of the present disclosure;
[0026] Figure 16 It is a block diagram of an electronic device for implementing the method according to an embodiment of the present disclosure. Detailed implementation manners
[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well - known functions and structures are omitted in the following description.
[0028] Artificial Intelligence (AI) is a technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Currently, AI technology has the advantages of high automation, high precision, and low cost, and has been widely applied.
[0029] Cloud Computing is a type of distributed computing. It refers to decomposing huge data computing programs into countless small programs through the network "cloud", and then processing and analyzing these small programs through a system composed of multiple servers to obtain results and return them to users. In the early stage of cloud computing, simply speaking, it was simple distributed computing, which solved task distribution and merged calculation results. Through this technology, it is possible to complete the processing of tens of thousands of data within a very short time (a few seconds), thus achieving powerful network services.
[0030] The following describes the mixing method, device, and electronic device according to the embodiments of the present disclosure with reference to the accompanying drawings.
[0031] Figure 1 It is a schematic diagram of the process of the mixing method according to the first embodiment of the present disclosure.
[0032] As Figure 1 shown, the mixing method according to the embodiments of the present disclosure may specifically include the following steps:
[0033] S101. Receive multiple to-be-mixed audio streams corresponding to multiple venues, where the to-be-mixed audio streams include venue identifiers of the venues.
[0034] Specifically, the execution subject of the audio mixing method in this disclosure embodiment can be the audio mixing device provided by this disclosure embodiment. The audio mixing device can be a hardware device with data information processing capabilities and / or necessary software for driving the hardware device to work. Optionally, the execution subject can include workstations, servers, computers, user terminals, and other devices. Among them, user terminals include but are not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle terminals, etc.
[0035] This disclosure embodiment can be applied to mixing audio for multiple large venues simultaneously. For example, virtual venues based on Virtual Reality (VR) technology, where there are 100,000 users online simultaneously in the virtual venue. Users can hold meetings or launch new products in the virtual venue, ensuring that every user in the venue can hear the sounds in the venue as if they were on the spot.
[0036] Each venue corresponds to audio streams sent by multiple clients. These audio streams are used as to-be-mixed audio streams for mixing. Each to-be-mixed audio stream includes the venue identifier of the venue corresponding to the audio stream. The venue identifier is used to uniquely identify the venue, such as the Identity Document (ID) corresponding to the venue.
[0037] S102. Perform mixing processing on the to-be-mixed audio streams including the same venue identifier to generate a target mixed audio stream corresponding to the venue.
[0038] In this disclosure embodiment, each to-be-mixed audio stream in all received to-be-mixed audio streams is parsed to extract the venue identifier in the to-be-mixed audio stream. The to-be-mixed audio streams including the same venue identifier are subjected to mixing processing to generate a target mixed audio stream corresponding to the venue. Thus, non-interfering mixing can be performed on multiple venues to generate target mixed audio streams corresponding to each venue.
[0039] S103. Send the target mixed audio stream to multiple clients in the venue.
[0040] In this application embodiment, the target mixed audio stream corresponding to the venue is pushed back to the venue, that is, pushed back to the clients corresponding to each user in the venue, so that each user in the venue can hear the sounds in the entire venue.
[0041] For example, as Figure 2 shown, when two venues, Venue A and Venue B, are opened simultaneously, each client in Venue A and Venue B (such as Figure 2Multiple terminal graphics in the medium venue A represent that the client sends the audio stream to be mixed corresponding to the user based on the Real-time Transport Protocol (RTP). The audio streams to be mixed in venue A and venue B are received simultaneously (e.g., Figure 2 n audio streams to be mixed corresponding to venue A and n audio streams to be mixed corresponding to venue B in the medium venue A). In the embodiments of the present disclosure, by identifying the venue identifier in each audio stream to be mixed, the audio streams to be mixed including the same venue identifier are mixed, so as to distinguish the audio processing processes of different venues, and at the same time, mix the audio of multiple venues. After mixing, the target mixed audio stream corresponding to the venue is pushed back to multiple clients in the venue (e.g., Figure 2 The 1 target mixed audio stream pushed back to venue A shown).
[0042] In summary, the mixing method of the embodiments of the present disclosure receives multiple audio streams to be mixed corresponding to multiple venues. The audio streams to be mixed include the venue identifier of the venue. The audio streams to be mixed including the same venue identifier are mixed to generate the target mixed audio stream corresponding to the venue, and the target mixed audio stream is sent to multiple clients in the venue. The present disclosure mixes multiple audio streams to be mixed including the same venue identifier by identifying the venue identifier, obtains the target mixed audio stream of the venue, thereby realizing accurate mixing of different venues, and pushing the target mixed audio stream back to the corresponding venue to realize the process of mixing the audio streams of multiple large venues.
[0043] Figure 3 It is a schematic flowchart of the mixing method according to the second embodiment of the present disclosure.
[0044] As Figure 3 shown, on the basis of the embodiment shown in Figure 1 the mixing method of the embodiments of the present disclosure may specifically include the following steps:
[0045] S301, receive multiple audio streams to be mixed corresponding to multiple venues. The audio streams to be mixed include the venue identifier of the venue.
[0046] S302, divide the multiple audio streams to be mixed into multiple groups of audio streams to be mixed according to the load balancing principle.
[0047] In the embodiments of the present disclosure, the multiple received audio streams to be mixed can be divided into multiple groups of audio streams to be mixed through load balancing, so as to evenly distribute the traffic to multiple nodes, and multiple nodes jointly complete the processing of the multiple audio streams to be mixed.
[0048] S303, send the group of audio streams to be mixed to the corresponding transcoding server for transcoding processing to generate the transcoded group of audio streams to be mixed.
[0049] In the embodiments of the present disclosure, multiple to-be-mixed audio groups that have been divided are respectively sent to the transcoding servers corresponding to the audio groups, and the audio groups are transcoded to generate to-be-mixed audio groups after transcoding. For example, the to-be-mixed audio stream encoded in the opus format in the received to-be-mixed audio group is transcoded into a to-be-mixed audio stream encoded in the pcm format, so as to mix the to-be-mixed audio stream after transcoding.
[0050] S304, mix the to-be-mixed audio streams including the same venue identifier in the to-be-mixed audio group after transcoding to generate a target mixed audio stream corresponding to the venue.
[0051] In the embodiments of the present disclosure, based on the to-be-mixed audio group after transcoding, the to-be-mixed audio streams are parsed, and the to-be-mixed audio streams including the same venue identifier are mixed together to generate a target mixed audio stream for the venue corresponding to the venue identifier.
[0052] S305, send the target mixed audio stream to multiple clients in the venue.
[0053] Specifically, step S301 is the same as the above step S101, and step S305 is the same as the above step S103, which will not be elaborated here.
[0054] Figure 4 It is a schematic flowchart of a mixing method according to the third embodiment of the present disclosure.
[0055] As Figure 4 shown, on the basis of the embodiment shown in Figure 1 the mixing method of the embodiments of the present disclosure may specifically include the following steps:
[0056] S401, receive multiple to-be-mixed audio streams corresponding to multiple venues, where the to-be-mixed audio streams include the venue identifiers of the venues.
[0057] S402, based on the multi-layer mixing service, mix the to-be-mixed audio streams including the same venue identifier to generate a target mixed audio stream corresponding to the venue.
[0058] In the embodiments of the present disclosure, to disperse the mixing volume, that is, the number of audios to be mixed, and reduce the computational amount of a single mixing device, the to-be-mixed audio streams of multiple venues are mixed based on the multi-layer mixing service.
[0059] Among them, the number of layers of the multi-layer mixing service can be determined through the following process: respectively obtain the total number of to-be-mixed audio streams corresponding to a single venue, select the maximum value from the total number of to-be-mixed audio streams corresponding to multiple venues, and determine the number of layers of the multi-layer mixing service according to the maximum value.
[0060] In addition, in addition to dividing the mixing service layers according to the maximum number of audio streams to be mixed (i.e., the above maximum value), the mixing service layers can also be divided according to the maximum number of online people that the venue can accommodate, and according to the mixing volume that each layer of mixing service can undertake, and according to a variety of parameters such as the maximum number of audio streams to be mixed, the maximum number of online people, and the mixing volume that each layer of mixing service can undertake to jointly divide the mixing service layers. The present disclosure does not make any limitations.
[0061] In some embodiments, the distributed mixing server undertakes the mixing volume that each layer of mixing service needs to undertake. For example, the total number of layers of multiple layers of mixing service is N, and the i-th layer of mixing service includes K i distributed mixing servers, where i is equal to or greater than 1 and equal to or less than N. Among them, K i can be determined according to parameters such as the mixing volume that the i-th layer needs to undertake and the mixing calculation capabilities of the mixing servers included in this layer.
[0062] For example, when mixing for two venues with 100,000 people online at the same time, the mixing service is divided into two layers. For example, 1000 * 100 = 100,000. The first layer of mixing service includes 100 mixing servers, and the second layer of mixing service includes 1 mixing server to remix the mixed audio including the same venue identifier output by the 100 mixing servers in the first layer to generate a target mixed audio stream for each venue.
[0063] S403, send the target mixed audio stream to multiple clients in the venue.
[0064] Specifically, step S401 is the same as the above step S101, and step S403 is the same as the above step S103, which will not be elaborated here.
[0065] On the basis of the above embodiments, as Figure 5 shown, in step S402, "Based on multiple layers of mixing service, mix the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue", which may specifically include the following steps:
[0066] S501, for the i-th layer of mixing service, based on K i mixing servers, perform in-layer mixing processing on the audio streams to be mixed including the same venue identifier in the audio streams to be mixed corresponding to the i-th layer of mixing service, and obtain K i mixed audio streams of the i-th layer corresponding to the venue. The mixed audio streams of the i-th layer corresponding to the venue include the venue identifier.
[0067] In the embodiment of the present disclosure, the multiple mixing servers included in each layer of the mixing service mix the audio streams to be mixed that they receive and include the same venue identifier, and generate a mixed audio stream of this layer (i.e., the mixed audio stream of the i-th layer) corresponding to the venue, and the mixed audio stream of this layer includes the venue identifier. Thus, each venue can obtain K in the i-th layer of the mixing service. i The number of mixed audio streams at this layer is the same as that received by each mixing server. The number of audios to be mixed can be the same or different, and can be allocated according to the mixing capability and mixing status of the mixing server. The disclosure does not limit this. The number of mixing can be added to the frame header of the audio frame corresponding to the mixed audio stream, that is, the number of audio frames to be mixed from which the mixed audio frame is obtained.
[0068] S502, in response to i being equal to N, K i The i-th layer mixed audio stream corresponding to the venue is determined as the target mixed audio stream corresponding to the venue.
[0069] In the embodiment of the present disclosure, if i is equal to N, that is, when it is the last layer of mixing service, the K obtained by each venue in the mixing service of this layer is N The mixed audio stream of this layer (i.e., the mixed audio stream of the Nth layer) is determined as the target mixed audio stream. The number of target mixed audio streams corresponding to each venue can be one or more. If the number of target mixed audio streams is multiple, the last layer of mixing service can include multiple mixing servers, and each mixing server outputs a target mixed audio stream.
[0070] S503, in response to i not being equal to N, K i The i-th layer mixed audio stream corresponding to the venue is determined as the audio stream to be mixed corresponding to the i+1-th layer mixing service, and the i+1-th layer mixing process is performed to generate the target mixed audio stream corresponding to the venue after the N-th layer mixing process.
[0071] In the embodiment of the present disclosure, if i is not equal to N, that is, the mixing service at this layer is not the last mixing service, the K corresponding to each venue at this layer is i The i-th layer mixed audio stream is determined as the audio stream to be mixed corresponding to the i+1-th layer mixing service, and the next layer mixing service is performed until the target mixed audio stream corresponding to each venue is generated after the N-layer mixing is completed.
[0072] For example, Figure 6As shown in the figure, based on a two-layer mixing service (the first-layer mixing service includes 100 mixing servers, and the second-layer mixing service includes 1 mixing server), the audio to be mixed corresponding to two venues (100,000 audio streams to be mixed for each venue, that is, 200,000 audio streams to be mixed in total) is mixed: The client sends traffic to the Content Delivery Network (CDN) through the RTP protocol. The CDN server performs load balancing and distributes it to the transcoding servers in the receiving cluster for transcoding. The transcoded audio streams to be mixed are stored in a database (such as a Key-Value database, also known as a Redis database). The database index, timestamp, and mixing number corresponding to the audio stream to be mixed (the mixing number can be understood as how many of these audio streams to be mixed are mixed into one mixed audio stream. For example, the number of audio streams that a single mixing server in the first-layer mixing service needs to mix. Different from the above mixing number, the mixing number is used to indicate the number of audio streams to be mixed soon, and the mixing number is used to indicate the number of audio streams mixed in the previous time) are stored in a distributed message system (such as a high-throughput distributed publish-subscribe message system kafka). The first-layer mixing service is notified through kafka. The 100 mixing servers of the first-layer mixing service mix the audio streams to be mixed according to the venue identifier (for example, each mixing server mixes 1,000 audio streams to be mixed for two venues respectively, generating 1 first-layer mixed audio stream corresponding to each venue, that is, obtaining 2 first-layer mixed audio streams). Similarly, the 200 mixed audio streams obtained by the first-layer mixing service are stored in the Redis database as the audio streams to be mixed corresponding to the second mixing service layer. The database index, timestamp, and mixing number corresponding to each mixed audio stream (such as the number of audio streams that a single mixing server in the second-layer mixing service needs to mix) are notified to the second-layer mixing service through kafka. One mixing server in the second-layer mixing service mixes 100 of the above audio streams to be mixed corresponding to two venues respectively according to the venue identifier, obtaining 1 target mixed audio stream for each venue, and the target mixed audio stream is sent to multiple clients in the corresponding venue in real time. Thus, the first-layer mixing service mixes 100,000 audio streams to be mixed for each venue into 100 first-layer mixed audios corresponding to each venue according to the venue identifier, and then the second-layer mixing service mixes these 100 first-layer mixed audios to obtain 1 target mixed audio corresponding to each venue.
[0073] If the duration of a single service request is 20 ms, for a scenario with 100,000 online users in one venue, the highest Queries-per-second (QPS) that can be achieved in the embodiments of the present disclosure is QPS = 5,000,000 / s. A mixing method applicable to voice interaction in large venues is provided, which meets the requirements of large mixing calculation amount and high processing efficiency, and can reduce the required CDN service cost at the same time.
[0074] In some embodiments, the mixing server includes a plurality of mixing components, which correspond to a plurality of conference venues one by one. The mixing components are used to mix the audio streams to be mixed that contain the corresponding conference venue identifiers. In the i-th layer of mixed audio stream corresponding to a conference venue, in addition to including the conference venue identifier corresponding to the conference venue, it also includes the identifier of the mixing component that generates the mixed audio stream.
[0075] In some embodiments, the transcoding service in the receiving cluster can also be combined with the first-layer mixing service, thereby saving the consumption of Kafka and Redis and the latency caused by network transmission, and reducing costs.
[0076] Based on the above embodiments, as Figure 7 shown, in step S501, "performing mixing processing at this layer on the audio streams to be mixed that include the same conference venue identifier in the audio streams to be mixed corresponding to the i-th layer mixing service, and obtaining the i-th layer mixed audio streams corresponding to K i conference venues", specifically, it may include the following steps:
[0077] S701, preprocessing the audio streams to be mixed that include the same conference venue identifier in the audio streams to be mixed corresponding to the i-th layer mixing service.
[0078] In the embodiments of the present disclosure, the transcoding service in the receiving cluster is combined with the first-layer mixing service. The first-layer mixing service receives the audio streams to be mixed distributed by the load balancer and performs transcoding and other processing on the audio streams to be mixed.
[0079] As a feasible implementation manner, a voice review process can be added to the first-layer mixing service to perform content review on each audio stream of the user. Data cleaning and other operations can also be performed on the multiple audio streams to be mixed received before transcoding.
[0080] In the embodiments of the present disclosure, each mixing server sends, through the control layer, the audio streams to be mixed that include the same conference venue identifier in the multiple received audio streams to be mixed to the mixing component corresponding to the conference venue identifier.
[0081] In some embodiments, as Figure 8As shown, in response to i being equal to 1, that is, in the first-layer mixing service, each mixing component creates a decoding resource for each audio stream to be mixed received (that is, one decoder for each user. For the audio streams to be mixed received in the next time period, the users can be distinguished by judging the identifiers corresponding to the users in the audio stream, and the audio stream is sent to the corresponding decoder), and performs voice review on the audio streams to be mixed after the decoding operation to complete the preprocessing process of the first-layer mixing. Among them, mixing component A is the mixing component corresponding to venue A, and mixing component B is the mixing component corresponding to venue B.
[0082] In addition, the preprocessing process may further include timestamp sorting of the audio frames corresponding to the audio streams to be mixed.
[0083] In some embodiments, as Figure 9 shown, in response to i not being equal to 1, that is, in other layer mixing services except the first-layer mixing service, the control layer sends the audio streams to be mixed including the same venue identifier to the mixing component corresponding to the venue identifier. Each mixing component receives the audio streams to be mixed through a receiver, and preprocesses the received audio streams to be mixed: sorts the audio streams to be mixed including the same venue identifier in the audio streams to be mixed corresponding to the i-th layer mixing service according to the identifier of the mixing component that generated the audio streams to be mixed (the audio streams to be mixed corresponding to the i-th layer mixing service) in the (i - 1)-th layer.
[0084] As Figure 10 shown, if the previous layer mixing service includes three mixing servers, before the mixing process of the current layer mixing service, the three audio streams to be mixed received by the audio component within a preset time interval (such as within 0 - 20 ms) (which can also be understood as the audio frames corresponding to the audio streams to be mixed. For example, Figure 10 the audio stream corresponding to mixing component 1 in Figure 10 can be understood as the mixed audio stream generated by the mixer in mixing component 1 in the previous layer mixing service) are sorted according to the identifier of the mixing component, and the three audio streams to be mixed are stored in the corresponding areas in the sorting order, and these sorted audio streams to be mixed are used as the audio stream group corresponding to timestamp 000. The mixer performs mixing every 60 ms, so the receiver needs to send the three received audio stream groups to the mixer in the order of timestamps. Among them, if only two audio streams to be mixed are received in one reception due to transmission delay, as
[0085] S702, perform mixing processing on the preprocessed audio stream to be mixed to obtain the i-th layer mixed audio stream corresponding to K i conference venues.
[0086] In the embodiments of the present disclosure, the mixer in each mixing component is used to perform mixing processing on the preprocessed audio stream to be mixed to obtain the i-th layer mixed audio stream corresponding to K i conference venues in the same conference venue. As Figure 11 shown, one mixed audio stream generated by the mixer of one mixing component is obtained by the mixer performing mixing on three groups of audio streams to be mixed corresponding to three timestamps in Figure 10 respectively. That is, the mixed audio stream includes the voice data output by the upper-layer mixing components 1, 2, and 3 within 0 to 60 ms.
[0087] Based on the above embodiments, as Figure 12 shown, in response to i being equal to N, before the step "perform mixing processing on the audio stream to be mixed including the audio stream with the same conference venue identifier in the i-th layer mixing service to obtain the i-th layer mixed audio stream corresponding to K i conference venues" in step S501, the following steps may further be included:
[0088] S1201, detect whether the conference venue is in an active state.
[0089] In the embodiments of the present disclosure, as Figure 13 shown, the mixing server included in the last layer of the multi-layer mixing service can be used to manage the state of the conference venue, detect whether the conference venue is in an active state through the state management module, and control the control layer to perform conference venue management through the state management module. For example, determine the active state of the conference venue by detecting whether there is valid voice in the conference venue within a period of time.
[0090] S1202, in response to the conference venue being in an active state, perform mixing processing at this layer.
[0091] In the embodiments of the present disclosure, if the conference venue is in an active state, then perform mixing processing at this layer. The specific processing process is similar to the above embodiments and will not be elaborated here.
[0092] After the mixing processing at this layer is completed, the target mixed audio corresponding to each conference venue is sent to the RTC server through the Figure 13 transmitter shown in, and the RTC server sends the target mixed audio to multiple clients of the conference venue. For example, disassemble the 0 to 60 ms mixed audio stream in Figure 11 into three parts, and send one mixed audio stream to the RTC server every 20 ms.
[0093] If the venue is not active, release the mixing resources corresponding to the venue (such as the receiver, mixer, and transmitter shown in Figure 13 ).
[0094] In some embodiments, the last-layer mixing server stores the key fields corresponding to the venue, such as the venue identifier, the last reporting time of the venue, the active flag of the venue, and the venue key, etc. The active state of the venue can be characterized by the active flag. When each tool (receiver, mixer, and transmitter) works, it will check this flag to perform active release, thus avoiding the concurrency problems that occur in external release.
[0095] In summary, the mixing method of the embodiments of the present disclosure receives multiple audio streams to be mixed corresponding to multiple venues. The audio streams to be mixed include the venue identifier of the venue, mixes the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue, and sends the target mixed audio stream to multiple clients in the venue. The present disclosure identifies the venue identifier, mixes multiple audio streams to be mixed including the same venue identifier to obtain the target mixed audio stream of the venue, thereby realizing accurate mixing of different venues, pushing the target mixed audio stream back to the corresponding venue to implement the process of mixing the audio streams of multiple large venues, providing a mixing method applicable to voice interaction in large venues, meeting the requirements of large mixing calculation amount and high processing efficiency, and reducing the mixing cost.
[0096] Figure 14 It is a block diagram of a mixing device according to a first embodiment of the present disclosure.
[0097] As shown in Figure 14 , the mixing device 1400 of the embodiments of the present disclosure includes: a receiving module 1401, a mixing module 1402, and a sending module 1403.
[0098] The receiving module 1401 is configured to receive multiple audio streams to be mixed corresponding to multiple venues, and the audio streams to be mixed include the venue identifier of the venue.
[0099] The mixing module 1402 is configured to mix the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue.
[0100] The sending module 1403 is configured to send the target mixed audio stream to multiple clients in the venue.
[0101] It should be noted that the above explanation of the embodiments of the mixing method also applies to the mixing device of the embodiments of the present disclosure, and the specific process will not be elaborated here.
[0102] In summary, the mixing device according to the embodiments of the present disclosure receives multiple to-be-mixed audio streams corresponding to multiple venues. The to-be-mixed audio streams include venue identifiers of the venues. The to-be-mixed audio streams including the same venue identifier are mixed to generate a target mixed audio stream corresponding to the venue, and the target mixed audio stream is sent to multiple clients in the venue. The present disclosure identifies the venue identifier, mixes multiple to-be-mixed audio streams including the same venue identifier to obtain the target mixed audio stream of the venue, thereby achieving precise mixing of different venues, and pushing the target mixed audio stream back to the corresponding venue to implement the process of mixing audio streams for multiple large venues.
[0103] Figure 15 It is a block diagram of a mixing device according to the second embodiment of the present disclosure.
[0104] As Figure 15 shown, the mixing device 1500 according to the embodiments of the present disclosure includes: a receiving module 1501, a mixing module 1502, and a sending module 1503.
[0105] Among them, the receiving module 1501 has the same structure and function as the receiving module 1401 in the previous embodiment, the mixing module 1502 has the same structure and function as the mixing module 1402 in the previous embodiment, and the sending module 1503 has the same structure and function as the sending module 1403 in the previous embodiment.
[0106] Further, the mixing module 1502 includes: a second mixing sub-module, configured to mix the to-be-mixed audio streams including the same venue identifier based on a multi-layer mixing service to generate a target mixed audio stream corresponding to the venue.
[0107] Further, the i-th layer of the mixing service includes K i distributed mixing servers, where i is equal to or greater than 1 and equal to or less than N, and N is the total number of layers of the multi-layer mixing service. The second mixing sub-module includes: a mixing unit, configured to, for the i-th layer of the mixing service, based on K i mixing servers, perform in-layer mixing processing on the to-be-mixed audio streams including the same venue identifier in the to-be-mixed audio streams corresponding to the i-th layer of the mixing service to obtain the i-th layer of mixed audio streams corresponding to K i venues, and the i-th layer of mixed audio streams corresponding to the venue includes the venue identifier; a first determination unit, configured to, in response to i being equal to N, determine the i-th layer of mixed audio streams corresponding to K i venues as the target mixed audio stream corresponding to the venue; and a second determination unit, configured to, in response to i not being equal to N, use the K iThe i-th layer of mixed audio stream corresponding to the venue is determined as the audio stream to be mixed corresponding to the (i + 1)-th layer of mixing service, and the (i + 1)-th layer of mixing process is performed to generate the target mixed audio stream corresponding to the venue after N layers of mixing processes.
[0108] Further, the i-th layer of mixed audio stream corresponding to the venue further includes the identifier of the mixing component for generating the mixed audio stream.
[0109] Further, the mixing unit includes: a preprocessing subunit for preprocessing the audio streams to be mixed including the same venue identifier in the audio streams to be mixed corresponding to the i-th layer of mixing service; and a first mixing subunit for mixing the preprocessed audio streams to be mixed to obtain the i-th layer of mixed audio streams corresponding to K i venues.
[0110] Further, in response to i being equal to 1, the preprocessing includes: a decoding operation and a voice review operation; in response to i not being equal to 1, the preprocessing includes: sorting the audio streams to be mixed including the same venue identifier in the audio streams to be mixed corresponding to the i-th layer of mixing service according to the identifier of the mixing component.
[0111] Further, in response to i being equal to N, the mixing unit further includes: a detection subunit for detecting whether the venue is in an active state; a release subunit for releasing the mixing component corresponding to the venue in response to the venue not being in an active state; and a second mixing subunit for performing the mixing process of this layer in response to the venue being in an active state.
[0112] Further, the apparatus 1500 further includes: an acquisition module 1504 for acquiring the total number of audio streams to be mixed corresponding to a single venue; a determination module 1505 for determining the number of layers of the multi-layer mixing service according to the maximum value among the total numbers of audio streams to be mixed corresponding to multiple venues.
[0113] Further, the mixing module 1502 includes: a division sub-module for dividing multiple audio streams to be mixed into multiple groups of audio streams to be mixed according to the load balancing principle; a transcoding sub-module for sending the groups of audio streams to be mixed to the corresponding transcoding servers for transcoding to generate the transcoded groups of audio streams to be mixed; and a first mixing sub-module for mixing the audio streams to be mixed including the same venue identifier in the transcoded groups of audio streams to be mixed to generate the target mixed audio stream corresponding to the venue.
[0114] In summary, the mixing device according to the embodiments of the present disclosure receives multiple to-be-mixed audio streams corresponding to multiple conference venues. The to-be-mixed audio streams include the venue identifiers of the conference venues. It mixes the to-be-mixed audio streams including the same venue identifier to generate a target mixed audio stream corresponding to the venue, and sends the target mixed audio stream to multiple clients in the venue. The present disclosure identifies the venue identifier, mixes multiple to-be-mixed audio streams including the same venue identifier to obtain the target mixed audio stream of the venue, thereby achieving precise mixing of different conference venues, pushing the target mixed audio stream back to the corresponding venue to implement the process of mixing audio streams for multiple large conference venues, providing a mixing method applicable to voice interaction in large conference venues, meeting the requirements of large mixing calculation amount and high processing efficiency, and reducing the mixing cost.
[0115] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0116] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0117] Figure 16 FIG. shows a schematic block diagram of an exemplary electronic device 1600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0118] As Figure 16 shown, the electronic device 1600 includes a computing unit 1601, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1602 or the computer program loaded from the storage unit 1608 into the random access memory (RAM) 1603. In the RAM 1603, various programs and data required for the operation of the electronic device 1600 can also be stored. The computing unit 1601, the ROM 1602, and the RAM 1603 are connected to each other through a bus 1604. The input / output (I / O) interface 1605 is also connected to the bus 1604.
[0119] Multiple components in the electronic device 1600 are connected to the I / O interface 1605, including: an input unit 1606, such as a keyboard, a mouse, etc.; an output unit 1607, such as various types of displays, speakers, etc.; a storage unit 1608, such as a magnetic disk, an optical disc, etc.; and a communication unit 1609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1609 allows the electronic device 1600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0120] The computing unit 1601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1601 executes the various methods and processes described above, such as Figures 1 to 13 the mixing method shown. For example, in some embodiments, the mixing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1600 via the ROM 1602 and / or the communication unit 1609. When the computer program is loaded into the RAM 1603 and executed by the computing unit 1601, one or more steps of the semantic parsing method described above can be executed. Alternatively, in other embodiments, the computing unit 1601 can be configured to execute the mixing method in any other suitable way (e.g., by means of firmware).
[0121] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0122] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0123] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0124] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0125] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.
[0126] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0127] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the steps of the mixing method as shown in the above embodiments of the present disclosure.
[0128] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.
[0129] The above specific embodiments do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A mixing method, comprising: Receiving a plurality of audio streams to be mixed corresponding to a plurality of venues, wherein the audio streams to be mixed include the venue identifiers of the venues; Performing mixing processing on the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue; And Sending the target mixed audio stream to a plurality of clients in the venue; Performing mixing processing on the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue, including: Based on a multi-layer mixing service, perform mixing processing on the to-be-mixed audio streams including the same venue identifier to generate a target mixed audio stream corresponding to the venue; wherein, the i-th layer of the mixing service includes distributed mixing servers, and each mixing server includes a plurality of mixing components for performing mixing processing on the to-be-mixed audio streams containing the corresponding venue identifier, and the plurality of mixing components correspond to the plurality of venues one by one, where i is equal to or greater than 1 and equal to or less than N, and N is the total number of layers of the multi-layer mixing service.
2. The mixing method according to claim 1, wherein, The performing mixing processing on the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue further includes: Dividing the plurality of audio streams to be mixed into a plurality of groups of audio streams to be mixed according to the load balancing principle; Sending the groups of audio streams to be mixed to corresponding transcoding servers for transcoding processing to generate transcoded groups of audio streams to be mixed; and Performing mixing processing on the audio streams to be mixed including the same venue identifier in the transcoded groups of audio streams to be mixed to generate the target mixed audio stream corresponding to the venue.
3. The mixing method according to claim 1, wherein, The performing mixing processing on the audio streams to be mixed including the same venue identifier based on a multi-layer mixing service to generate the target mixed audio stream corresponding to the venue includes: For the i-th layer of mixing service, based on of the mixing servers, perform in-layer mixing processing on the audio streams to be mixed corresponding to the i-th layer of mixing service that include the same venue identifier, and obtain the i-th layer of mixed audio streams corresponding to the venues, where the i-th layer of mixed audio streams corresponding to the venues includes the venue identifier; In response to the i being equal to N, set the i-th layer mixed audio stream corresponding to the venue as the target mixed audio stream corresponding to the venue; and In response to the condition that i is not equal to N, the i-th layer mixed audio stream corresponding to the venue is determined as the audio stream to be mixed corresponding to the (i + 1)-th layer mixing service, and (i + 1)-th layer mixing processing is performed to generate the target mixed audio stream corresponding to the venue after N-layer mixing processing.
4. The mixing method according to claim 1, wherein, The i-th layer mixed audio stream corresponding to the venue further includes the identifier of the mixing component that generates the mixed audio stream.
5. The mixing method according to claim 3, wherein, Performing mixing processing on the audio streams to be mixed in the i-th layer mixing service that include the audio streams with the same venue identifier to obtain the i-th layer mixed audio streams corresponding to the venues, including: Performing preprocessing on the audio streams to be mixed including the same venue identifier in the audio streams to be mixed corresponding to the i-th layer mixing service; and Perform mixing processing on the preprocessed audio stream to be mixed to obtain the i-th layer mixed audio stream corresponding to the venue 6. The mixing method according to claim 5, wherein, In response to i being equal to 1, the preprocessing includes: a decoding operation and a voice review operation; In response to i not being equal to 1, the preprocessing includes: Sorting the audio streams to be mixed including the same venue identifier in the audio streams to be mixed corresponding to the i-th layer mixing service according to the identifier of the mixing component.
7. The mixing method according to claim 3, wherein, In response to the fact that i is equal to N, before performing mixing processing on the audio streams to be mixed in the i-th layer mixing service that include the audio streams with the same venue identifier to obtain the mixed audio stream of the i-th layer corresponding to the venue, it further includes: Detecting whether the venue is in an active state; In response to the venue not being in an active state, releasing the mixing component corresponding to the venue; In response to the venue being in an active state, performing mixing processing at this layer.
8. The mixing method according to claim 1, further comprising: Obtaining the total number of the audio streams to be mixed corresponding to a single venue; Determining the number of layers of the multi-layer mixing service according to the maximum value among the total numbers of the audio streams to be mixed corresponding to the plurality of venues.
9. A mixing device, comprising: A receiving module, configured to receive a plurality of audio streams to be mixed corresponding to a plurality of venues, wherein the audio streams to be mixed include the venue identifiers of the venues; A mixing module, configured to perform mixing processing on the audio streams to be mixed including the same venue identifier to generate a target mixed audio stream corresponding to the venue; And A sending module, configured to send the target mixed audio stream to a plurality of clients in the venue; The mixing module includes: A second mixing sub-module, configured to perform mixing processing on the to-be-mixed audio streams including the same venue identifier based on a multi-layer mixing service, so as to generate a target mixed audio stream corresponding to the venue; wherein, the i-th layer of mixing service includes distributed mixing servers, and each of the mixing servers includes a plurality of mixing components for performing mixing processing on the to-be-mixed audio streams including the corresponding venue identifier, and the plurality of mixing components correspond to the plurality of venues one by one, where i is equal to or greater than 1 and equal to or less than N, and N is the total number of layers of the multi-layer mixing service.
10. The mixing device according to claim 9, wherein, The mixing module further includes: A dividing sub-module, configured to divide the plurality of audio streams to be mixed into a plurality of groups of audio streams to be mixed according to the load balancing principle; A transcoding sub-module, configured to send the to-be-mixed audio group to a corresponding transcoding server for transcoding processing to generate a transcoded to-be-mixed audio group; and A first mixing sub-module, configured to mix the to-be-mixed audio streams including the same venue identifier in the transcoded to-be-mixed audio group to generate the target mixed audio stream corresponding to the venue.
11. The mixing device according to claim 9, wherein, The second mixing sub-module includes: A mixing unit, for the i-th layer mixing service, based on of the mixing servers, perform in-layer mixing processing on the audio streams to be mixed corresponding to the i-th layer mixing service that include the same venue identifier, and obtain i-th layer mixed audio streams corresponding to the venues, where the i-th layer mixed audio streams corresponding to the venues include the venue identifier; A first determination unit, configured to, in response to i being equal to N, determine the i-th layer mixed audio stream corresponding to the number of the venues as the target mixed audio stream corresponding to the venue; and A second determination unit, configured to, in response to the i not being equal to N, determine the i-th layer mixed audio stream corresponding to of the venues as the audio stream to be mixed corresponding to the (i + 1)-th layer mixing service, and perform (i + 1)-th layer mixing processing, so as to generate the target mixed audio stream corresponding to the venues after N-layer mixing processing. venues as the audio stream to be mixed corresponding to the (i + 1)-th layer mixing service, and perform (i + 1)-th layer mixing processing, so as to generate the target mixed audio stream corresponding to the venues after N-layer mixing processing.
12. The mixing device according to claim 9, wherein, The i-th layer mixed audio stream corresponding to the venue further includes the identifier of the mixing component that generates the mixed audio stream.
13. The mixing device according to claim 11, wherein, The mixing unit includes: A preprocessing sub-unit, configured to preprocess the to-be-mixed audio streams including the same venue identifier in the to-be-mixed audio stream corresponding to the i-th layer mixing service; and The first mixing subunit is configured to perform mixing processing on the to-be-mixed audio stream after preprocessing to obtain the i-th layer mixed audio stream corresponding to the said venue.
14. The mixing device according to claim 13, wherein, In response to i being equal to 1, the preprocessing includes: a decoding operation and a voice review operation; In response to i not being equal to 1, the preprocessing includes:[[]] Sorting the to-be-mixed audio streams including the same venue identifier in the to-be-mixed audio stream corresponding to the i-th layer mixing service according to the identifier of the mixing component.
15. The mixing device according to claim 11, wherein, In response to i being equal to N, the mixing unit further includes:[[]] A detection sub-unit, configured to detect whether the venue is in an active state; A release sub-unit, configured to release the mixing component corresponding to the venue in response to the venue not being in an active state; A second mixing sub-unit, configured to perform mixing processing at this layer in response to the venue being in an active state.
16. The mixing device according to claim 9, further comprising:[[]] An acquisition module, configured to acquire the total number of the to-be-mixed audio streams corresponding to a single venue; A determination module, configured to determine the number of layers of the multi-layer mixing service according to the maximum value among the total numbers of the to-be-mixed audio streams corresponding to multiple venues.
17. An electronic device, comprising:[[]] At least one processor; And A memory communicatively connected to the at least one processor; wherein,[[]] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
19. A computer program product, comprising a computer program, and the computer program implements the steps of the method according to any one of claims 1-8 when being executed by a processor.
Citation Information
Patent Citations
Distributed audio mixing method for conference call
CN108712584A
Sound mixing method and device
CN114217996A