Method for converting mono to multi-channel based on audio processing and related device
By calculating the speaker gain and converting mono data into multi-channel data, the problem of stereo data transmission delay in video conferencing was solved, achieving multi-channel effects while reducing network bandwidth.
Patent Information
- Application Number
- CN202310069278.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-02-06
AI Technical Summary
In video conferencing, stereo or multi-channel data transmission may cause network latency, affecting real-time performance, and existing technologies struggle to achieve multi-channel effects while reducing network bandwidth.
By determining the position of the target sound image, the gain value of each speaker is calculated based on the vector relationship between the position and gain value of multiple speakers. The mono data is then multiplied by the gain value of each speaker to obtain the channel data for playback.
Without increasing network bandwidth, multi-channel and stereo effects were achieved, improving the audio quality of video conferencing.
Smart Images

Figure CN116055985B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio, and in particular to a method for converting mono sound into multi sound based on audio processing and related equipment. BACKGROUND
[0002] At present, in the process of video conference, users often have the demand for audio playing with stereo effect or directional space sense. However, the transmission of stereo data or multi sound data may cause delay, thereby affecting the real-time performance of video conference.
[0003] Therefore, how to reduce the network bandwidth of video conference while obtaining multi sound effect is a technical problem to be solved by those skilled in the art. SUMMARY
[0004] The present application is provided to overcome the defects of the prior art, and provides a method for converting mono sound into multi sound based on audio processing and related equipment, which can reduce the network bandwidth of video conference while obtaining multi sound effect.
[0005] According to an aspect of the present application, a method for converting mono sound into multi sound based on audio processing is provided, comprising:
[0006] determining the position of a target sound image based on conference information of video conference;
[0007] calculating the gain value of each loudspeaker based on the vector relationship among the positions of the plurality of loudspeakers, the gain value of each loudspeaker and the position of the target sound image;
[0008] multiplying the received mono sound data with the gain value of each loudspeaker respectively, and playing the channel data of each loudspeaker.
[0009] In some embodiments of the present application, the position of the target sound image is determined based on the position information carried in the message sent by the remote participant in the conference information.
[0010] In some embodiments of the present application, the position information carried in the message by the remote participant in the conference is obtained based on the microphone array positioning of the remote participant.
[0011] In some embodiments of the present application, the position of the target sound image is determined based on the position of the remote participant in the display screen in the conference information.
[0012] In some embodiments of the present application, the gain value of each loudspeaker is calculated based on the vector relationship among the positions of the plurality of loudspeakers, the gain value of each loudspeaker and the position of the target sound image, comprising:
[0013] obtaining a position coordinate of the target sound image in a Cartesian coordinate system, the Cartesian coordinate system taking a set audience position as an origin;
[0014] obtaining position coordinates of the plurality of loudspeakers in a Cartesian coordinate system;
[0015] calculating a gain value of each loudspeaker based on the position coordinate of the target sound image and the position coordinates of the plurality of loudspeakers.
[0016] In some embodiments of the present application, the gain value of each loudspeaker is calculated according to the following formula:
[0017]
[0018] wherein g is the gain value of each loudspeaker, [p1...pn] is the position coordinate of the target sound image, is the position coordinate of the plurality of loudspeakers, and n is 2 or 3.
[0019] In some embodiments of the present application, the calculating the gain value of each loudspeaker based on the vector relationship between the position of the plurality of loudspeakers, the gain value of each loudspeaker and the position of the target sound image comprises:
[0020] performing amplitude gain normalization processing on the gain value of each loudspeaker.
[0021] According to still another aspect of the present application, there is provided an apparatus for converting a mono sound into a multi-channel sound based on audio processing, comprising:
[0022] a determining module configured to determine a position of a target sound image based on conference information of a video conference;
[0023] a gain value determining module configured to calculate a gain value of each loudspeaker based on a vector relationship between the position of the plurality of loudspeakers, the gain value of each loudspeaker and the position of the target sound image;
[0024] a channel data calculating module configured to multiply the received mono sound data with the gain value of each loudspeaker respectively, and play the multiplied data as channel data of each loudspeaker.
[0025] According to still another aspect of the present application, there is also provided an electronic device, comprising: a processor; a storage medium having a computer program stored thereon, the computer program being executed by the processor to perform the steps as described above.
[0026] According to still another aspect of the present application, there is also provided a storage medium having a computer program stored thereon, the computer program being executed by a processor to perform the steps as described above.
[0027] Therefore, the solution provided in this application has the following advantages compared with the prior art:
[0028] In video conferencing, by determining the location of the target sound image, the gain value of each speaker can be calculated based on the positions of multiple speakers, the vector relationship between the gain value of each speaker and the position of the target sound image, and then the received mono data is multiplied by the gain value of each speaker to obtain the channel data for playback. Therefore, participating terminals do not need to transmit stereo or multi-channel data; they only need to transmit mono data, which can then be converted into multi-channel data. This saves network bandwidth while providing multi-channel and stereo effects. Attached Figure Description
[0029] The above and other features and advantages of this application will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.
[0030] Figure 1 A flowchart is shown of a method for converting mono to multichannel based on audio processing according to an embodiment of this application.
[0031] Figure 2 A schematic diagram illustrating the calculation of gain values based on two loudspeakers according to an embodiment of this application is shown.
[0032] Figure 3 A schematic diagram is shown illustrating the calculation of gain values for two loudspeakers selected from a plurality of loudspeakers based on the location of the target acoustic image, according to an embodiment of this application.
[0033] Figure 4 A schematic diagram illustrating the calculation of gain values based on three speakers according to an embodiment of this application is shown.
[0034] Figure 5 A flowchart of a video conference according to an embodiment of this application is shown.
[0035] Figure 6 A block diagram of an apparatus for converting mono to multichannel based on audio processing according to an embodiment of this application is shown.
[0036] Figure 7 This illustration schematically depicts a computer-readable storage medium according to an exemplary embodiment of the present disclosure.
[0037] Figure 8 The schematic diagram illustrates an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0038] Example implementations are now described with reference to the drawings. Example implementations can, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the inventive aspects to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations.
[0039] In addition, the accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application. In the drawings:
[0040] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the steps. For example, some steps can be further decomposed, and some steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.
[0041] In order to overcome the defects of the prior art described above, the present application provides a method for converting a single channel into a multi-channel based on audio processing and related equipment, which can reduce the network bandwidth of a video conference while obtaining a multi-channel effect.
[0042] Reference will now be made to Figure 1 , Figure 1 A flowchart of a method for converting a single channel into a multi-channel based on audio processing according to an embodiment of the present application is shown. The method for converting a single channel into a multi-channel based on audio processing includes:
[0043] Step S110: determining the position of a target sound image based on the conference information of the video conference.
[0044] Specifically, within the space where the plurality of loudspeakers are located, the virtual sound source to be formed by the multi-channel audio played by the plurality of loudspeakers (i.e. to make the listener experience that the sound is from the virtual sound source) is the target sound image.
[0045] Specifically, the present application is applied in a video conference, so that the position of the target sound image can be determined based on the transmission data of the conference information of the video conference or the display settings of the video conference.
[0046] In some embodiments, the position of the target sound image can be determined based on the conference information of the remote participant sending a message carrying the position information. When the remote participant is provided with a plurality of microphones to form a microphone array, the sound source positioning can be based on the microphone array, or the direction of the microphone with the largest proportion in the microphone array mixing algorithm can be taken as the sound source. Thus, the position information can be obtained more efficiently through the microphone matrix. After the remote participant obtains the position of the target sound image, the position information can be sent to the near participant through a control message transmission protocol. In some specific implementations, the position information can be transmitted based on a control message transmission protocol, such as a Realtime Transport Control Protocol (RTCP). For example, a position field can be added to the message to transmit the position information. In other specific implementations, the position information can also be transmitted through another private protocol. The position information is transmitted through the message, and the data amount of the position information is small, so that the network bandwidth is not occupied, thereby improving the transmission efficiency. The present application can realize more changes, which will not be described here.
[0047] In other embodiments, in a video conference, each remote participant is displayed on the display screen of the near participant, so that the position of the target sound image can be determined based on the position of the remote participant in the display screen in the conference information. For example, when the participant terminal is located on the left side of the display screen, the position of the target sound image can be determined as the left side; when the participant terminal is located on the right side of the display screen, the position of the target sound image can be determined as the right side.
[0048] Step S120: calculating the gain value of each loudspeaker based on the position of the plurality of loudspeakers, the gain value of each loudspeaker, and the vector relationship between the position of the target sound image.
[0049] Specifically, according to the Haas effect or the de Boer effect, the position of the target sound image is the vector sum of the position of each loudspeaker and the gain value thereof. Thus, the position of each loudspeaker and the position of the target sound image can be converted to the Cartesian coordinate system, thereby facilitating the calculation of the gain value of each loudspeaker. Specifically, step S120 can include: obtaining the position coordinates of the target sound image in the Cartesian coordinate system, the Cartesian coordinate system taking the set audience position as the origin; obtaining the position coordinates of the plurality of loudspeakers in the Cartesian coordinate system; and calculating the gain value of each loudspeaker based on the position coordinates of the target sound image and the position coordinates of the plurality of loudspeakers.
[0050] Further, step S120 further includes a step of performing amplitude gain normalization processing on the gain value of each loudspeaker to ensure that the target sound image is located in the target sound image activity region range formed between the plurality of loudspeakers and the set audience position.
[0051] Step S130: multiplying the received monaural data with the gain value of each speaker respectively, and playing as the channel data of each speaker.
[0052] The method for converting monaural data into multi-channel data based on audio processing provided in the present application can obtain the gain value of each speaker based on the position of the multi-speakers, the vector relationship between the gain value of each speaker and the position of the target sound image, and multiply the received monaural data with the gain value of each speaker respectively, and play as the channel data of each speaker. Thus, each terminal in the conference only needs to transmit monaural data, and can convert the monaural data into multi-channel data, save network bandwidth, and obtain multi-channel and stereo sound effect.
[0053] Reference will be made to Figure 2 , Figure 2 Fig. 1 shows a schematic diagram for calculating the gain value based on two speakers according to an embodiment of the present application. As shown in Fig. 1, a Cartesian coordinate system is formed with the audience position set, and two speakers are a left channel speaker and a right channel speaker. Figure 2
[0054] Based on the vector relationship, the position of the target sound image P = g1*U1+ g2*U2, wherein U1 is the coordinate of the left channel speaker, U2 is the coordinate of the right channel speaker, g1 is the gain value of the left channel speaker, and g2 is the gain value of the right channel speaker. The formula can be converted into P T = g*U12, wherein, P = [p1, p2] T , p1 and p2 are the x and y axis coordinates of the target sound image in the Cartesian coordinate system, and U12 = [u11, u12]T. , wherein u11 and u12 are the x and y axis coordinates of the left channel speaker in the Cartesian coordinate system, and u21 and u22 are the x and y axis coordinates of the right channel speaker in the Cartesian coordinate system.
[0055] Thus, the gain value can be calculated based on the following formula:
[0056]
[0057] Further, when calculating the gain value, the gain value can also be subjected to amplitude gain normalization, that is, g1 2 + g2 2 = 1. Thus, the gain value g1 of the left channel speaker and the gain value g2 of the right channel speaker can be solved.
[0058] Reference will be made to Figure 3 , Figure 3 Fig. 1 shows a schematic diagram of calculating gain values of two loudspeakers based on the position of a target sound image according to an embodiment of the present application. In this embodiment, three horizontally placed loudspeakers are provided at the near end, thus, two loudspeakers closer to the target sound image can be determined based on the position of the target sound image, and the gain values of the two loudspeakers are calculated by using the following formula: Figure 2 The gain values of the two loudspeakers are calculated by using the calculation method of the gain values as shown in Fig. 1. The present application is not limited to this, and more loudspeakers are also within the protection scope of the present application.
[0059] In the following, the calculation of the gain values of three loudspeakers will be described with reference to Figure 4 , Figure 4 Fig. 2 shows a schematic diagram of calculating gain values of three loudspeakers according to an embodiment of the present application. As shown in Fig. 2, a Cartesian coordinate system is formed with the audience position, and three loudspeakers are channel1, channel2 and channel3, and the three loudspeakers form a three-dimensional space. Figure 4
[0060] Based on the vector relationship, the position of the target sound image P = g1*U1+g2*U2+g3*U3, wherein U1 is the coordinate of the loudspeaker channel1, U2 is the coordinate of the loudspeaker channel2, U3 is the coordinate of the loudspeaker channel3, g1 is the gain value of the loudspeaker channel1, g2 is the gain value of the loudspeaker channel2, and g3 is the gain value of the loudspeaker channel3. The formula can be converted to P = g*U123, wherein, T P = [p1, p2, p3] T , p1, p2, p3 are respectively the x, y, z axis coordinates of the target sound image in the Cartesian coordinate system, wherein u11, u12, u13 are respectively the x, y, z coordinates of the loudspeaker channel1 in the Cartesian coordinate system, u21, u22, u23 are respectively the x, y, z coordinates of the loudspeaker channel2 in the Cartesian coordinate system, and u31, u32, u33 are respectively the x, y, z coordinates of the loudspeaker channel3 in the Cartesian coordinate system.
[0061] Thus, the gain values can be calculated based on the following formula:
[0062]
[0063] Further, when calculating the gain values, the gain values can also be subjected to amplitude gain normalization, that is, to make g1 2 + g2 2 + g3 2 = 1. Thus, the gain value g1 of the loudspeaker channel 1, the gain value g2 of the loudspeaker channel 2, and the gain value g3 of the channel 3 can be solved.
[0064] Referring to the following Figure 5 , Figure 5 A flowchart of a video conference according to an embodiment of the present application is shown.
[0065] As Figure 5 shown, the microphone array of the near-end conference terminal collects voice data of the sound source at the near-end.
[0066] Step S101: The sound source position is located by the microphone array at the near-end conference terminal, and the voice data is subjected to 3A (AEC echo control, ANS voice noise reduction, AGC voice enhancement) voice pre-processing.
[0067] Step S102: The position information of the sound source is added to the reserved field in the Real-time Transport Control Protocol (RTCP) packet or the custom private protocol, and is sent to the far-end conference terminal.
[0068] Step S103: The single-channel voice data is transmitted based on the Real-time Transport (RTP) protocol, and is sent to the far-end conference terminal.
[0069] The near-end conference terminal receives the RTCP packet or the custom private protocol packet and the single-channel voice data sent by the far-end conference terminal.
[0070] Step S104: The single-channel voice data and the position information are extracted.
[0071] Step S105: When there are multiple parties in the conference, the position information of each participant in the display screen of the application layer of the near-end conference terminal is extracted.
[0072] Step S106: The position information of the target sound image is determined.
[0073] Specifically, when the number of the far-end conference terminals is one, the position information extracted in step S104 can be used as the position information of the target sound image. When the number of the far-end conference terminals is multiple, the position information of the participant currently speaking in the display screen determined in step S105 can be used as the position information of the target sound image.
[0074] Step S107: The single-channel data and the position information of the target sound image are used for upmixing, and the corresponding channel data is played by each loudspeaker.
[0075] Thus, in the embodiment, the two parties do not need to transmit stereo data or multi-channel data, but only need to transmit mono data and position information, and stereo data and multi-channel data are obtained by up-mixing at the receiving end, so that stereo effect or directional space sense is obtained while network bandwidth is saved; in a multi-party conference, display position information of each party in the display in the application layer is extracted as position information of a target sound image, so that audio playing effect is in harmony with the picture of the speaker by gain value calculation of each loudspeaker; considering that the near-end conference participant mainly views the display screen of the near-end, so when the position signal from the far-end conference and the display position information of each party in the picture in the multi-party conference both exist, the display position information in the picture is used as the position information of the target sound image.
[0076] The above exemplary illustrates the multiple implementation manners of the present application, and the present application is not limited thereto, and in each embodiment, the addition, omission or sequence change of the steps are within the protection scope of the present application; each embodiment can be realized alone or in combination.
[0077] The following will be described in combination with Figure 5 The device for converting mono to multi-channel based on audio processing provided by the present application is described. The device for converting mono to multi-channel based on audio processing 200 comprises a determination module 210, a gain value determination module 220 and a channel data calculation module 230.
[0078] The determination module 210 is configured to determine the position of a target sound image based on conference information of a video conference;
[0079] The gain value determination module 220 is configured to calculate the gain value of each loudspeaker based on the vector relationship among the positions of the multiple loudspeakers, the gain value of each loudspeaker and the position of the target sound image;
[0080] The channel data calculation module 230 is configured to multiply the received mono data with the gain value of each loudspeaker respectively, and play as the channel data of each loudspeaker.
[0081] In the device for converting mono to multi-channel based on audio processing provided by the present application, in a video conference, the position of a target sound image is determined, so that the gain value of each loudspeaker can be calculated based on the vector relationship among the positions of the multiple loudspeakers, the gain value of each loudspeaker and the position of the target sound image, and the received mono data is multiplied with the gain value of each loudspeaker respectively, and played as the channel data of each loudspeaker. Thus, each terminal of the conference does not need to transmit stereo data or multi-channel data, but only needs to transmit mono data, so that mono data can be converted to multi-channel data, and multi-channel and stereo effect can be obtained while network bandwidth is saved.
[0082] The application can realize the device for converting a single sound track into a multi-sound track based on audio processing and the face detection device by software, hardware, firmware and any combination thereof. The splitting, merging and adding of modules are within the protection scope of the application without violating the concept of the application.
[0083] In the exemplary embodiments of the present disclosure, a computer readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, can realize the steps of the method for converting a single sound track into a multi-sound track based on audio processing in any one of the above embodiments. In some possible implementation manners, various aspects of the application can also be realized in the form of a program product, which includes program codes, and the program codes are configured to make a terminal device execute the steps according to various exemplary embodiments of the application described in the above method part of the present description based on audio processing for converting a single sound track into a multi-sound track when the program product runs on the terminal device.
[0084] Reference Figure 7 As shown, the program product 800 configured to realize the above method according to the embodiments of the application can adopt a portable compact disc read-only memory (CD-ROM) and includes program codes, and can run on a terminal device, such as a personal computer. However, the program product of the application is not limited thereto, and in the present document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, device or apparatus.
[0085] The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0086] The computer readable storage medium can include a data signal transported over a carrier wave and can be baseband or propagated along with carriers. The program code embodied on the computer readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0087] The program code can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0088] In an exemplary embodiment of the present disclosure, an electronic device is also provided, which can include a processor, and a memory configured to store executable instructions of the processor. Wherein the processor is configured to perform the steps of the method for converting a mono sound into a multi sound based on audio processing in any one of the above embodiments via executing the executable instructions.
[0089] Those skilled in the art can understand that each aspect of the present application can be implemented as a system, a method or a program product. Therefore, each aspect of the present application can be specifically implemented as follows: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining software and hardware aspects, which can be collectively referred to as "circuitry", "module" or "system".
[0090] The electronic device 600 according to this embodiment of the present application will be described below with reference to Figure 8 Figure 8 The displayed electronic device 600 is only an example and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0091] AsFigure 8 As shown, the electronic device 600 is in the form of a general computing device. Components of the electronic device 600 can include, but are not limited to, at least one processing unit 610, at least one memory unit 620, a bus 630 that connects different system components including the memory unit 620 and the processing unit 610, a display unit 640, and the like.
[0092] The memory unit stores program codes which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present application described in the above method part of the present specification based on audio processing for converting a mono track into a multi-track. For example, the processing unit 610 can perform the steps as shown in FIG. 6. Figure 1
[0093] The memory unit 620 can include a readable medium in the form of a volatile memory unit, such as a random access memory (RAM) 6201 and / or a cache memory unit 6202, and can further include a read-only memory (ROM) 6203.
[0094] The memory unit 620 can further include program / utility 6204 having a set of the program modules 6205, such as an operating system, one or more application programs, other program modules, and program data, and each of these examples, or some combination thereof, can include implementation of a network environment.
[0095] The bus 630 can be representative of one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, a processor or local bus using any of a variety of bus structures, and the like.
[0096] The electronic device 600 can also communicate with one or more external devices 700 such as a keyboard, a pointing device, a Bluetooth device, or a CD player, and can communicate with one or more devices that enable a user to interact with the electronic device 600 and / or any devices (e.g., a router, a modem, or the like) that enable the electronic device 600 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 650. Still yet, the electronic device 600 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or the Internet through network adapter 660. The network adapter 660 can communicate with the other components of the electronic device 600 via bus 630. It should be understood that although not shown, other hardware and / or software components could be used in conjunction with the electronic device 600. These include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0097] Through the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by software in combination with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or on a network, and includes a plurality of instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to perform the above-mentioned method of converting a monaural sound into a multi-channel sound based on audio processing according to the embodiments of the present disclosure.
[0098] In the video conference, the position of the target sound image is determined, so that the gain values of the speakers are calculated based on the vector relationship between the positions of the plurality of speakers, the gain values of the speakers, and the position of the target sound image, and the received monaural data is multiplied by the gain values of the speakers respectively, and played as channel data of the speakers. Thus, each terminal participating in the conference does not need to transmit stereo data or multi-channel data, but only needs to transmit monaural data, so that the monaural data can be converted into multi-channel data, thereby saving network bandwidth and obtaining multi-channel and stereo effect.
[0099] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the aspects disclosed herein. It is intended that the present disclosure cover any and all variations of the present disclosure that come within the scope of the claims and their equivalents. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the disclosure indicated by the following claims.
Claims
1. A method for converting mono to multi-channel audio based on audio processing, the method being applied to a near-end of a meeting, characterized in that, include: In a video conference, receive messages carrying sound source location information and mono audio data sent by remote participants; Based on the message and the mono audio data, the location of the target sound image is determined, wherein determining the location of the target sound image includes: When the number of participating remote ends is one, the location of the target acoustic image is determined based on the location information carried in the message; When there are multiple participants at the remote end, extract the position information of each participant in the display screen from the application layer at the near end to determine the position of the target sound image; The gain value of each speaker is calculated based on the vector relationship between the positions of multiple speakers, the gain value of each speaker, and the determined position of the target sound image. The mono audio data is multiplied by the gain value to obtain the channel data for each speaker for playback.
2. The method for converting mono to multi-channel audio based on audio processing as described in claim 1, characterized in that, The remote participant obtains its location based on the microphone array of the remote participant using the location information carried in the message.
3. The method for converting mono to multi-channel audio based on audio processing as described in claim 1, characterized in that, The calculation of the gain value of each speaker based on the vector relationship between the positions of multiple speakers, the gain value of each speaker, and the position of the target sound image includes: Obtain the position coordinates of the target sound image in a Cartesian coordinate system with the set audience position as the origin; Obtain the position coordinates of the plurality of speakers in the Cartesian coordinate system; Based on the position coordinates of the target sound image and the position coordinates of the multiple loudspeakers, the gain value of each loudspeaker is calculated.
4. The method for converting mono to multi-channel audio based on audio processing as described in claim 3, characterized in that, The gain value of each speaker is calculated according to the following formula: , Where g is the gain value of each speaker, The coordinates of the target acoustic image are given. Let n be the position coordinates of the plurality of speakers, where n is 2 or 3.
5. The method for converting mono to multi-channel audio based on audio processing as described in claim 1, characterized in that, The calculation of the gain value of each speaker based on the vector relationship between the positions of multiple speakers, the gain value of each speaker, and the position of the target sound image includes: Amplitude gain normalization is performed on the gain values of each speaker.
6. A device for converting mono to multi-channel audio based on audio processing, characterized in that, Implementing the method for converting mono to multichannel based on audio processing as described in any one of claims 1 to 5, comprising: The module is configured to determine the location of the target audio-visual image based on the meeting information of the video conference. The gain value determination module is configured to calculate the gain value of each speaker based on the position of multiple speakers, the vector relationship between the gain value of each speaker and the position of the target sound image; The channel data calculation module is configured to multiply the received mono data by the gain value of each speaker to obtain the channel data for each speaker for playback.
7. An electronic device, characterized in that, The electronic device includes: processor; A memory, on which a computer program is stored, is executed by the processor during runtime: The method for converting mono to multichannel based on audio processing as described in any one of claims 1 to 5.
8. A storage medium, characterized in that, The storage medium stores a computer program, which is executed by a processor: The method for converting mono to multichannel based on audio processing as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Three-dimensional audio signal generation method and system for non-spherical speaker array
CN105392102A
Immersive conference system based on multi-excitation flat panel loudspeaker
CN112584299A
Information processing method, electronic equipment, conference system and medium
CN115002401A