Image and sound localization method and system thereof
By dividing the acquisition area in the video conference room and using cascaded routers and audio processors, the problems of accurate matching and real-time performance of audio-visual synchronization technology in complex environments were solved, enabling an immersive audio-visual experience in scenarios of different scales.
Patent Information
- Application Number
- CN202410139812.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-01-31
AI Technical Summary
Existing audio-visual synchronization technology struggles to achieve accurate sound and image matching in complex environments, suffers from insufficient real-time performance, and its hardware limitations restrict system scalability, making it difficult to adapt to meeting scenarios of different sizes.
The video conference room is divided into multiple acquisition areas, each equipped with an independent camera and microphone. A cascaded router structure and audio processor are used to select the sound-emitting area through preset rules to ensure consistency between sound and image. An immersive experience is achieved by adjusting the volume of the speaker equipment.
It achieves accurate sound and image matching in complex environments, improves real-time performance and system scalability, provides a personalized audio and video experience, and adapts to meeting scenarios of different sizes.
Smart Images

Figure CN117998055B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video communication technology, and more specifically, to a method and system for audio-visual synchronization. Background Technology
[0002] Audio-visual synchronization is an audio-visual technology that aims to achieve spatial consistency between sound and image, enabling viewers to experience an immersive experience where the sound and video source are consistent. It is commonly used in scenarios such as video conferencing, multimedia presentations, and remote training to improve the user's perceptual consistency and interactive experience.
[0003] In audio-visual synchronization technology, the system uses intelligent audio and video processing and analysis to associate sound with images, ensuring that sound from a specific area originates from the corresponding video feed. For example, in a video conference, if someone speaks in a certain area of the screen, audio-visual synchronization technology ensures that the person's voice is played through the appropriate speaker, rather than being generated in another area.
[0004] However, existing technologies still have some limitations in the field of audio-visual synchronization:
[0005] Current audio-visual synchronization technology still has limitations in achieving highly accurate sound and image matching. Especially in complex environments, such as multi-person conferences or noisy background noise, accurately tracing the sound source and related video can become even more difficult.
[0006] For applications with high real-time requirements, such as video conferencing, existing technologies may face challenges in real-time audio-video synchronization and spatial consistency. Processing latency can lead to asynchrony between sound and image, degrading the user experience.
[0007] In addition, existing audio-visual synchronization systems are limited by hardware settings, resulting in poor scalability in large conference rooms and making them difficult to adapt to scenarios of different sizes. Summary of the Invention
[0008] This invention provides a method and system for audio-visual synchronization. By dividing the video conference room into multiple acquisition areas, each with an independent camera and microphone, more accurate sound and image matching is achieved. The remote equipment includes at least two display areas and at least two speaker devices. By adjusting the speaker volume, audio and video are synchronized, resulting in a more personalized and immersive audio-visual experience. The cascaded router structure effectively improves the system's scalability, helping to adapt to scenarios of different sizes, especially in large conference rooms, overcoming the limitations of hardware settings in existing technologies. By setting up the connection structure between the router and the microphone, as well as the audio processor, each sound-emitting area is ensured to independently process audio and video signals, reducing the impact of complex environments on system accuracy. Through multi-channel processing by the audio processor, the system can process audio signals more effectively, improving real-time performance and reducing processing latency. According to preset rules, the system can adaptively select at least one acquisition area with participant audio output as the sound-emitting area, ensuring that the sound originates from an active area and improving the audio-visual synchronization effect.
[0009] In a first aspect, the present invention provides a method for audio-visual co-location, characterized in that the method comprises:
[0010] Divide the video conference room into at least two acquisition areas;
[0011] Each camera captures video of its corresponding capture area, so that each capture area corresponds to one video stream, resulting in multiple video streams.
[0012] At least one row of seats should be provided for participants in each data collection area;
[0013] Each microphone captures the audio of the participant in their corresponding seat;
[0014] The number of routers per row is set according to the number of microphones per row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas of the row in which the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor.
[0015] Select at least one collection area where the participants’ voices are being output as the sound area according to preset rules;
[0016] The audio processor is used to process the audio of all participants in each of the vocal areas, so that each vocal area corresponds to a processed audio stream, thus obtaining multiple audio streams.
[0017] Send multiple video and multiple audio streams to the remote location.
[0018] Secondly, the present invention also provides an audio-visual synchronization system, characterized in that the system comprises: a region division device, at least two cameras, at least two microphones, a seating arrangement device, a routing arrangement device, a selection device, an audio processor, a transmitting device, and a remote end; wherein
[0019] The area division device is used to divide the video conference room into at least two acquisition areas;
[0020] Each camera is used to capture video of a corresponding capture area, so that each capture area corresponds to one video stream, thus obtaining multiple video streams.
[0021] The seating arrangement device is used to set up at least one row of seats for participants in each collection area;
[0022] Each microphone is used to collect audio from the participant at the corresponding seat;
[0023] The routing setting device is used to set the number of routers in each row according to the number of microphones in each row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas in the row where the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor.
[0024] The selection device is used to select at least one collection area where the sound output of a participant exists as the sound output area according to a preset rule.
[0025] The audio processor is used to process the audio of all participants in each of the vocal areas separately, so that each vocal area corresponds to one processed audio, thus obtaining multiple audio streams.
[0026] The transmitting device is used to transmit multiple video and multiple audio streams to a remote location.
[0027] The audio-visual synchronization method and system provided by this invention: First, by dividing the video conference room into multiple acquisition areas, each equipped with an independent camera and microphone, precise matching of sound and image is achieved. In the remote device, at least two display areas and two speaker devices are configured. By adjusting the speaker volume, the sound source is made consistent with the image, providing a personalized and immersive audio-visual experience. Second, the hardware setup using a cascaded router structure helps to flexibly adapt to meeting scenarios of different sizes. Users can easily add or remove acquisition areas without modifying the existing connection structure. Third, by setting up the connection structure between the router and the microphone, as well as the audio processor, it is ensured that each sound-emitting area independently processes audio and video signals, reducing the impact of complex environments on system accuracy. Fourth, according to preset rules, the system can adaptively select at least one acquisition area with participant audio output as the sound-emitting area, ensuring that the sound originates from an active area and improving the audio-visual synchronization effect. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart of the audio-visual co-location method provided in the embodiments of the present invention;
[0030] Figure 2 This is a block diagram of the audio-visual co-location system provided in an embodiment of the present invention;
[0031] Figure 3 This is a schematic diagram of the video conference room acquisition area provided in an embodiment of the present invention;
[0032] Figure 4 This is a rendering of a remote video display device provided in an embodiment of the present invention. Detailed Implementation
[0033] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Invention Overview
[0035] As mentioned above, the present invention provides a method and system for audio-visual synchronization, which effectively processes multi-channel audio and multi-screen video, ensuring that each sound source is consistent with the corresponding video, and solving the limitations of the prior art in complex scenarios.
[0036] Exemplary methods
[0037] Figure 1 This is a flowchart of the acoustic-image co-location method provided in an embodiment of the present invention, which includes the following steps:
[0038] S101: Divide the video conference room into at least two acquisition areas.
[0039] S102: Each camera captures video of its corresponding capture area, so that each capture area corresponds to one video stream, thus obtaining multiple video streams.
[0040] S103: Set up at least one row of seats for participants in each collection area.
[0041] For example, such as Figure 3 As shown, the video conference room is divided into three acquisition areas 1-3. Camera 1 acquires video from area 1, camera 2 acquires video from area 2, and camera 3 acquires video from area 3, obtaining three independent video signals. Two rows of seats for participants are set up in each acquisition area.
[0042] S104: Each microphone captures the audio of the attendee at their corresponding seat.
[0043] Each microphone can correspond to one seat, or it can correspond to multiple seats in the same collection area.
[0044] In summary, by dividing the video conference room into multiple acquisition areas, each with its own camera and microphone, preparations are made for achieving accurate sound and image matching.
[0045] S105: Set the number of routers per row according to the number of microphones per row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas of the row in which the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor.
[0046] Configure a certain number of routers for each row of microphones, ensuring that each router has a sufficient number of interfaces, not less than the corresponding number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas within the same row, and at least one interface of different routers is connected to different microphones within the same acquisition area.
[0047] Preferably, the router is an MX204. MX204 routers are typically used to handle large volumes of network traffic and connections, offering high performance and scalability. They support various network protocols and functions, including routing, switching, security, and Quality of Service (QoS), to meet complex network requirements.
[0048] For example, such as Figure 3As shown, there are 3 microphones in the first row, and 1 router MX204 is set up. There are 6 microphones in the second row, and 2 routers MX204 are set up. Interface number 1 of each router MX204 is responsible for area 1, interface number 2 of each router MX204 is responsible for area 2, and interface number 3 of each router MX204 is responsible for area 3. Interfaces with the same number on different routers are responsible for the same collection area.
[0049] It should be noted that at least one interface of different routers can be connected to different microphones in the same acquisition area, and the same numbered interface corresponding to the same acquisition area is only one implementation method.
[0050] Each of these routers forms a cascade, allowing them to work collaboratively and easily integrate microphone signals from the same acquisition area. Furthermore, router cascading enhances system scalability and flexibility, enabling the system to easily add or remove acquisition areas by increasing the number of routers without modifying the existing connection structure.
[0051] For example, such as Figure 3 As shown, if a third row of seats needs to be added to this video conference room, then nine microphones corresponding to the nine seats in the third row (three microphones in each acquisition area) and three routers connected to the nine microphones are added. Interface number 1 on the three routers is responsible for area 1, interface number 2 for area 2, and interface number 3 for area 3. In short, no modifications are needed to the existing router connection structure in the first two rows; only the corresponding hardware deployment is required on the newly added third row, achieving flexible expansion of the acquisition area while maintaining the stability of the existing structure.
[0052] S106: Select at least one acquisition area with the sound output of a participant as the sound output area according to the preset rules.
[0053] This step determines which acquisition area's audio will be processed and transmitted.
[0054] The preset rules can be to select the collection area with the largest human voice volume as the sound generation area, or to select the collection area with a human voice volume greater than a certain set value as the sound generation area, or to select all collection areas where there is sound output from the participants as the sound generation area.
[0055] By dynamically selecting the speaking area, the voices of participants who actively speak can be prioritized, thereby improving the quality of the meeting and the user experience. At the same time, by using rules to restrict the selection of areas with low volume or irrelevant information as speaking areas, noise and unnecessary interference can be reduced.
[0056] S107: The audio processor is used to process the audio of all participants in each of the sound areas, so that each sound area corresponds to one processed audio, thus obtaining multiple audio streams.
[0057] Specifically, to ensure that the audio from each sound-emitting area can be processed independently, the router interface number corresponding to the sound-emitting area is first determined to obtain the sound-emitting interface number; then, the audio processor processes the audio of the participants from different routers corresponding to the sound-emitting interface numbers, so that each sound-emitting area corresponds to one processed audio, thus obtaining multiple audio streams.
[0058] For example, such as Figure 3 As shown, the three acquisition areas 1-3 correspond to different interface numbers of the MX204 router. Specifically, the three microphones CH1 in area 1 connect to interface number 1 (orange dashed line) of the three MX204 routers; the three microphones CH2 in area 2 connect to interface number 2 (blue dashed line); and the three microphones CH3 in area 3 connect to interface number 3 (green dashed line). Assuming participants are simultaneously speaking in all three acquisition areas, acquisition areas 1-3 with participant audio output are selected as the sound output areas according to preset rules. The audio processor obtains and processes the participant audio from area 1 via interface 1 of the three MX204 routers; it also obtains and processes the participant audio from area 2 via interface 2 of the three MX204 routers; and it obtains and processes the participant audio from area 3 via interface 3 of the three MX204 routers. After processing, the audio processor outputs three independent audio signals, corresponding to the processed audio from areas 1-3 respectively. The audio from different areas is processed synchronously to ensure that the processed audio is synchronized in time.
[0059] To optimize audio quality and merge multiple participants' audio into a single audio stream, specific audio processing methods employed include, but are not limited to, noise reduction, equalization, and mixing.
[0060] In summary, the multi-channel processing of the audio processor improves real-time performance and reduces processing latency. The audio processor independently processes multiple audio and video signals from each sound-producing area into one audio stream, which can more effectively reduce the mutual interference of audio information from multiple acquisition areas in complex environments, ensure the consistency and clarity of the sound, and achieve the orderly integration of audio signals from various areas into one audio stream.
[0061] S108: Send multiple video and multiple audio streams to the remote end.
[0062] The remote end includes a video display device and at least two speaker devices.
[0063] The video display device includes at least two display areas. One of the display areas can be a separate screen (e.g., ...). Figure 4 As shown, the remote video display device includes three display areas (each of which is a screen), or it can be one area on a screen (for example, the two windows on the screen of a regular video terminal, such as a laptop, are two display areas).
[0064] A sound amplification device is a device used to amplify and play sound, including but not limited to loudspeakers, horns, speakers, and headphones.
[0065] The display area and the speaker may or may not have a corresponding relationship.
[0066] When there is no corresponding relationship between the two, the system determines whether each display area of the remote video display device plays one of the received video streams according to the user's needs. If a certain display area plays one of the video streams and the acquisition area is a sound-emitting area, the system adjusts the volume of at least two speaker devices according to a certain ratio so that the source of the playback sound is the display area.
[0067] For example, such as Figure 4 As shown, the three screens at the far end are three display areas, each corresponding to displaying the three video feeds of the participants from the sending end. The participant is speaking on the right screen. Assuming there is no direct correspondence between the far-end display areas and the speaker equipment, the volumes of the left and right speakers at the far end are adjusted to 10% and 80% respectively to ensure that the sound in the right display area is more prominent and clear, while keeping the volume of the left display area relatively low so as not to interfere with the auditory experience.
[0068] When there is a correspondence between the two, that is, each display area corresponds to at least one speaker device, the system determines whether each display area of the remote video display device plays one of the received video streams according to the user's needs. If a certain display area plays one of the video streams and that area is a sound-emitting area, then the speaker device corresponding to that display area plays the audio of that sound-emitting area.
[0069] For example, such as Figure 4 As shown, the three screens at the remote end are three display areas, each corresponding to displaying the three video feeds of the participants from the sending end. The participant is speaking on the right screen. Assuming that each screen at the remote end has a speaker corresponding to it, the speaker at the bottom of the right screen will play the participant's audio.
[0070] In summary, the remote device includes at least two display areas and at least two speaker devices. By adjusting the volume of the speaker devices or playing corresponding audio through the corresponding speaker devices, the source of the played sound and video is ensured to be consistent, providing users with a more consistent and immersive audio and video experience.
[0071] In addition, for situations where the speaker is not on the screen or when sharing audio content, the sound can be adjusted to be emitted from the center screen, or all speakers can output the audio at the same volume.
[0072] Exemplary System
[0073] Accordingly, embodiments of the present invention also provide an audio-visual synchronization system. Figure 2 This is a block diagram of the audio-visual co-location system provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the system 100 provided in this embodiment includes:
[0074] The device includes a zone division device 101, at least two cameras 102, at least two microphones 103, a seating arrangement device 104, a routing arrangement device 105, a selection device 106, an audio processor 107, a transmitting device 108, and a remote device 109; wherein...
[0075] The area division device 101 is used to divide the video conference room into at least two acquisition areas;
[0076] Each of the cameras 102 is used to capture video of a corresponding capture area, so that each capture area corresponds to one video stream, thereby obtaining multiple video streams.
[0077] The seating arrangement device 104 is used to set at least one row of seats for participants in each collection area;
[0078] Each of the microphones 103 is used to collect the audio of the participants in the corresponding seats;
[0079] The routing setting device 105 is used to set the number of routers in each row according to the number of microphones 103 in each row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones 103 in different acquisition areas in the row where the router is located, so as to route the audio of the participants collected by the microphones 103 in different acquisition areas to the audio processor.
[0080] The selection device 106 is used to select at least one collection area where the sound output of a participant exists as the sound output area according to a preset rule.
[0081] The audio processor 107 is used to process the audio of all participants in each of the sound areas separately, so that each sound area corresponds to one processed audio, thereby obtaining multiple audio streams.
[0082] The transmitting device 108 is used to transmit multiple video and multiple audio signals to a remote end 109.
[0083] The remote end 109 includes a video display device and at least two speaker devices;
[0084] The video display device includes at least two display areas.
[0085] The remote end 109 also includes a determining unit 110 and an adjusting unit 111;
[0086] The determining unit 110 is used to determine, according to user needs, whether each display area of the video display device at the remote end 109 should play one of the received video streams:
[0087] If a certain display area plays one of the video streams and the acquisition area is a sound-emitting area, the adjustment unit 111 is used to adjust the volume of at least two speaker devices according to a certain ratio so that the source of the playback sound is the display area.
[0088] If each display area corresponds to at least one speaker device, the remote end 109 further includes a determination unit 110 and a playback unit 112;
[0089] The determining unit 110 is used to determine, according to user needs, whether each display area of the video display device at the remote end 109 should play one of the received video streams:
[0090] If a certain display area is playing one of the video streams and that area is also a sound-emitting area, then the playback unit 112 is used to select the speaker device corresponding to that display area to play the audio of that sound-emitting area.
[0091] A display area is a screen or a region on a screen.
[0092] Each of the routers forms a cascade.
[0093] At least one interface of the different routers is connected to different microphones 103 in the same acquisition area.
[0094] The preset rules include:
[0095] Select the area with the highest volume of human voice as the vocal area;
[0096] Select the area where the human voice volume is greater than a certain set value as the sound emission area; or
[0097] Select all areas where the participants' voices are being captured as the sound output area.
[0098] The audio processor 107 also includes:
[0099] A unit for determining the router interface number corresponding to the sound-emitting area to obtain the sound-emitting interface number;
[0100] This unit is used to process the audio of participants from different routers with corresponding audio interface numbers, so that each audio area corresponds to one processed audio stream, thus obtaining multiple audio streams.
[0101] It should be noted that although the operation of the audio-visual synchronization method of the present invention is described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0102] Furthermore, although several devices, units, or modules of the audio-visual synchronization system have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in a single module. Conversely, the features and functions of a single module described above can be further divided and embodied by multiple modules.
[0103] While the spirit and principles of the invention have been described with reference to several specific embodiments, it should be understood that the invention is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for ease of description. The invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
[0104] This invention provides:
[0105] 1. A method for audio-visual co-location, characterized in that the method comprises:
[0106] Divide the video conference room into at least two acquisition areas;
[0107] Each camera captures video of its corresponding capture area, so that each capture area corresponds to one video stream, resulting in multiple video streams.
[0108] At least one row of seats should be provided for participants in each data collection area;
[0109] Each microphone captures the audio of the participant in their corresponding seat;
[0110] The number of routers per row is set according to the number of microphones per row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas of the row in which the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor.
[0111] Select at least one collection area where the participants’ voices are being output as the sound area according to preset rules;
[0112] The audio processor is used to process the audio of all participants in each of the vocal areas, so that each vocal area corresponds to a processed audio stream, thus obtaining multiple audio streams.
[0113] Send multiple video and multiple audio streams to the remote location.
[0114] 2. The audio-visual synchronization method according to item 1, characterized in that the remote end includes a video display device and at least two speaker devices;
[0115] The video display device includes at least two display areas.
[0116] 3. The audio-visual synchronization method according to item 2, characterized in that the method further includes: determining whether each display area of the remote video display device plays one of the received video streams according to user requirements:
[0117] If a certain display area is playing one of the video streams and that acquisition area is a sound-generating area, then the volume of at least two speaker devices is adjusted according to a certain ratio so that the source of the playback sound is that display area.
[0118] 4. The audio-visual synchronization method according to item 2, characterized in that each display area corresponds to at least one speaker device, and the method further includes: determining whether each display area of the remote video display device plays one of the received video streams according to user requirements.
[0119] If a certain display area is playing one of the video streams and that area is also a sound-emitting area, then the speaker device corresponding to that display area will play the audio from that sound-emitting area.
[0120] 5. The audio-visual synchronization method according to any one of items 2-4, characterized in that one of the display areas is a screen or a region on a screen.
[0121] 6. The audio-visual co-location method according to any one of items 1-4, characterized in that each of the routers is cascaded.
[0122] 7. The audio-visual synchronization method according to item 6, characterized in that at least one interface of different routers is connected to different microphones in the same acquisition area.
[0123] 8. The method for audio-visual co-location according to any one of items 1-4, characterized in that the preset rules include:
[0124] Select the area with the highest volume of human voice as the vocal area;
[0125] Select the area where the human voice volume is greater than a certain set value as the sound emission area; or
[0126] Select all areas where the participants' voices are being captured as the sound output area.
[0127] 9. The method of audio-visual synchronization according to item 7, characterized in that, the step of using the audio processor to process the audio of all participants in each of the vocal regions separately, so that each vocal region corresponds to one processed audio path, and obtaining multiple audio paths, specifically comprises:
[0128] Determine the router interface number corresponding to the sound-emitting area to obtain the sound-emitting interface number;
[0129] The audio processor processes the audio of the participants from different routers according to their corresponding sound interface numbers, so that each sound area corresponds to one processed audio stream, thus obtaining multiple audio streams.
[0130] 10. A simultaneous audio-visual system, characterized in that the system comprises: a region division device, at least two cameras, at least two microphones, a seating arrangement device, a routing arrangement device, a selection device, an audio processor, a transmitting device, and a remote end; wherein
[0131] The area division device is used to divide the video conference room into at least two acquisition areas;
[0132] Each camera is used to capture video of a corresponding capture area, so that each capture area corresponds to one video stream, thus obtaining multiple video streams.
[0133] The seating arrangement device is used to set up at least one row of seats for participants in each collection area;
[0134] Each microphone is used to collect audio from the participant at the corresponding seat;
[0135] The routing setting device is used to set the number of routers in each row according to the number of microphones in each row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas in the row where the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor.
[0136] The selection device is used to select at least one collection area where the sound output of a participant exists as the sound output area according to a preset rule.
[0137] The audio processor is used to process the audio of all participants in each of the vocal areas separately, so that each vocal area corresponds to one processed audio, thus obtaining multiple audio streams.
[0138] The transmitting device is used to transmit multiple video and multiple audio streams to a remote location.
[0139] 11. The audio-visual co-location system according to item 10, characterized in that the remote end includes a video display device and at least two speaker devices;
[0140] The video display device includes at least two display areas.
[0141] 12. The audio-visual co-location system according to item 11, characterized in that the distal end further includes a determining unit and an adjusting unit;
[0142] The determining unit is used to determine, based on user requirements, whether each display area of the remote video display device should play one of the received video streams:
[0143] If a certain display area is playing one of the video streams and the acquisition area is a sound-emitting area, the adjustment unit is used to adjust the volume of at least two speaker devices according to a certain ratio so that the source of the playback sound is the display area.
[0144] 13. The audio-visual co-location system according to item 11, characterized in that, if each display area corresponds to at least one speaker device, the far end further includes a determining unit and a playback unit;
[0145] The determining unit is used to determine, based on user requirements, whether each display area of the remote video display device should play one of the received video streams:
[0146] If a certain display area is playing one of the video streams and that area is also a sound-emitting area, the playback unit is used to select the speaker device corresponding to that display area to play the audio from that sound-emitting area.
[0147] 14. The audio-visual synchronization system according to any one of claims 11-13, characterized in that one of the display areas is a screen or an area on a screen.
[0148] 15. The audio-visual co-location system according to any one of claims 10-13, characterized in that each of the routers is cascaded.
[0149] 16. The audio-visual synchronization system according to claim 15, characterized in that at least one interface of each of the different routers is connected to different microphones in the same acquisition area.
[0150] 17. The audio-visual co-location system according to any one of items 10-13, characterized in that the preset rules include:
[0151] Select the area with the highest volume of human voice as the vocal area;
[0152] Select the area where the human voice volume is greater than a certain set value as the sound emission area; or
[0153] Select all areas where the participants' voices are being captured as the sound output area.
[0154] 18. The audio-visual synchronization system according to claim 16, characterized in that the audio processor further comprises:
[0155] A unit for determining the router interface number corresponding to the sound-emitting area to obtain the sound-emitting interface number;
[0156] This unit is used to process the audio of participants from different routers with corresponding audio interface numbers, so that each audio area corresponds to one processed audio stream, thus obtaining multiple audio streams.
Claims
1. A method for locating audio-visual images, characterized in that, The method includes: Divide the video conference room into at least two acquisition areas; Each camera captures video of its corresponding capture area, so that each capture area corresponds to one video stream, resulting in multiple video streams. At least one row of seats should be provided for participants in each data collection area; Each microphone captures the audio of the participant in their corresponding seat; The number of routers per row is set according to the number of microphones per row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas of the row in which the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor. Select at least one collection area where the participants’ voices are being output as the sound area according to preset rules; The audio processor is used to process the audio of all participants in each of the vocal areas, so that each vocal area corresponds to a processed audio stream, thus obtaining multiple audio streams. Send multiple video and multiple audio streams to the remote location.
2. The method for audio-visual co-location according to claim 1, characterized in that, The remote end includes a video display device and at least two speaker devices; The video display device includes at least two display areas.
3. The method for audio-visual co-location according to claim 2, characterized in that, The method further includes: determining, based on user needs, whether each display area of the remote video display device should play one of the received video streams. If a certain display area is playing one of the video streams and that acquisition area is a sound-generating area, then the volume of at least two speaker devices is adjusted according to a certain ratio so that the source of the playback sound is that display area.
4. The method for audio-visual co-location according to claim 2, characterized in that, Each display area corresponds to at least one speaker device, and the method further includes: determining whether each display area of the remote video display device plays one of the received video streams based on user requirements. If a certain display area is playing one of the video streams and that area is also a sound-emitting area, then the speaker device corresponding to that display area will play the audio from that sound-emitting area.
5. The method for audio-visual co-location according to any one of claims 2-4, characterized in that, A display area is a screen or a region on a screen.
6. The method for audio-visual co-location according to any one of claims 1-4, characterized in that, Each of the routers forms a cascade.
7. The method for audio-visual co-location according to claim 6, characterized in that, At least one interface of each router is connected to different microphones in the same acquisition area.
8. The method for audio-visual co-location according to any one of claims 1-4, characterized in that, The preset rules include: Select the area with the highest volume of human voice as the vocal area; Select the area where the human voice volume is greater than a certain set value as the sound emission area; or Select all areas where the participants' voices are being captured as the sound output area.
9. The method for audio-visual co-location according to claim 7, characterized in that, The audio processor is used to process the audio of all participants in each of the vocal regions separately, so that each vocal region corresponds to one processed audio stream. The specific steps to obtain multiple audio streams are as follows: Determine the router interface number corresponding to the sound-emitting area to obtain the sound-emitting interface number; The audio processor processes the audio of the participants from different routers according to their corresponding sound interface numbers, so that each sound area corresponds to one processed audio stream, thus obtaining multiple audio streams.
10. A sound-image synchronization system, characterized in that, The system includes: a region division device, at least two cameras, at least two microphones, a seating arrangement device, a routing arrangement device, a selection device, an audio processor, a transmitting device, and a remote end; wherein... The area division device is used to divide the video conference room into at least two acquisition areas; Each camera is used to capture video of a corresponding capture area, so that each capture area corresponds to one video stream, thus obtaining multiple video streams. The seating arrangement device is used to set up at least one row of seats for participants in each collection area; Each microphone is used to collect audio from the participant at the corresponding seat; The routing setting device is used to set the number of routers in each row according to the number of microphones in each row. The number of interfaces of each router is not less than the number of acquisition areas. Different interfaces of each router are connected to microphones in different acquisition areas in the row where the router is located, so as to route the audio of the participants collected by the microphones in different acquisition areas to the audio processor. The selection device is used to select at least one collection area where the sound output of a participant exists as the sound output area according to a preset rule. The audio processor is used to process the audio of all participants in each of the vocal areas separately, so that each vocal area corresponds to one processed audio, thus obtaining multiple audio streams. The transmitting device is used to transmit multiple video and multiple audio streams to a remote location.
11. The audio-visual synchronization system according to claim 10, characterized in that, The remote end includes a video display device and at least two speaker devices; The video display device includes at least two display areas.
12. The audio-visual co-location system according to claim 11, characterized in that, The remote end also includes a determining unit and an adjusting unit; The determining unit is used to determine, based on user requirements, whether each display area of the remote video display device should play one of the received video streams: If a certain display area is playing one of the video streams and the acquisition area is a sound-emitting area, the adjustment unit is used to adjust the volume of at least two speaker devices according to a certain ratio so that the source of the playback sound is the display area.
13. The audio-visual synchronization system according to claim 11, characterized in that, If each display area corresponds to at least one speaker device, the remote end also includes a determination unit and a playback unit; The determining unit is used to determine, based on user requirements, whether each display area of the remote video display device should play one of the received video streams: If a certain display area is playing one of the video streams and that area is also a sound-emitting area, the playback unit is used to select the speaker device corresponding to that display area to play the audio from that sound-emitting area.
14. The audio-visual synchronization system according to any one of claims 11-13, characterized in that, A display area is a screen or a region on a screen.
15. The audio-visual synchronization system according to any one of claims 10-13, characterized in that, Each of the routers forms a cascade.
16. The audio-visual synchronization system according to claim 15, characterized in that, At least one interface of each router is connected to different microphones in the same acquisition area.
17. The audio-visual synchronization system according to any one of claims 10-13, characterized in that, The preset rules include: Select the area with the highest volume of human voice as the vocal area; Select the area where the human voice volume is greater than a certain set value as the sound emission area; or Select all areas where the participants' voices are being captured as the sound output area.
18. The audio-visual synchronization system according to claim 16, characterized in that, The audio processor also includes: A unit for determining the router interface number corresponding to the sound-emitting area to obtain the sound-emitting interface number; This unit is used to process the audio of participants from different routers with corresponding audio interface numbers, so that each audio area corresponds to one processed audio stream, thus obtaining multiple audio streams.
Citation Information
Patent Citations
Video call equipment and audio gain method
CN112423191A
Dispatching command system
CN215420342U