Conference system and method, electronic equipment and storage medium
By obtaining and sharing audio loudness data, the problems of echo and speaker recognition in the same space meeting of multiple devices are solved, and efficient meeting collaboration and accurate meeting records are achieved.
Patent Information
- Application Number
- CN202510453358.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-24
AI Technical Summary
When multiple devices use the same physical space for meetings at the same time, it is easy to have problems such as echoes and difficult to identify the speaker's identity.
The terminal device attribute acquisition module obtains the space where each device is located, the audio attribute acquisition module obtains the audio loudness data, and the audio sharing control module selects the audio stream with the highest audio loudness in the same space to share it with the device participating in the conference, and records the audio data and screen sharing content in real time.
It realizes avoiding echoes in the same space conference scenario of multiple devices and accurately identifying the speakers, improving the collaboration efficiency of the meeting and the accuracy of recording.
Smart Images

Figure CN120201155A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electronic communication, and relates to a conference system, and particularly to a multi-device conference system, method, electronic device, and storage medium. Background Art
[0002] With the widespread application of online conferences, improving the quality and efficiency of conferences has received increasing attention. When multiple people in the same physical room are talking together, only one device can turn on the MIC and speaker. Otherwise, a lot of echo will be heard, which brings a lot of inconvenience; at the same time, it brings the big problem of difficult to identify who is speaking.
[0003] In view of this, there is an urgent need to design a new conference system today to overcome the above-mentioned defects existing in the existing conference systems. Summary of the Invention
[0004] The present invention provides a conference system, method, electronic device, and storage medium, which can be applicable to the scenario of multiple devices in the same space. Multiple users in the same space can select their own terminal devices without generating echo; and the conference process can be recorded.
[0005] To solve the above technical problems, according to one aspect of the present invention, the following technical solution is adopted:
[0006] A conference system, the conference system includes:
[0007] A terminal device attribute acquisition module for acquiring the set attribute information of each terminal device, where the attribute information includes the space where the terminal device is located;
[0008] An audio attribute acquisition module for acquiring the attribute information of each terminal device for acquiring audio, where the attribute information of the audio includes audio loudness data; and
[0009] An audio sharing control module for selecting at least one audio stream with the highest audio loudness data in the same space and sharing it to the set terminal devices participating in the conference according to the audio loudness data of each terminal device acquired by the audio attribute acquisition module; if the shared audio is acquired through terminal device A and terminal device B is in the set same space as terminal device A, then the audio is not shared to terminal device B.
[0010] As an implementation manner of the present invention, the conference system further includes:
[0011] A shared audio data recording module for real-time recording of the audio data shared by the audio sharing control module to each terminal device participating in the conference;
[0012] A screen sharing control module for sharing the content displayed in a set screen area of a first terminal device with at least one second terminal device;
[0013] A shared display content recording module for real-time recording of the content shared and displayed by the screen sharing control module; and
[0014] A dynamic operation recording module for obtaining dynamic operations on the shared display content and recording the dynamic operations in real time.
[0015] As an embodiment of the present invention, the terminal device attribute acquisition module includes a space declaration unit. Each terminal device sends the physical space where it participates in the meeting to the conference system through the space declaration unit; the terminal device attribute acquisition module obtains the physical space where each terminal device is located according to the information declared by each terminal device.
[0016] As an embodiment of the present invention, the conference system includes a server and at least one terminal device. The server is respectively connected to each terminal device; the server includes the terminal device attribute acquisition module, the audio attribute acquisition module and the audio sharing control module.
[0017] As an embodiment of the present invention, the conference system further includes:
[0018] A device control module for controlling the operation of the device;
[0019] A voice recognition module for recognizing voice information.
[0020] According to another aspect of the present invention, the following technical solution is adopted: A conference method, the conference method includes:
[0021] A terminal device attribute acquisition step: acquiring the set attribute information of each terminal device, where the attribute information includes the space where the terminal device is located;
[0022] An audio attribute acquisition step: acquiring the attribute information of the audio obtained by each terminal device, where the attribute information of the audio includes audio loudness data; and
[0023] An audio sharing control step: according to the audio loudness data of the audio obtained by each terminal device acquired in the audio attribute acquisition step, selecting at least one audio stream with the highest audio loudness data in the same space and sharing it to the set terminal devices participating in the meeting; if the shared audio is obtained through terminal device A and terminal device B is in the same set space as terminal device A, then this audio is not shared to terminal device B.
[0024] As an embodiment of the present invention, the conference method further includes:
[0025] Steps for sharing audio data recording: Record in real time the audio data shared in the audio sharing control steps and shared to each terminal device participating in the meeting;
[0026] Steps for screen sharing control: Share the content displayed in the set screen area of the first terminal device to at least one second terminal device;
[0027] Steps for sharing display content recording: Record in real time the content shared by the screen sharing control module; and
[0028] Steps for dynamic operation recording: Obtain dynamic operations on the shared display content and record the dynamic operations in real time.
[0029] As an embodiment of the present invention, in the terminal device attribute acquisition steps, each terminal device sends the physical space where it participates in the meeting to the conference system; the physical space where each terminal device is located is obtained according to the information declared by each terminal device.
[0030] According to another aspect of the present invention, the following technical solution is adopted: An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0031] According to another aspect of the present invention, the following technical solution is adopted: A storage medium stores computer program instructions thereon. When the computer program instructions are executed by a processor, the steps of the above method are implemented.
[0032] The beneficial effects of the present invention are as follows: The conference system, method, electronic device, and storage medium proposed by the present invention can be applied to scenarios where multiple devices are in the same space. Multiple users in the same space can select their own terminal devices without generating echoes; and the conference process can be recorded.
[0033] The present invention clearly tells the server which device the sound comes from when the sound is transmitted to the audio server; at the same time, each device is tagged to represent which physical room each participant comes from, so as to solve the problem of echoes generated by multiple devices in the same room.
[0034] The present invention has the following beneficial effects: (1) Simplify the participation process: Support free joining of multiple devices without complex settings; (2) Improve collaboration efficiency: Accurate speaker detection and task allocation functions; (3) Seamless integration experience: Bridge the gap between online and face-to-face meetings; (4) Wide applicability: Applicable to various scenarios such as enterprise meetings, education and training, and remote collaboration. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the composition of the conference system in an embodiment of the present invention.
[0036] Figure 2 This is a flowchart of the meeting method in an embodiment of the present invention.
[0037] Figure 3 This is a schematic diagram of the composition of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0038] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0039] In order to further understand the present invention, the preferred implementation manners of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.
[0040] The description of this part only focuses on several typical embodiments, and the present invention is not limited to the scope described in the embodiments. The mutual replacement of the same or similar prior art means and some technical features in the embodiments is also within the scope of the description and protection of the present invention.
[0041] The expression of the steps in each embodiment in the specification is only for convenience of description, and the implementation manner of the present application is not limited by the order of step implementation.
[0042] "Connection" in the specification includes both direct connection and indirect connection.
[0043] The present invention discloses a meeting system. Figure 1 This is a schematic diagram of the composition of the meeting system in an embodiment of the present invention; please refer to Figure 1 , the meeting system includes: a terminal device attribute acquisition module 1, an audio attribute acquisition module 2, and an audio sharing control module 3.
[0044] The terminal device attribute acquisition module 1 is used to acquire the set attribute information of each terminal device, and the attribute information includes the space where the terminal device is located.
[0045] In an embodiment of the present invention, the terminal device attribute acquisition module 1 includes a space declaration unit, and each terminal device sends the physical space where it participates in the meeting to the meeting system through the space declaration unit; the terminal device attribute acquisition module acquires the physical space where each terminal device is located according to the information declared by each terminal device. For example, the system can provide a graphical user interface (GUI) for participants to independently declare their physical room identifiers (such as "Meeting Room 401").
[0046] In addition, it is also possible to obtain whether the terminal devices are in the same physical space through an automatic recognition method. The method may include: first, judging whether they are in the same large physical area according to the network IPs of the terminal devices; if not in the same large area, it is judged that the terminal devices are not in the same physical space. For terminal devices with the same network IP, it can be preliminarily judged that they are in the same large physical area; then, according to whether the audio data obtained by each terminal device is similar, it is judged whether they are in the same physical space (such as the same room).
[0047] The audio attribute acquisition module 2 is used to acquire the attribute information of the audio obtained by each terminal device, and the attribute information of the audio includes audio loudness data.
[0048] The audio sharing control module 3 is used to select at least one audio stream with the highest audio loudness data in the same space and share it with the set terminal devices participating in the meeting according to the audio loudness data of the audio obtained by each terminal device acquired by the audio attribute acquisition module; if the shared audio is obtained through terminal device A and terminal device B is in the set same space as terminal device A, then this audio is not shared with terminal device B.
[0049] In a usage scenario of the present invention, the conference system includes a server, and the server dynamically selects the audio stream with the highest volume in the devices in the same physical room and transmits it to the participants in other rooms.
[0050] In an embodiment of the present invention, the conference system further includes: a shared audio data recording module 4, a screen sharing control module 5, a shared display content recording module 6, and a dynamic operation recording module 7.
[0051] The shared audio data recording module 4 is used to record in real time the audio data shared by the audio sharing control module to each terminal device participating in the meeting. The screen sharing control module 5 is used to share the content displayed in the set area of the screen of the set first terminal device to at least one second terminal device. The shared display content recording module 6 is used to record in real time the content shared and displayed by the screen sharing control module. The dynamic operation recording module 7 is used to obtain the dynamic operations on the shared display content and record the dynamic operations in real time.
[0052] Finally, the conference system can generate a meeting record in real time, accurately mark the identities of the speakers, and is applicable to online and face-to-face meetings.
[0053] In an embodiment of the present invention, the conference system includes a server and at least one terminal device, and the server is respectively connected to each terminal device; the server includes the terminal device attribute acquisition module 1, the audio attribute acquisition module 2, the audio sharing control module 3, the shared audio data recording module 4, the screen sharing control module 5, the shared display content recording module 6, and the dynamic operation recording module 7.
[0054] In an embodiment of the present invention, the conference system further includes: a main control module 8, a device control module 9, and a voice recognition module. The main control module 8 is respectively connected to the terminal device attribute acquisition module 1, the audio attribute acquisition module 2, the audio sharing control module 3, the shared audio data recording module 4, the screen sharing control module 5, the shared display content recording module 6, the dynamic operation recording module 7, the device control module 8, and the voice recognition module 9. The device control module 8 is used to control the operation of the device; the voice recognition module 9 is used to recognize voice information.
[0055] The present invention further discloses a conference method. Figure 2 It is a flowchart of the conference method in an embodiment of the present invention; please refer to Figure 2 and the conference method includes:
[0056]
Step S1
[0057] In an embodiment of the present invention, in the terminal device attribute acquisition step, each terminal device sends the physical space where it participates in the conference to the conference system; the physical space where each terminal device is located is obtained according to the information reported by each terminal device.
[0058] In addition, it is also possible to obtain whether the terminal devices are in the same physical space by an automatic recognition method; the method may include: first, judge whether they are in the same physical large area according to the network IP of the terminal devices; if they are not in the same large area, it is judged that the terminal devices are not in the same physical space. For terminal devices with the same network IP, it can be initially judged that they are in the same physical large area; then, according to whether the audio data obtained by each terminal device is similar, judge whether they are in the same physical space (such as the same room).
[0059]
Step S2
[0060]
Step S3
[0061] In an embodiment of the present invention, the conference method further includes:
[0062]
Step S4
[0063]
Step S5
[0064]
Step S6
[0065]
Step S7
[0066] The above method is not only limited to being carried out step by step according to the above steps. Some steps can be carried out simultaneously, and some steps can adjust the execution order.
[0067] In a usage scenario of the present invention, the present invention provides a seamless multi-device conferencing method, including the following steps: allowing participants to join the meeting using a computer or mobile device, whether in the same physical room or remotely participating; capturing the voice of each participant through their independent device to ensure 100% accurate speaker detection; dynamically routing the audio to all participants while preventing echo by blocking audio transmission to devices in the same physical room; integrating online and face-to-face meeting functions into a unified system.
[0068] In another usage scenario of the present invention, the present invention discloses a method for enhancing meeting collaboration, including the following steps: allowing participants to share screens or documents from any device without manual coordination; automatically assigning tasks or action items according to the speech content of the identified speakers; providing real-time analysis data, such as speaking time and participation, to improve meeting efficiency.
[0069] In yet another usage scenario of the present invention, the present invention discloses a method for integrating online and face-to-face meetings, including the following steps: capturing the audio from physical room devices and remote participants through a unified server; ensuring seamless communication between in-room and remote participants through accurate speaker detection and echo prevention mechanisms; supporting applications such as real-time captioning, multilingual translation, and voice-activated control to provide a fully integrated meeting experience.
[0070] In yet another usage scenario of the present invention, the present invention discloses a method for compliance and documentation, including the following steps: automatically recording and annotating the speech content of specific participants for legal or regulatory purposes; generating meeting records with speaker identities and timestamps for auditing or dispute resolution.
[0071] In yet another usage scenario of the present invention, the present invention discloses a method for supporting accessibility and inclusivity, comprising the following steps: providing real-time captions or subtitles, accurately labeling the identities of speakers to facilitate the participation of people with hearing impairments; supporting multi-language conferences, automatically translating and annotating the speech content to specific speakers.
[0072] In yet another usage scenario of the present invention, the present invention discloses a method for voice-activated conference control, comprising the following steps: allowing participants to use voice commands to control conference room devices (such as display screens, lights, or projectors); accurately identifying speakers and executing commands according to their permissions or roles.
[0073] In yet another usage scenario of the present invention, the present invention discloses a method for dynamic conference summarization and follow-up, comprising the following steps: automatically generating a conference summary, annotating key points and the identities of speakers; allocating action items or tasks according to the speech content of speakers; sending follow-up emails or notifications containing the summary content and assigned tasks.
[0074] In an embodiment of the present invention, five employees of a company participate in a hybrid conference: three participate in the conference using laptops and mobile phones in "Conference Room 401", and two participate remotely using computers. The server detects the room identifier and dynamically selects the speaker with the highest volume (such as Alice's laptop). Alice's audio is transmitted to remote participants, and the in-room devices do not receive her audio through the system (to prevent echo). At the same time, the system generates real-time captions, tracks the speaking time, and assigns tasks according to the speech content. Remote participants can share their screens and use voice commands to control the in-room display screen, while in-room participants collaborate naturally.
[0075] The present invention also discloses an electronic device, Figure 3 is a schematic diagram of the composition of the electronic device in an embodiment of the present invention; please refer to Figure 3 , at the hardware level, the electronic device includes a memory, a processor, and at least one communication interface; the processor can be a microprocessor, and the memory can include a memory, such as a random access memory (RAM), and can also include a non-volatile memory, etc. Of course, the electronic device can also be provided with other hardware according to needs.
[0076] The processor, communication interface, and memory can be interconnected through an internal bus. The memory is used to store programs (which can include an operating system program and application programs); the programs can include program codes, and the program codes can include computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0077] In one embodiment, the processor may read the corresponding program from the non-volatile memory into the memory and then run it; the processor can execute the program stored in the memory and is specifically used to perform the following operations (as Figure 2 shown):
[0078]
Step S1
[0079]
Step S2
[0080]
Step S3
[0081]
Step S4
[0082]
Step S5
[0083]
Step S6
[0084]
Step S7
[0085] The present invention further discloses a storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the following steps of the method of the present invention are implemented (as Figure 2 shown):
[0086]
Step S1
[0087]
Step S2
[0088]
Step S3
[0089]
Step S4
[0090]
Step S5
[0091]
Step S6
[0092]
Step S7
[0093] In summary, the conference system and method proposed by the present invention can be applied to the scenario of multiple devices in the same space. Multiple users in the same space can select their own terminal devices without generating echoes; and the conference process can be recorded.
[0094] The present invention clearly tells the server which device it comes from when transmitting the sound to the audio server; at the same time, each device is tagged to represent which physical room each participant comes from, so as to solve the problem that echoes will be generated when multiple devices are in the same room.
[0095] The conference system provided by the present invention allows participants to freely use computer or mobile phone devices to join the conference, whether in the same physical room or remotely. The system captures the voice of each participant through an independent device of each participant, ensures 100% accurate speaker detection, and dynamically routes the audio to all participants while preventing echoes. The system integrates online and face-to-face conference functions, supports real-time applications such as conference recording, analysis and task assignment, significantly improves the effectiveness of the conference, simplifies the participation process, improves the collaboration experience, and bridges the gap between physical and virtual conference environments.
[0096] The present invention has the following functions and characteristics:
[0097] Device-independent participation mechanism: Participants can use any device (computer, mobile phone or tablet) to join the conference, whether in the same physical room or remotely.
[0098] Automatic speaker detection: Capturing the voice of each participant through their individual devices to ensure 100% accurate speaker detection.
[0099] Echo prevention mechanism: The server dynamically selects the highest volume audio stream of the devices in the same physical room and transmits it to the participants in other rooms; preventing the selected audio stream from being transmitted back to the devices in the same room to avoid echoes.
[0100] Online and face-to-face meeting integration: The system unifies and integrates the audio and video streams from the devices in the physical room and remote participants into a single meeting environment; supporting applications such as real-time captioning, multilingual translation, and voice-activated control for all participants.
[0101] Real-time application functions, including: ① Meeting recording: Automatically generating and accurately annotating the identity of the speaker; ② Data analysis: Tracking and visualizing metrics such as speaking time, participation frequency, and contribution patterns; ③ Task assignment: Assigning action items based on the speech content of the identified speaker.
[0102] Accessibility and compliance functions: Real-time captioning and multilingual translation ensure inclusivity; meeting records with speaker identity and timestamps support compliance and legal documentation requirements.
[0103] Voice-activated control: Participants can use voice commands to control the meeting room devices (such as displays, lights), and the system ensures safe execution through accurate speaker detection.
[0104] It should be noted that the present application can be implemented in software and / or a combination of software and hardware; for example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium; for example, a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of the present application can be implemented using hardware; for example, as a circuit that cooperates with the processor to execute each step or function.
[0105] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.
[0106] The description and application of the present invention herein are illustrative and are not intended to limit the scope of the present invention to the above embodiments. The effects or advantages involved in the embodiments may not be reflected in the embodiments due to various interferences, and the description of the effects or advantages is not used to limit the embodiments. Modifications and changes to the disclosed embodiments are possible, and various components of substitution and equivalence of the embodiments are known to those of ordinary skill in the art. Those skilled in the art should clearly understand that the present invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the present invention. Other modifications and changes can be made to the disclosed embodiments without departing from the scope and spirit of the present invention.
Claims
1. A conference system, characterized in that: The conference system comprises: A terminal device attribute acquisition module is used to acquire the set attribute information of each terminal device, wherein the attribute information includes the space where the terminal device is located; An audio attribute acquisition module, used to acquire attribute information of audio acquired by each terminal device, wherein the attribute information of the audio includes audio loudness data; and The audio sharing control module is used to obtain audio loudness data of each terminal device obtained by the audio attribute acquisition module, and select at least one audio stream with the highest audio loudness data in the same space to share with the set terminal devices participating in the meeting; if the shared audio is obtained through terminal device A, and terminal device B is in the same set space as terminal device A, the audio will not be shared with terminal device B.
2. The conference system according to claim 1, characterized in that: The conference system further comprises: A shared audio data recording module is used to record in real time the audio data shared by the audio sharing control module to each terminal device participating in the conference; A screen sharing control module, used for sharing the content displayed in the setting area of the screen of the first terminal device with at least one second terminal device; a shared display content recording module, used for recording in real time the shared display content of the screen sharing control module; and The dynamic operation recording module is used to obtain dynamic operations on the shared display content and record the dynamic operations in real time.
3. The conference system according to claim 1, characterized in that: The terminal device attribute acquisition module includes a space reporting unit, through which each terminal device sends the physical space where it participates in the meeting to the conference system; the terminal device attribute acquisition module obtains the physical space where each terminal device is located based on the information reported by each terminal device.
4. The conference system according to claim 1, characterized in that: The conference system includes a server and at least one terminal device, and the server is connected to each terminal device respectively; the server includes the terminal device attribute acquisition module, the audio attribute acquisition module and the audio sharing control module.
5. The conference system according to claim 1, characterized in that: The conference system further comprises: Equipment control module, used to control the operation of equipment; The speech recognition module is used to recognize speech information.
6. A conference method, characterized in that: The conference methods include: Terminal device attribute acquisition step: acquiring setting attribute information of each terminal device, wherein the attribute information includes the space where the terminal device is located; Audio attribute acquisition step: acquiring attribute information of audio acquired by each terminal device, wherein the audio attribute information includes audio loudness data; and Audio sharing control step: obtaining audio loudness data from each terminal device according to the audio attribute acquisition step, selecting at least one audio stream with the highest audio loudness data in the same space to share with the set terminal devices participating in the meeting; if the shared audio is obtained through terminal device A, and terminal device B is in the same set space as terminal device A, the audio will not be shared with terminal device B.
7. The conference method according to claim 6, characterized in that: The conference method further comprises: Shared audio data recording step: recording in real time the audio data shared by the audio sharing control step to each terminal device participating in the conference; Screen sharing control step: sharing the content displayed in the setting area of the screen of the first terminal device with at least one second terminal device; A shared display content recording step: recording the shared display content of the screen sharing control module in real time; and Dynamic operation recording step: obtaining dynamic operations on shared display content, and recording the dynamic operations in real time.
8. The conference method according to claim 6, characterized in that: In the terminal device attribute acquisition step, each terminal device sends the physical space where it participates in the conference to the conference system; and the physical space where each terminal device is located is acquired based on the information reported by each terminal device.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 6 to 8 are implemented.
10. A storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 6 to 8 are implemented.
Citation Information
Patent Citations
Network teaching method and system
CN105405325A
Group session-based audio playing method and device, group session-based equipment management method and device and computer equipment
CN113516991A
Audio signal processing method, readable medium and electronic equipment
CN116684785A
Teleconferencing configuration based on proximity information
US20080160976A1
Potential echo detection and warning for online meeting
US20180077205A1