System and method for stable recording of presentations

By combining regional audio signal acquisition with image acquisition, the problem of unstable speeches during seminars was solved, enabling accurate recording and clear output of speakers, and reducing crosstalk and noise interference.

CN114566181BActive Publication Date: 2026-04-17YUNJIACLOUD COM
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNJIACLOUD COM
Filing Date
2021-12-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing speech recognition systems struggle to reliably record speeches during seminars, especially when multiple people are speaking simultaneously, which can lead to crosstalk and noise interference, making it difficult for playback devices to accurately reproduce the corresponding speaker's voice information.

Method used

The method of combining regional sound source signal acquisition with image acquisition is adopted. Sound source information and personnel image information are obtained through sound acquisition module and image acquisition module. Noise and crosstalk are processed by processing module to achieve matching between speaker and sound source. Stable sound source output is achieved through audio output module.

Benefits of technology

It achieved stable recording of seminar presentations, reduced crosstalk and noise interference, and ensured the accuracy and clarity of the audio output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114566181B_ABST
    Figure CN114566181B_ABST
Patent Text Reader

Abstract

This invention discloses a system and method for stably recording speeches at seminars, including sound acquisition modules distributed throughout the seminar venue, each with its own acquisition area and connected to a processing module; an image acquisition module that acquires image information of personnel at the seminar venue and is also connected to the processing module; and a processing module that processes the audio source information to remove noise and crosstalk to obtain a stable audio source signal, and then matches the speaker with the audio source based on the processed audio source signal and the personnel image information before outputting the audio source. This invention achieves stable recording of seminar speeches by using sound acquisition modules to acquire audio source signals in designated areas, combining this with image acquisition modules to acquire images of speakers, and processing the acquired information to perform audio source gain, noise recognition, and crosstalk recognition processing to obtain stable audio source information. This stable audio source is then matched with the speaker and output, thus reducing crosstalk and noise interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound signal processing technology, and in particular to a system and method for stably recording speeches at seminars. Background Technology

[0002] Current speech recognition systems rely on several prerequisites to ensure high accuracy, the most important of which include: 1) the speaker's speech must be stable, clear, and easily captured by the pickup device; 2) speakers must speak one at a time; multiple simultaneous speeches can lead to incorrect recognition results, the most common being that the recognition result of someone else's speech appears on one's own recognition screen; 3) the recording environment must be relatively quiet, with minimal white noise and a uniform sound pickup pattern. Currently, there are two main strategies to optimize for these issues:

[0003] Firstly, adjusting the hardware facilities at each location and limiting the speaker's distance ensures that the speaker's voice is captured by the correct microphone, effectively acquiring the sound signal and preventing it from being picked up by other microphones, thus optimizing the recording process. Furthermore, each hardware device has multiple parameters, including microphone sensitivity, pickup range, and threshold, for real-time adjustment. However, this method is limited by the specific environment and recording equipment, lacking generalization ability. It requires tedious observation, recording, and repeated testing for each location. Moreover, due to the variability of scene conditions and the differences in voices and speaking habits among different target speakers, hardware adjustments often fail to address the aforementioned problems.

[0004] Secondly, regarding crosstalk, typical speech recognition systems, after acquiring the digital signals from each channel, perform pre-designed mathematical transformations and strategy calculations to predict the output channel of the recognition result for that frame and output accordingly. This method separates the sound signal acquisition and crosstalk recognition output during the crosstalk process, and cannot suppress or eliminate the crosstalk source portion of the sound source. Therefore, this method relies on the hardware acquisition results and is easily affected by the on-site hardware facilities and environmental structure. For example, the reflection and diffraction phenomena of the sound source can make the sound source acquired by some microphones more susceptible to crosstalk.

[0005] Currently, most speeches at seminars require recording and broadcasting. However, when multiple people speak simultaneously, crosstalk and noise occur, making it difficult for the broadcasting equipment to stably play the corresponding speaker's voice information.

[0006] For example, Chinese patent CN202010497438.0 discloses a method and apparatus for capturing meeting audio, recording meeting minutes, and presenting meeting minutes. It records speech by separating the speaker's voice; however, it still cannot solve the problem of unstable audio capture during the speaking process, making it difficult for playback devices to stably play the corresponding speaker's voice information. Summary of the Invention

[0007] This invention primarily addresses the problem of the difficulty in reliably recording speeches at seminars in existing technologies; it provides a system and method for reliably recording speeches at seminars.

[0008] The aforementioned technical problem of this invention is mainly solved by the following technical solution: a system for stable recording of speeches at a seminar, comprising several sound acquisition modules distributed and installed at the seminar site, each sound acquisition module being divided into acquisition areas, each sound acquisition module including several microphones for acquiring seminar audio source information and converting it into digital signals, each microphone being connected to a processing module; several image acquisition modules paired with the sound acquisition modules to acquire image information of personnel at the seminar site, and connected to the processing module; the processing module partitioning and marking the acquisition areas, acquiring audio source information from the microphones in each acquisition area and personnel image information from the image acquisition modules, processing the audio source information for noise and crosstalk to obtain a stable audio source signal, and matching the speaker with the audio source based on the processed audio source signal and personnel image information before outputting the audio source. By acquiring regional audio source signals from the seminar through the sound acquisition modules, combining this with image acquisition modules for speaker image acquisition, and processing the acquired information to obtain stable audio source information, which is then matched with the speaker for stable audio source output, stable recording of seminar speeches is achieved, reducing crosstalk and noise interference.

[0009] Preferably, the system also includes a mounting bracket, which comprises a mounting base fixed to a wall or ground and a rotating shaft rotatably mounted on the mounting base. The image acquisition module includes a camera and a gyroscope. The mounting base has several mounting slots for mounting the microphone. The camera is mounted on the rotating shaft, and the gyroscope is mounted on the camera to detect the camera's rotation angle. The rotating shaft is connected to a motor, and both the motor and the gyroscope are connected to an MCU. The MCU acquires the microphone's audio source information and controls the motor to rotate the rotating shaft based on this information, enabling the camera to capture an image of the speaker. This system, utilizing the gyroscope, motor, and rotating shaft, allows the camera to quickly align with the speaker and capture their image in real time, resulting in faster and better matching between the speaker and the audio source.

[0010] Preferably, the processing module includes an audio source signal preprocessing module, an audio source gain module, a noise recognition module, a crosstalk recognition module, an image processing module, and an audio output module. The audio source signal preprocessing module is connected to the sound acquisition module, the audio source gain module is connected to the audio source signal preprocessing module, the noise recognition module is connected to the audio source gain module, the crosstalk recognition module is connected to the noise recognition module, the image processing module is connected to both the crosstalk recognition module and the image acquisition module, and the audio output module is connected to both the image processing module and the crosstalk recognition module. The audio source signal preprocessing module extracts audio source signal features, the audio source gain module amplifies the audio source, the noise recognition module identifies noise, and the crosstalk recognition module identifies crosstalk, providing a stable audio source signal.

[0011] Preferably, an electromagnet is installed in the mounting slot, and a permanent magnet is installed on the microphone. The electromagnet attracts or repels the permanent magnet. A slot is provided on the side of the microphone, and a locking block is provided on the side wall of the mounting slot. The locking block and the slot match to secure the microphone in the mounting slot. When the electromagnet is energized, it becomes magnetic, and its north and south poles can change with the direction of the current. When the electromagnet is controlled to attract the permanent magnet, the magnetic attraction force secures the microphone in the mounting slot. The locking block and the slot are used for installation and positioning.

[0012] Preferably, the locking block is an arc-shaped locking block. When the locking block is arc-shaped, if a microphone replacement problem occurs, the microphone needs to be removed. The electromagnet changes its north and south poles, causing the electromagnet to repel the permanent magnet. The electromagnetic repulsion force is greater than the friction force between the arc-shaped locking block and the slot, causing the microphone to pop out of the mounting slot, making microphone replacement convenient.

[0013] Preferably, the locking block is a rectangular locking block, and the side wall of the mounting slot is provided with a storage slot for accommodating the rectangular locking block. A spring is installed inside the rectangular locking block. When the spring is not energized, it is in its natural state, allowing the rectangular locking block to engage with the slot. When the spring is energized, it retracts into the storage slot. When the spring retracts completely, the rectangular locking block and slot no longer have a limiting effect during operation, making microphone replacement faster. Furthermore, when the electromagnet is suddenly de-energized, the microphone will not fall out of the mounting slot due to bumps or other reasons, thus improving safety.

[0014] This invention also provides a method for stably recording speeches at a seminar, comprising the following steps:

[0015] The sound acquisition module acquires signals from different sound sources, and the image acquisition module acquires images of the speaker.

[0016] Feature extraction is performed on the acquired sound source signals;

[0017] The audio source signal is subjected to adaptive audio source gain, noise identification, and crosstalk identification.

[0018] The optimized audio signal is matched with the speaker before being output as an audio source.

[0019] As a preferred method, the adaptive audio source gain method is as follows: obtain the audio source signal in a certain audio source channel of the current frame, obtain the K-frame historical frame signal of the audio source signal provider of the current frame, and input the K+1 frame audio source signal into the feedforward memory network to obtain the amplified audio source signal.

[0020] Preferably, the crosstalk identification method is as follows: similarity calculation is performed on the feature data of each channel; for channels with high similarity, a time-series Markov process is used to align the digital signals in time, and similar channels with a time delay are identified. These similar channels with a time delay are then determined as crosstalk channels.

[0021] As a preferred approach, anomaly detection is performed by taking the features of each channel at the current time and the features in historical time frames to identify the process of the microphone suddenly collecting sound, and calculating the probability that this process is crosstalk. The crosstalk probability and the crosstalk channel determination result are weighted and calculated to obtain the final crosstalk recognition result.

[0022] The beneficial effects of this invention are as follows: by using a sound acquisition module to collect regional sound source signals for the seminar, and combining this with an image acquisition module to collect images of the speakers, the processing module performs sound source gain, noise identification, and crosstalk identification processing on the collected information to obtain stable sound source information. After matching the information with the speaker, stable sound source output is achieved, thus realizing stable recording of seminar speeches and reducing crosstalk and noise interference. Attached Figure Description

[0023] Figure 1 This is a system structure block diagram of an embodiment of the present invention.

[0024] Figure 2 This is a flowchart of a method according to an embodiment of the present invention.

[0025] The diagram shows: 1. Sound acquisition module, 2. Image acquisition module, 3. Sound source signal preprocessing module, 4. Sound source gain module, 5. Noise recognition module, 6. Crosstalk recognition module, 7. Image processing module, and 8. Audio output module. Detailed Implementation

[0026] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0027] It should be noted that in the following description, reference is made to the accompanying drawings, which illustrate several embodiments of the present invention. It should be understood that other embodiments may also be used, and changes in mechanical composition, structure, electrical system, and operation may be made without departing from the spirit and scope of the invention. The following detailed description should not be considered limiting, and the scope of the embodiments of the invention is defined only by the claims of the published patents. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. Spatially related terms, such as “upper,” “lower,” “left,” “right,” “below,” “below,” “lower part,” “above,” “upper part,” etc., may be used herein to illustrate the relationship between one element or feature shown in the figures and another element or feature.

[0028] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," "fixing," and "holding" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0029] Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context indicates otherwise. The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data used can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising,” “including,” indicate the presence of the stated features, operations, elements, components, items, kinds, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. It should be further understood that the terms “or” and “and / or” as used herein are interpreted as inclusive, or mean any one or any combination thereof. Thus, “A, B, or C” or “A, B, and / or C” means “any one of: A; B; C; A and B; A and C; B and C; A, B, and C.” An exception to this definition will only occur if the combination of elements, functions, or operations is inherently mutually exclusive in some way.

[0030] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the invention.

[0031] Example: A system for stably recording speeches at a seminar, such as Figure 1 As shown, the system includes a sound acquisition module 1, an image acquisition module 2, a sound source signal preprocessing module 3, a sound source gain module 4, a noise recognition module 5, a crosstalk recognition module 6, an image processing module 7, and an audio output module 8. The sound source signal preprocessing module is connected to the sound acquisition module, the sound source gain module is connected to the sound source signal preprocessing module, the noise recognition module is connected to the sound source gain module, the crosstalk recognition module is connected to the noise recognition module, the image processing module is connected to both the crosstalk recognition module and the image acquisition module, and the audio output module is connected to both the image processing module and the crosstalk recognition module.

[0032] The system includes several sound acquisition modules, which are distributed and installed at the seminar venue. Each sound acquisition module is divided into acquisition areas, and each sound acquisition module includes several microphones to collect the seminar's audio source information and convert it into digital signals.

[0033] The image acquisition module, consisting of several units, is paired with the sound acquisition module to collect image information of people at the seminar.

[0034] This invention also provides a method for stably recording speeches at seminars, such as... Figure 2 As shown, it includes the following steps:

[0035] S1: The sound acquisition module acquires signals from different sound sources, and the image acquisition module acquires images of the speakers.

[0036] S2: Extract features from the acquired audio source signal; determine the volume index of the original signal based on the provided preset feature quantization and analysis. For example, for a certain channel's data l_in^t, which has 4000 numbers, take a window size of 200 and a total number of windows of 20, and obtain window data with a dimension of 20×200. This data is then fused through feature fusion and finally expressed as a vector l_out^t={0.152,-0.236,…,1.768} with a length of 512. After performing the above operation on all channels, feature data with a dimension of 4×512 will be output.

[0037] S3: Adaptive audio source gain, noise identification, and crosstalk identification are performed on the audio source signal. The audio source gain method is as follows: acquire the audio source signal in a certain audio source channel of the current frame, and acquire K frames of historical frame signals from the audio source signal provider of this frame. Input the K+1 frames of audio source signals into the feedforward memory network to obtain the amplified audio source signal. Specifically, for the acoustic digital signal L_in^t={l_1^t,l_2^t,…,l_4000^t} of a speaker in the current frame, in addition to words, K frames of historical frame signals L_in^(t-1), L_in^(t-2),…,L_in^(tK) of the speaker are also needed. A total of K+1 frames of signal are passed through the feedforward memory network. The network passes through multiple layers of feedforward neural network and memory network, and outputs a floating-point vector h^l={h_1^l,…,h_H^l} of length H representing the local acoustic information features and a dimension of H. The floating-point vector h^g = {h_1^g,…,h_H^g} representing the historical acoustic information features is used. After weighted activation, the output is a floating-point number p = Relu(W_l h^l + W_g h^g + b), where Relu is an activation function, and W and b are trained parameters. For example, if p = 0.5, then L_out = 0.5∙L_in. According to the gain strategy provided by this method, the sound source of each channel will determine its unique gain parameters in real time based on its own characteristics. The gained sound source has a clear and stable listening effect, without inaudible or plosive sounds. In practical applications, it also plays a significant role in improving recognition efficiency and preserving sound source information.

[0038] The noise identification method is as follows: based on the feature extraction results of each sound source channel, sound source classification is performed, a classification model is established, a noise threshold is set, the digital signal power of the sound source channel is calculated, and the classification model result is output. Specifically, for the acoustic digital signal of a certain channel, its power is first calculated, such as watt_in^t=20. Then, after feature data extraction, it is entered into the classification model, and the probability that it is a noise source is output as p_in^t=0.6. If either of the two exceeds the given threshold, the channel is determined to be an environmental noise source. The classification model is set by a big data fusion classification model established by integrating data features based on historical data or a large amount of experimental data, which has a certain degree of reliability.

[0039] The crosstalk identification method is as follows: Similarity calculation is performed on the feature data of each channel. For channels with high similarity, a temporal Markov process is used to align the digital signals in time, identifying similar channels with a temporal delay. These similar channels with a temporal delay are then identified as crosstalk channels. Anomaly detection is performed on the features of each channel at the current time and in historical time frames to identify the process of the microphone suddenly picking up sound. The probability of this process being crosstalk is calculated, and the crosstalk probability and the crosstalk channel identification result are weighted to obtain the final crosstalk identification result. Specifically, there are feature data for four channels, F_in^t={〖f1〗_in^t,〖f2〗_in^t,〖f3〗_in^t,〖f4〗_in^t}. Similarity calculation is performed pairwise. Assuming that the similarity of channels 2, 3, and 4 is high, s_2,3=80%, s_2,4=85%, s_3,4=73%,... That is, channels 2 and 3 are 80% similar, channels 2 and 4 are 85% similar, and channels 3 and 4 are 73% similar. Then, the original acoustic digital signals acquired from the three channels are time-series aligned. The alignment process involves calculating the time periods in which the similar parts between similar channel pairs occur and selecting the alignment path with the highest probability. Assuming that after alignment, channels 3 and 4 are both delayed compared to channel 2, then channels 3 and 4 are crosstalk channels relative to channel 2. For a certain channel, the feature data at that time frame is a 512-dimensional f_in^t = {0.121, -1.423, ..., 1.845}. K frames of historical feature data f_in^(t-1), f_in^(t-2), ..., f_in^(tK) are taken, and for this K+1... The features of the frames are used to model a feedforward memory network in a temporal sequence. The result will output a probability, which represents the probability of crosstalk occurring based on the data pattern of K historical frames, such as p=0.87. The crosstalk occurrence probability and the crosstalk channel determination result are weighted and calculated to obtain the final crosstalk recognition result. The crosstalk channels are marked, and the weighted values ​​are obtained based on actual experiments.

[0040] S4: After matching the optimized audio source signal with the speaker, the audio source is output. At a specific time frame, all microphone channel audio sources have three processing results: normal audio source, ambient noise audio source, and crosstalk audio source. The recognition result for the ambient noise audio source is set to empty, the recognition result for the crosstalk audio source is sent to the image processing unit, and the normal audio source is input into the speech recognition module for speech recognition. Combined with the speaker's facial information collected by the image acquisition unit, the normal audio source is matched with the speaker's role. The matching method is as follows: the image processing unit acquires the image information collected by the image acquisition unit, the crosstalk channel audio source, and the normal audio source; it extracts the same voiceprint information between the crosstalk and normal audio sources; it obtains two times by recording the audio source information to mark the sound source location; it combines the image information to confirm the sound source role and performs normal audio source matching, outputting the matching result; the audio output module combines the speaker matching result with the normal audio source to output a stable audio signal.

[0041] In another embodiment of the present invention, a mounting bracket is provided, which includes a mounting base fixed to a wall or ground and a rotating shaft rotatably mounted on the mounting base. The image acquisition module includes a camera and a gyroscope. The mounting base is provided with several mounting slots for mounting a microphone. The camera is mounted on the rotating shaft, and the gyroscope is mounted on the camera for detecting the rotation angle of the camera. The rotating shaft is connected to a motor, and both the motor and the gyroscope are connected to an MCU. The MCU acquires the sound source information of the microphone and controls the motor to rotate the rotating shaft according to the acquired sound source information so that the camera can capture an image of the speaker.

[0042] An electromagnet is installed in the mounting slot, and a permanent magnet is installed on the microphone. The electromagnet attracts or repels the permanent magnet. A slot is provided on the side of the microphone, and a block is provided on the side wall of the mounting slot. The block and the slot match to snap the microphone into the mounting slot. The block is an arc-shaped block.

[0043] When an electromagnet is energized, it becomes magnetic, and its north and south poles can change depending on the direction of the current. When the electromagnet is controlled to attract the permanent magnet, the magnetic attraction force fixes the microphone in place in the mounting slot. The locking block and slot are used for installation and limiting. When the locking block is arc-shaped, if the microphone needs to be replaced, it needs to be removed. The electromagnet changes its north and south poles, causing the electromagnet and permanent magnet to repel each other. The electromagnetic repulsion force is greater than the friction force between the arc-shaped locking block and the slot, causing the microphone to pop out of the mounting slot, making it easy to replace the microphone.

[0044] In another embodiment of the present invention, the card block is a rectangular card block, and the side wall of the mounting groove is provided with a storage groove for storing the rectangular card block. A spring is provided inside the rectangular card block. When the spring is not energized, it is in a natural state, so that the rectangular card block is engaged with the card groove. When the spring is energized, the spring retracts and enters the storage groove.

[0045] When the spring is energized and retracts, it is completely retracted into the storage slot. The rectangular block and slot no longer have a starting limit function, making it faster to change the microphone. At the same time, when the electromagnet is suddenly de-energized, the microphone will not fall out of the mounting slot due to bumps or other reasons, making it safer.

[0046] Based on this, the locking block of the present invention can also be set in the shape of a right triangle, with its hypotenuse facing outward and its straight side facing the bottom of the mounting groove. It is connected to the bottom of the storage groove through an energized spring. When the microphone is placed into the mounting groove, the energized spring is compressed, causing the right triangle locking block to retract into the storage groove. After the microphone enters the mounting groove, its locking slot corresponds to the locking block, and the energized spring returns to its original position, realizing the limiting locking. When replacing, the energized spring is energized and retracts, and the microphone is taken out for replacement.

[0047] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Other variations and modifications are possible without departing from the technical solutions described in the claims.

Claims

1. A method for stably recording speeches at a seminar, characterized by: Includes the following steps: The sound acquisition module acquires signals from different sound sources, and the image acquisition module acquires images of the speaker. Feature extraction is performed on the acquired sound source signals; The audio source signal is subjected to adaptive audio source gain, noise identification, and crosstalk identification. The crosstalk identification method is as follows: similarity calculation is performed on the feature data of each channel; for channels with high similarity, the digital signals are time-aligned through a temporal Markov process to identify similar channels with a backward time delay, and these similar channels with a backward time delay are identified as crosstalk channels; anomaly detection is performed on the features of each channel at the current time and the features in historical time frames to identify the process of the microphone suddenly collecting sound, and the probability of this process being crosstalk is calculated; the crosstalk probability and the crosstalk channel identification result are weighted and calculated to obtain the final crosstalk identification result; The optimized audio signal is matched with the speaker before being output as an audio source.

2. The method for stably recording speeches at a seminar according to claim 1, characterized in that, The method of adaptive audio source gain is as follows: obtain the audio source signal in a certain audio source channel of the current frame, obtain the K-frame historical frame signal of the audio source signal provider of the current frame, and input the K+1 frame audio source signal into the feedforward memory network to obtain the gained audio source signal.

3. A system for stably recording speeches at seminars, characterized in that: The method for stably recording speeches at a seminar as described in any one of claims 1-2 includes: Several sound acquisition modules are set up and distributed at the seminar site. Each sound acquisition module is divided into acquisition areas. Each sound acquisition module includes several microphones for acquiring the seminar's sound source information and converting it into digital signals. All of the microphones are connected to the processing module. The image acquisition module, consisting of several units, is paired with the sound acquisition module to collect image information of people at the seminar and connect to the processing module. The processing module divides the acquisition area into zones and marks them. It acquires the audio source information transmitted by the microphone in each acquisition area and the personnel image information transmitted by the image acquisition module. After noise and crosstalk processing of the audio source information, it obtains a stable audio source signal. Based on the processed audio source signal and personnel image information, it matches the speaker with the audio source and then outputs the audio source.

4. The system for stably recording speeches at a seminar according to claim 3, characterized in that, It also includes a mounting bracket, which includes a mounting base fixed to a wall or ground and a rotating shaft rotatably mounted on the mounting base. The image acquisition module includes a camera and a gyroscope. The mounting base is provided with several mounting slots for mounting the microphone. The camera is mounted on the rotating shaft, and the gyroscope is mounted on the camera to detect the rotation angle of the camera. The rotating shaft is connected to a motor, and both the motor and the gyroscope are connected to an MCU. The MCU acquires the sound source information of the microphone and controls the motor to rotate the rotating shaft based on the acquired sound source information so that the camera can capture an image of the speaker.

5. The system for stably recording speeches at a seminar according to claim 3, characterized in that, The processing module includes an audio source signal preprocessing module, an audio source gain module, a noise recognition module, a crosstalk recognition module, an image processing module, and an audio output module. The audio source signal preprocessing module is connected to the sound acquisition module, the audio source gain module is connected to the audio source signal preprocessing module, the noise recognition module is connected to the audio source gain module, the crosstalk recognition module is connected to the noise recognition module, the image processing module is connected to both the crosstalk recognition module and the image acquisition module, and the audio output module is connected to both the image processing module and the crosstalk recognition module.

6. The system for stably recording speeches at a seminar according to claim 4, characterized in that, An electromagnet is installed in the mounting slot, and a permanent magnet is installed on the microphone. The electromagnet attracts or repels the permanent magnet. A slot is provided on the side of the microphone, and a locking block is provided on the side wall of the mounting slot. The locking block and the slot match to lock the microphone into the mounting slot.

7. The system for stably recording speeches at a seminar according to claim 6, characterized in that, The card block is an arc-shaped card block.

8. The system for stably recording speeches at a seminar according to claim 6, characterized in that, The card block is a rectangular card block, and the side wall of the mounting groove is provided with a storage groove for storing the rectangular card block. A spring is provided inside the rectangular card block. When the spring is not energized, it is in a natural state, allowing the rectangular card block to engage with the card groove. When the spring is energized, the spring retracts and enters the storage groove.

9. The system for stably recording speeches at a seminar according to claim 6, characterized in that, The card block is set as a right triangle with its hypotenuse facing outwards and its straight side facing the bottom of the mounting slot.

10. The system for stably recording speeches at a seminar according to claim 6, characterized in that, When the electromagnet attracts the permanent magnet, the microphone is fixedly clipped into the mounting slot by the magnetic attraction.

Citation Information

Patent Citations

  • Conference sound acquisition method and device, conference record method and device and conference record presentation method and device

    CN111739553A

  • Text record-based video archiving device and method

    CN108712624A

  • Method and device for generating conference record and conference terminal

    CN110232925A

  • Anti-crosstalk method, device and equipment based on multi-pickup scene

    CN112151036A

  • Method and device for adjusting volume of audio equipment, electronic device and medium

    CN112698808A