Audio array interaction method and system for network conference

Through audio array technology and deep learning algorithms, the sound signals are processed in online meetings, the target tone is generated and the speech content is converted, which solves the problems of non-speaking content interference and speech fluency, and improves user experience and meeting efficiency.

CN119028367BActive Publication Date: 2025-05-23广州帝声电子有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410968773.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-05-23
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

The existing online conferencing system has not yet effectively solved the problem of non-speaking content interference and speech fluency in terms of audio processing, affecting the user experience.

Method used

The audio array technology collects and processes sound signals, generates target tones, and converts speech content into speech content using target tones through speech conversion sub-mode and text conversion sub-mode. At the same time, a pure vocal mode is constructed using adaptive noise cancellation algorithm and deep learning algorithm to filter out speech noise.

Benefits of technology

It improves the audio processing effect in online meetings, reduces the pressure on speakers to speak, enhances the auditory experience of participants, and improves the stability and efficiency of the meeting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119028367B_ABST
    Figure CN119028367B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of audio processing technology, and discloses an audio array interaction method and system for network conferencing. The audio array technology is used to pre-collect and process user sound signals to form a target timbre, and the speaker's language speech content and text speech content are converted into speech voice using the target timbre for output, thereby reducing the speaker's speaking pressure and improving the auditory experience of the participants, so that the participants can more clearly understand what the speaker wants to express. In addition, the present invention also constructs a pure human voice model based on a deep learning algorithm to filter out speech noise that is irrelevant to the speech content emitted by the speaker during the speech process, further improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular to an audio array interaction method and system for network conferencing. Background Art

[0002] An audio array is a system composed of multiple microphones used to capture and process sound signals. By combining and analyzing the sound signals received by different microphones, the audio array can enhance the sound in a certain direction while suppressing noise in other directions. The core of this is that it can achieve beamforming, which is a method of enhancing or suppressing sound in a specific direction by adjusting the phase and amplitude of the signal received by the microphone.

[0003] Audio array technology is often used to improve the accuracy of speech recognition and the clarity of voice communication. For example, when conducting a network conference, there are usually multiple participants. At this time, it is necessary to locate the speaker so that other participants can clearly hear the speech without being disturbed by other noises. Current network conferencing systems have made significant progress in audio processing. Many systems have integrated basic noise suppression and echo cancellation functions, but there has been no effective improvement in non-speech content interference and speech fluency, which affects the user experience. Summary of the invention

[0004] In view of this, the object of the present invention is to provide an audio array interaction method and system for network conferencing, so as to improve the experience of participants after effectively processing the timbre and speech content in the audio in the network conference.

[0005] The first aspect of the present invention discloses an audio array interaction method for network conferencing, the method comprising the following steps:

[0006] S01. Voice collection, before the meeting starts, the user collects his voice from multiple angles through the audio array to obtain sound samples;

[0007] S02. Performing an audio processing operation on the collected sound sample to obtain a target sound sample, and saving the target sound sample;

[0008] S03. Generate a target timbre according to the target sound sample, and construct a timbre selection mode according to the target timbre;

[0009] S04. During the meeting, the user activates and uses the target timbre by clicking the timbre selection mode;

[0010] S05. Continue to click the conversion selection mode to perform the conversion operation to convert the speech content into the speech content using the target tone.

[0011] Furthermore, the conversion selection mode in step S05 includes a speech conversion sub-mode and a text conversion sub-mode.

[0012] Furthermore, the voice conversion sub-mode converts the user's real-time voice speech content into speech content using the target timbre and then outputs the speech.

[0013] Furthermore, the text conversion sub-mode converts the text speech content pre-input by the user into speech content using the target timbre and then outputs the speech in voice.

[0014] Furthermore, when a user speaks, an audio array is used to collect sound signals in the room, and the position of the speaker in the meeting is determined in real time through a beamforming algorithm. The current speaker is accurately identified using time difference and intensity difference analysis.

[0015] Furthermore, after identifying the current speaker, audio signal processing technology is used to analyze the indoor sound signals collected by the audio array in real time, identify the ambient noise and the speaker's voice signal, and dynamically adjust the noise filtering parameters through an adaptive noise elimination algorithm to obtain the target voice.

[0016] Furthermore, the method also includes constructing a pure human voice mode, and the user clicks on the pure human voice mode to filter out speech noise on the target voice, and the speech noise includes coughing and sneezing during the speech process.

[0017] Furthermore, the process of constructing a pure vocal model includes:

[0018] Prepare an audio data set containing speech content and speech noise, and annotate the data set to distinguish speech content from non-speech content; wherein the audio data set includes a training audio data set and a verification audio data set;

[0019] Select a deep learning architecture to train a pure human voice model, use the training audio data set as input data of the training model, use the speech content distinguished from the data set as output data of the training model, and train to obtain a first pure human voice model;

[0020] Inputting the verification audio data set into the first clean human voice model, verifying whether the result output by the first clean human voice model reaches a preset value compared with the speech content distinguished from the verification audio data set, and if the preset value is reached, using the first clean human voice model as a trained target clean human voice model;

[0021] Construct a pure vocal model based on the target pure vocal model.

[0022] The second aspect of the present invention discloses an audio array interactive system for network conferencing, which is implemented based on the method disclosed in the first aspect, and includes a collection module, an audio processing module, a storage module, a construction module and a control module;

[0023] The acquisition module is used to allow users to collect their own voices from multiple angles through the audio array to obtain sound samples before the meeting starts;

[0024] The audio processing module is used to perform audio processing operations on the collected sound samples to obtain target sound samples;

[0025] The storage module is used to save the target sound sample;

[0026] The building module is used to generate a target timbre according to a target sound sample, and to build a timbre selection mode according to the target timbre;

[0027] The control module is used to provide an interface for the user to select an audio processing mode to control the system to process the audio accordingly; wherein the audio processing mode includes a timbre selection mode and a conversion selection mode;

[0028] When the meeting is in progress, the user activates and uses the target timbre by clicking on the timbre selection mode; and continues to click on the conversion selection mode, and controls the system through the control module to perform the conversion operation, converting the speech content into speech content using the target timbre.

[0029] Furthermore, the conversion selection mode includes a voice conversion sub-mode and a text conversion sub-mode.

[0030] Furthermore, the voice conversion sub-mode converts the user's real-time voice speech content into speech content using the target timbre and then outputs the speech.

[0031] Furthermore, the text conversion sub-mode converts the text speech content pre-input by the user into speech content using the target timbre and then outputs the speech in voice.

[0032] Furthermore, when a user speaks, the acquisition module is also used to use an audio array to collect sound signals in the room, determine the position of the speaker in the meeting in real time based on the beamforming algorithm through the audio processing module, and use time difference and intensity difference analysis to accurately identify the current speaker.

[0033] Furthermore, after identifying the current speaker, the audio processing module uses audio signal processing technology to analyze the indoor sound signals collected by the audio array in real time, identify the ambient noise and the speaker's voice signal, and dynamically adjust the noise filtering parameters through an adaptive noise elimination algorithm to obtain the target voice.

[0034] Furthermore, the audio processing mode also includes a pure human voice mode, and the construction module is also used to construct the pure human voice mode. The user clicks on the pure human voice mode to filter out speech noise on the target voice, and the speech noise includes coughing and sneezing during the speech process.

[0035] Furthermore, the process of constructing a pure vocal model includes:

[0036] Prepare an audio data set containing speech content and speech noise, and annotate the data set to distinguish speech content from non-speech content; wherein the audio data set includes a training audio data set and a verification audio data set;

[0037] Select a deep learning architecture to train a pure human voice model, use the training audio data set as input data of the training model, use the speech content distinguished from the data set as output data of the training model, and train to obtain a first pure human voice model;

[0038] Inputting the verification audio data set into the first clean human voice model, verifying whether the result output by the first clean human voice model reaches a preset value compared with the speech content distinguished from the verification audio data set, and if the preset value is reached, using the first clean human voice model as a trained target clean human voice model;

[0039] Construct a pure vocal model based on the target pure vocal model.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The present invention uses audio array technology to pre-collect and process user voice signals to form a target timbre, and converts the speaker's verbal speech content and text speech content into speech voice using the target timbre for output, thereby reducing the speaker's speaking pressure while improving the auditory experience of the attendees, allowing the attendees to more clearly understand what the speaker wants to express; in addition, the present invention also constructs a pure human voice model based on a deep learning algorithm to filter out speech noise that is irrelevant to the speech content emitted by the speaker during the speech process, further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of the economic application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0043] Figure 1 A schematic diagram of a flow chart of an audio array interactive method for network conferencing disclosed in an embodiment of the present invention;

[0044] Figure 2The present invention discloses a structural diagram of an audio array interactive system for network conferencing according to another embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments.

[0046] Embodiment 1

[0047] The first aspect of the present invention discloses an audio array interactive method for network conferencing, see Figure 1 , Figure 1 The present invention discloses a flow chart of an audio array interactive method for a network conference, the method comprising the following steps:

[0048] S01. Voice collection, before the meeting starts, the user collects his voice from multiple angles through the audio array to obtain sound samples;

[0049] S02. Performing audio processing operations on the collected sound samples to obtain target sound samples, and saving the target sound samples;

[0050] S03. Generate a target timbre according to the target sound sample, and construct a timbre selection mode according to the target timbre;

[0051] S04. During the meeting, the user activates and uses the target timbre by clicking the timbre selection mode;

[0052] S05. Continue to click the conversion selection mode to perform the conversion operation to convert the speech content into the speech content using the target tone.

[0053] Specifically, in an embodiment of the present invention, the user's voice is collected by an audio array including multiple microphones, and the collected sound signal is recorded. Multi-angle collection through the audio array can ensure the capture of complete and clear sound samples. After collecting the user's voice sample, the sound sample is first subjected to basic audio preprocessing, including but not limited to noise reduction and echo cancellation, and digital signal processing DSP technology is used to optimize the quality of the sound sample through filters and equalizers.

[0054] In addition, the audio processing software allows the user to adjust the frequency response, volume and equalization parameters of the sound sample to obtain the user's target timbre. The target timbre in the embodiment of the present invention is preferably the timbre closest to the user's real speech, but the number of target timbre is not limited to one.

[0055] In the embodiment of the present invention, using the sound closest to the user's actual speech as the target timbre can shorten the communication distance between participants, present a face-to-face communication atmosphere, and make it easier for participants to focus on the meeting discussion.

[0056] During the meeting, the user clicks to apply the target timbre as needed. Specifically, by integrating the timbre selection interface in the conference software, the user can select the target timbre by clicking a button or menu, and the timbre mode selected by the user is applied in real time to adjust the timbre of the current speech.

[0057] Furthermore, the conversion selection mode in step S05 includes a voice conversion sub-mode

[0058] The voice conversion sub-mode converts the user's real-time voice speech content into speech content using the target timbre and then outputs the speech.

[0059] Specifically, in the speech conversion sub-mode, the speech conversion model can be pre-trained based on the deep learning algorithm to convert the real-time speech into a speech signal of the target timbre, and then the converted speech signal is output to the conference system to achieve the target timbre conversion of the real-time speech.

[0060] In an embodiment of the present invention, by converting real-time speech into speech with a target timbre for output, it can be used when the speaker has difficulty speaking due to hoarseness, sore throat or coughing. On the one hand, it reduces the speaker's speaking pressure, and on the other hand, it allows participants to clearly hear the speech content.

[0061] Furthermore, the conversion selection mode in step S05 also includes a text conversion sub-mode.

[0062] The text conversion sub-mode converts the text speech content pre-input by the user into speech content using the target tone and then outputs it as a voice.

[0063] Specifically, in the text conversion sub-mode, text preprocessing operations are performed on the text content input by the user, including but not limited to grammatical analysis and text segmentation operations, and a speech synthesis model (such as Tacotron 2) is used to convert the text content into a speech signal of the target timbre, and the synthesized speech signal is output to the conference system to achieve the target timbre conversion of the text speech content.

[0064] In an embodiment of the present invention, by converting the pre-written text content of the user into speech content using the target tone and then outputting it in voice, the pressure of on-site speaking for some users who are not fluent in language expression, such as socially anxious speakers, can be reduced. In addition, the fluency of the speaker's sentence expression during speech can be improved, and the phenomenon of stumbling can be reduced. On the one hand, it can improve the auditory experience of the participants, and on the other hand, it can promote the stable progress of the meeting, allowing the speaker to accurately express the content of the speech while allowing the participants to clearly understand the meaning of the content that the speaker wants to express, reducing the situation where the meeting efficiency is affected by the poor expression of the speech content.

[0065] Furthermore, when a user speaks, an audio array is used to collect sound signals in the room, and the position of the speaker in the meeting is determined in real time through a beamforming algorithm. The current speaker is accurately identified using time difference and intensity difference analysis.

[0066] Furthermore, as a preferred implementation of Example 1 of the present invention, a video surveillance algorithm is used to assist in positioning when determining the position of the speaker. Specifically, a camera is used to collect image data of the conference room in real time, and people in the image are detected in real time, and the identity of the speaker is determined in combination with facial recognition technology. The facial recognition technology here can be implemented by constructing a facial recognition model, which is obtained by training the model with a data set including the facial expressions of people when speaking. The audio signal and the video signal are synchronized and fused, and the video data is used to assist audio beamforming to enhance the accuracy of speaker positioning.

[0067] In addition, in the embodiment of the present invention, environmental sensors (such as noise sensors, temperature sensors, and light sensors) are deployed to collect environmental data in real time, and a real-time environmental model of the conference room is constructed using the environmental data. The environmental model is used to analyze the impact of noise sources and environmental changes on audio signals. According to the environmental model and audio signal characteristics, the beamforming parameters are dynamically adjusted to further assist the speaker positioning.

[0068] Furthermore, after identifying the current speaker, audio signal processing technology is used to analyze the indoor sound signals collected by the audio array in real time, identify the ambient noise and the speaker's voice signal, and dynamically adjust the noise filtering parameters through an adaptive noise elimination algorithm to obtain the target voice.

[0069] Specifically, in an embodiment of the present invention, the parameters of the noise reduction algorithm can also be dynamically adjusted according to the above-mentioned environmental model and audio signal characteristics to improve the noise filtering effect. For example, the filter coefficients are adjusted in real time using an adaptive LMS algorithm to maximize the signal-to-noise ratio. Among them, the noise reduction algorithm is mainly a noise elimination method that combines time domain and frequency domain. For example, an adaptive LMS filter is used in the time domain, and a Wiener filter is used in the frequency domain. The advantages of multiple domains are combined to improve the noise reduction effect. In addition, in the multi-channel audio processing process, collaborative noise reduction technology is used to further reduce noise by using the signals of multiple microphones through coherence analysis and collaborative processing.

[0070] Furthermore, the method also includes constructing a pure human voice mode, and the user clicks on the pure human voice mode to filter out speech noise from the target voice. As a preferred implementation, speech noise mainly refers to noise that is irrelevant to the content of the speech and is emitted by the speaker during the speech process, including but not limited to coughing and sneezing during the speech process. In addition, speech noise can also include environmental noise, such as keyboard sounds, page turning sounds, background music, etc.

[0071] Furthermore, the process of constructing a pure vocal model includes:

[0072] Prepare an audio data set containing speech content and speech noise, and annotate the data set to distinguish speech content from non-speech content; wherein the audio data set includes a training audio data set and a verification audio data set;

[0073] Select a deep learning architecture to train a pure human voice model, use the training audio data set as input data of the training model, use the speech content distinguished from the data set as output data of the training model, and train to obtain a first pure human voice model;

[0074] Inputting the verification audio data set into the first clean human voice model, verifying whether the result output by the first clean human voice model reaches a preset value compared with the speech content distinguished from the verification audio data set, and if the preset value is reached, using the first clean human voice model as a trained target clean human voice model;

[0075] Construct a pure vocal model based on the target pure vocal model.

[0076] As another preferred implementation of the embodiment of the present invention, a pure human voice model is constructed through the U-Net deep learning architecture.

[0077] Specifically, assuming that the original audio signal is x(t), the original audio signal is normalized, and the normalized signal is:

[0078]

[0079] Where μ is the mean of the audio signal, σ is the standard deviation of the audio signal, and t represents a continuous time variable.

[0080] x norm (t) is expressed in the discrete time domain as x norm (n), n represents a discrete time point. The normalized audio signal is subjected to short-time Fourier transform STFT to obtain a spectrum diagram:

[0081]

[0082] Among them, w(τ-n) is the window function, τ is the frame index used to indicate the position of short-time Fourier transform in time, and f is the frequency.

[0083] Use the U-Net model to separate the pure speech components from the spectrogram. Input the noisy spectrogram X noisy (f,τ), output the spectrogram X of pure speech clean (f,τ).

[0084] Assuming that the model parameter of the U-Net model is θ, the model is expressed as:

[0085]

[0086] The mean square error (MSE) is used as the loss function to measure the difference between the model output and the target pure speech spectrogram:

[0087]

[0088] Where N is the number of training samples, It is the pure speech spectrogram output by the model.

[0089] The audio dataset X will be verified val (f,τ) inputs the trained model and obtains the model output:

[0090]

[0091] Calculate the signal-to-noise ratio (SNR) on the validation set to evaluate the model performance:

[0092]

[0093] The verified model, i.e. the target pure voice model, is integrated into the model to build a pure voice mode, which is then integrated into the conference system. Users can click the pure voice mode button to activate the mode and filter out the speech noise from the target voice.

[0094] In the embodiment of the present invention, by setting the pure human voice mode, the speaker's speech quality and the auditory experience of the participants can be further improved by filtering out the speech noise such as coughing that does not involve the expression of the speech content during the speaker's speech. It should be noted that the speech input in the pure human voice mode can be the target speech after noise filtering, and the background noise in the speech can be further filtered out through this operation.

[0095] Embodiment 2

[0096] The second aspect of the present invention discloses an audio array interactive system for network conferencing, see Figure 2 , Figure 2 It is a structural schematic diagram of an audio array interactive system for network conferencing disclosed in another embodiment of the present invention, the system includes a collection module, an audio processing module, a storage module, a construction module and a control module;

[0097] The acquisition module is used to allow users to collect their own voices from multiple angles through the audio array to obtain sound samples before the meeting starts;

[0098] The audio processing module is used to perform audio processing operations on the collected sound samples to obtain target sound samples;

[0099] The storage module is used to save the target sound sample;

[0100] The building module is used to generate a target timbre according to a target sound sample, and to build a timbre selection mode according to the target timbre;

[0101] The control module is used to provide an interface for the user to select an audio processing mode to control the system to process the audio accordingly; wherein the audio processing mode includes a timbre selection mode and a conversion selection mode;

[0102] When the meeting is in progress, the user activates and uses the target timbre by clicking on the timbre selection mode; and continues to click on the conversion selection mode, and controls the system through the control module to perform the conversion operation, converting the speech content into speech content using the target timbre.

[0103] Furthermore, the conversion selection mode includes a voice conversion sub-mode and a text conversion sub-mode.

[0104] Furthermore, the user's real-time speech content is converted into speech content using the target timbre through the speech conversion sub-mode and then output as speech.

[0105] Furthermore, the text conversion sub-mode converts the speech content pre-input by the user into speech content using the target timbre and then outputs the speech.

[0106] Furthermore, when a user speaks, the acquisition module is also used to use an audio array to collect sound signals in the room, and the audio processing module determines the position of the speaker in the meeting in real time based on the beamforming algorithm, and uses time difference and intensity difference analysis to accurately identify the current speaker.

[0107] Furthermore, after identifying the current speaker, the audio processing module uses audio signal processing technology to analyze the indoor sound signals collected by the audio array in real time, identify the ambient noise and the speaker's voice signal, and dynamically adjust the noise filtering parameters through an adaptive noise elimination algorithm to obtain the target voice.

[0108] Furthermore, the audio processing mode also includes a pure human voice mode. The construction module is also used to construct the pure human voice mode. The user clicks on the pure human voice mode to filter out speech noise on the target voice. The speech noise includes coughing and sneezing during the speech process.

[0109] Furthermore, the process of constructing a pure vocal model includes:

[0110] Prepare an audio data set containing speech content and speech noise, and annotate the data set to distinguish speech content from non-speech content; the audio data set includes a training audio data set and a verification audio data set;

[0111] Select a deep learning architecture to train a pure human voice model, use a training audio data set as input data of the training model, use speech content distinguished from the data set as output data of the training model, and train to obtain a first pure human voice model;

[0112] Inputting the verification audio data set into the first clean human voice model, verifying whether the result output by the first clean human voice model reaches a preset value compared with the speech content distinguished from the verification audio data set, and if the preset value is reached, using the first clean human voice model as the trained target clean human voice model;

[0113] Construct a pure vocal model based on the target pure vocal model.

[0114] It should be noted that the specific implementation process of the second embodiment is similar to that of the first embodiment and will not be repeated in the second embodiment.

[0115] Finally, it should be noted that the audio array interactive method and system for network conferencing disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio array interactive method for network conferencing, characterized in that: The method comprises: S01. Voice collection, before the meeting starts, the user collects his voice from multiple angles through the audio array to obtain sound samples; S02. Performing audio processing operations on the collected sound samples to obtain target sound samples, and saving the target sound samples; S03. Generate a target timbre according to the target sound sample, and construct a timbre selection mode according to the target timbre; the target timbre is the timbre closest to the user's real speech sound; S04. During the meeting, the user activates and uses the target timbre by clicking the timbre selection mode; S05. Continue to click on the conversion selection mode to convert the user's speech content into speech content using the target timbre for voice output; the conversion selection mode includes a voice conversion sub-mode and a text conversion sub-mode; wherein the voice conversion sub-mode is used to convert the user's real-time voice speech content into speech content using the target timbre for voice output, and the text conversion sub-mode is used to convert the user's pre-input text speech content into speech content using the target timbre for voice output; When a user speaks, the audio array is used to collect the sound signals in the room, and the position of the speaker in the meeting is determined in real time through the beamforming algorithm. The current speaker is accurately identified by using time difference and intensity difference analysis. When determining the position of the speaker, we also use video surveillance algorithms to assist in positioning, use cameras to collect image data from the conference room in real time, detect people in the image in real time, and use facial recognition technology to determine the speaker's identity; When determining the speaker's position, environmental sensors are deployed to collect environmental data in real time. The environmental data is used to build a real-time environmental model of the conference room. The environmental model is used to analyze the impact of noise sources and environmental changes on audio signals. Based on the environmental model and audio signal characteristics, beamforming parameters are dynamically adjusted to assist in speaker positioning. After identifying the current speaker, the audio signal processing technology is used to analyze the indoor sound signals collected by the audio array in real time, identify the environmental noise and the speaker's voice signal, and dynamically adjust the noise filtering parameters through the adaptive noise elimination algorithm to obtain the target voice; A pure voice mode is constructed based on the target pure voice model. The user clicks on the pure voice mode to filter out speech noise for the target voice. The speech noise refers to the noise emitted by the speaker during the speech process that is irrelevant to the speech content, including coughing and sneezing during the speech process.

2. The audio array interactive method for network conferencing according to claim 1, characterized in that: The process of constructing a pure vocal model includes: Prepare an audio data set containing speech content and speech noise, and annotate the data set to distinguish speech content from non-speech content; wherein the audio data set includes a training audio data set and a verification audio data set; Select a deep learning architecture to train a pure human voice model, use the training audio data set as input data of the training model, use the speech content distinguished from the data set as output data of the training model, and train to obtain a first pure human voice model; The verification audio data set is input into the first pure human voice model to verify whether the result output by the first pure human voice model reaches a preset value compared with the speech content distinguished from the verification audio data set. If the preset value is reached, the first pure human voice model is used as the trained target pure human voice model.

3. An audio array interactive system for network conferencing, the audio array interactive system is implemented based on the audio array interactive method according to any one of claims 1-2, characterized in that: The system includes an acquisition module, an audio processing module, a storage module, a construction module and a control module; The acquisition module is used to allow users to collect their own voices from multiple angles through the audio array to obtain sound samples before the meeting starts; The audio processing module is used to perform audio processing operations on the collected sound samples to obtain target sound samples; The storage module is used to save the target sound sample; The construction module is used to generate a target timbre according to a target sound sample, and to construct a timbre selection module according to the target timbre. The target timbre is the timbre that is closest to the user's actual speaking voice; The control module is used to provide an interface for the user to select an audio processing mode to control the system to process the audio accordingly; wherein the audio processing mode includes a timbre selection mode, a conversion selection mode and a pure voice mode; During a meeting, the user activates and uses the target timbre by clicking on the timbre selection mode; And continue to click on the conversion selection mode, and the control module controls the system to perform the conversion operation, converting the user's speech content into speech content using the target timbre and then outputting it in voice; the conversion selection mode includes a speech conversion sub-mode and a text conversion sub-mode; wherein the speech conversion sub-mode is used to convert the user's real-time speech content into speech content using the target timbre and then output it in voice, and the text conversion sub-mode is used to convert the user's pre-input text speech content into speech content using the target timbre and then output it in voice; When a user speaks, the acquisition module is also used to collect the sound signals in the room using the audio array, and the audio processing module determines the position of the speaker in the meeting in real time based on the beamforming algorithm, and uses time difference and intensity difference analysis to accurately identify the current speaker; When determining the position of the speaker, we also use video surveillance algorithms to assist in positioning, use cameras to collect image data from the conference room in real time, detect people in the image in real time, and use facial recognition technology to determine the speaker's identity; When determining the speaker's position, environmental sensors are deployed to collect environmental data in real time. The environmental data is used to build a real-time environmental model of the conference room. The environmental model is used to analyze the impact of noise sources and environmental changes on audio signals. Based on the environmental model and audio signal characteristics, beamforming parameters are dynamically adjusted to assist in speaker positioning. After identifying the current speaker, the audio processing module uses audio signal processing technology to analyze the indoor sound signals collected by the audio array in real time, identify the environmental noise and the speaker's voice signal, and dynamically adjust the noise filtering parameters through an adaptive noise elimination algorithm to obtain the target voice; The construction module is also used to construct a pure human voice mode. The user clicks on the pure human voice mode to filter out speech noise on the target voice. The speech noise includes coughing and sneezing during the speech process.

Citation Information

Patent Citations

  • Identity-removed teleconference processing method and device and intelligent terminal

    CN112004050A

  • Ring array pickup control method and device, storage medium and ring array

    CN113068101A