Conference interaction method and apparatus, computer device and storage medium
By performing wake word detection and voiceprint matching in online meetings, the system ensures that only users with speaking privileges can switch to speaking mode, thus solving the problem of voice leakage and improving meeting security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-08-15
- Publication Date
- 2026-04-21
AI Technical Summary
In online meetings, even when a speaker turns on their microphone, their voice, including parts unrelated to the meeting content, can still be transmitted, affecting meeting security.
By performing wake word detection and voiceprint matching on the local end, it ensures that only users with speaking permissions can switch to speaking mode, and collects and sends conference speaking voice data.
This effectively prevents voice leaks and improves meeting security.
Smart Images

Figure CN115312057B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a conference interaction method, apparatus, computer equipment, storage medium, and computer program product. Background Technology
[0002] With the development of computer technology, meeting formats have become increasingly diverse, no longer limited to participants gathering in a unified conference room. Remote audio and video conferencing allows for meetings across geographical locations, facilitating people's work and lives. However, during online meetings, speakers need to turn on their microphones. But once the microphone is on, all audio emitted by the speaker, including irrelevant remarks, will still be transmitted, leading to audio leakage and compromising meeting security. Summary of the Invention
[0003] Therefore, it is necessary to provide a meeting interaction method, device, computer equipment, computer-readable storage medium, and computer program product that can avoid voice leakage and improve meeting security in response to the above-mentioned technical problems.
[0004] Firstly, this application provides a meeting interaction method. The method includes:
[0005] When the meeting is in a non-speaking state awaiting voice wake-up on the local end, it responds to the voice wake-up trigger event and obtains the wake-up trigger voice data from the local end;
[0006] Perform wake word detection on the wake-up trigger voice data to obtain the wake word detection results;
[0007] The wake-up trigger voice data is matched with the standard voiceprint data of the account with speaking permissions associated with the local terminal to obtain the voiceprint matching result.
[0008] When the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match, the meeting will be switched to local speaking mode;
[0009] When the participant is in speaking mode on their local device, the system collects the audio data of the participant's speech and sends it to the meeting client.
[0010] Secondly, this application also provides a conference interaction device. The device includes:
[0011] The wake-up trigger voice acquisition module is used to acquire wake-up trigger voice data from the local end when the meeting is in a non-speaking state awaiting voice wake-up on the local end in response to a voice wake-up trigger event.
[0012] The wake word detection module is used to detect wake words in wake-up trigger voice data and obtain wake word detection results;
[0013] The voiceprint matching module is used to match the wake-up trigger voice data with the standard voiceprint data of the speaking permission account associated with the local terminal to obtain the voiceprint matching result.
[0014] The voice state switching module is used to switch the conference to the local speaking state when the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match.
[0015] The speech processing module is used to collect the speech data of the local device when the user is speaking during the meeting, and then send the speech data to the meeting client.
[0016] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0017] When the meeting is in a non-speaking state awaiting voice wake-up on the local end, it responds to the voice wake-up trigger event and obtains the wake-up trigger voice data from the local end;
[0018] Perform wake word detection on the wake-up trigger voice data to obtain the wake word detection results;
[0019] The wake-up trigger voice data is matched with the standard voiceprint data of the account with speaking permissions associated with the local terminal to obtain the voiceprint matching result.
[0020] When the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match, the meeting will be switched to local speaking mode;
[0021] When the participant is in speaking mode on their local device, the system collects the audio data of the participant's speech and sends it to the meeting client.
[0022] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0023] When the meeting is in a non-speaking state awaiting voice wake-up on the local end, it responds to the voice wake-up trigger event and obtains the wake-up trigger voice data from the local end;
[0024] Perform wake word detection on the wake-up trigger voice data to obtain the wake word detection results;
[0025] The wake-up trigger voice data is matched with the standard voiceprint data of the account with speaking permissions associated with the local terminal to obtain the voiceprint matching result.
[0026] When the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match, the meeting will be switched to local speaking mode;
[0027] When the participant is in speaking mode on their local device, the system collects the audio data of the participant's speech and sends it to the meeting client.
[0028] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0029] When the meeting is in a non-speaking state awaiting voice wake-up on the local end, it responds to the voice wake-up trigger event and obtains the wake-up trigger voice data from the local end;
[0030] Perform wake word detection on the wake-up trigger voice data to obtain the wake word detection results;
[0031] The wake-up trigger voice data is matched with the standard voiceprint data of the account with speaking permissions associated with the local terminal to obtain the voiceprint matching result.
[0032] When the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match, the meeting will be switched to local speaking mode;
[0033] When the participant is in speaking mode on their local device, the system collects the audio data of the participant's speech and sends it to the meeting client.
[0034] The aforementioned meeting interaction methods, devices, computer equipment, storage media, and computer program products, when the meeting is in a non-speaking state awaiting voice wake-up on the local end, perform wake-up word detection and voiceprint matching on the wake-up trigger voice data on the local end. When the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word, and the voiceprint matching result is a match, the meeting is switched to a speaking state on the local end, and the collected meeting speaking voice data is sent to the meeting client. During the meeting speaking control process, when the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word, and the voiceprint matching result is a match, it ensures that the speaking on the local end is effectively confirmed, which can prevent voice leakage and thus improve the security of the meeting. Attached Figure Description
[0035] Figure 1This is a diagram illustrating the application environment of a meeting interaction method in one embodiment;
[0036] Figure 2 This is a flowchart illustrating a meeting interaction method in one embodiment;
[0037] Figure 3 This is a schematic diagram of the voiceprint matching process in one embodiment;
[0038] Figure 4 This is a schematic diagram of the interface for a meeting to be woken up by voice in one embodiment;
[0039] Figure 5 This is a schematic diagram of the interface after the meeting is woken up by voice in one embodiment.
[0040] Figure 6 This is a schematic diagram of the interface for configuring the wake word in one embodiment;
[0041] Figure 7 This is a flowchart illustrating the meeting interaction method in another embodiment;
[0042] Figure 8 This is a schematic diagram of the wake word recognition feature extraction process in one embodiment;
[0043] Figure 9 This is a schematic diagram of the wake word detection model in one embodiment;
[0044] Figure 10 This is a schematic diagram of the voiceprint matching process in another embodiment;
[0045] Figure 11 This is a structural block diagram of a conference interaction device in one embodiment;
[0046] Figure 12 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0048] The meeting interaction method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, multiple terminals 102 communicate with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another server. A conference client can be installed on each terminal 102, allowing users to hold online meetings and communicate online. When the conference is in a non-speaking state awaiting voice wake-up on the local end of terminal 102, terminal 102 responds to a user-triggered voice wake-up event by performing wake-up word detection and voiceprint matching on the local wake-up trigger voice data. When the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word, and the voiceprint matching result is a match, terminal 102 switches the conference to a speaking state on the local end and sends the collected conference speaking voice data to other conference clients via server 104. Furthermore, the conference interaction method can be implemented independently by server 104, or it can be executed jointly by terminal 102 and server 104.
[0049] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, aircraft, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0050] In practical applications, wake-up word detection and voiceprint matching for wake-up trigger voice data can be achieved based on speech recognition technology in Artificial Intelligence (AI), such as by establishing a corresponding network model for processing. Artificial intelligence utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to obtain optimal results—theories, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0051] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning. Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech being one of the most promising methods. Natural Language Processing (NLP) is an important area within computer science and AI. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, and thus it is closely related to linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies. The meeting record processing solution provided in this application involves speech technology and natural language processing technology within artificial intelligence.
[0052] In one embodiment, such as Figure 2 As shown, a meeting interaction method is provided. This method is executed by a computer device, specifically, it can be executed by a computer device such as a terminal or a server alone, or it can be executed by both a terminal and a server. In this embodiment, the method is applied to... Figure 1Taking the terminal in the example, the explanation includes the following steps:
[0053] Step 202: When the meeting is in a non-speaking state awaiting voice wake-up on the local end, in response to the voice wake-up trigger event, the wake-up trigger voice data on the local end is obtained.
[0054] The meeting can include various forms of online conferences, such as audio conferences and video conferences. Participants can communicate remotely by speaking. The local client refers to the client running on the user's local machine. Specifically, when listening to or speaking during a conference through a conference client, the local client can be the conference client running locally on the device. Non-speaking state refers to the meeting state where the local client does not speak; that is, the local client does not use its microphone to speak, but instead listens to the speeches of other clients in the meeting. In a web conference, the conference client has both speaking and non-speaking states. In speaking state, the conference client uses its microphone, allowing the user to speak and have their voice amplified throughout the meeting. In non-speaking state, the conference client does not use its microphone, and does not capture and transmit voice messages; instead, it listens to the voice messages transmitted by other clients in the meeting. "Pending voice wake-up" means the meeting is in a non-speaking state on the local client, waiting for the user to wake it up via voice to switch to speaking state and allow the user to speak.
[0055] A voice wake-up trigger event is an event that triggers a wake-up call via voice. Voice wake-up trigger events can be pre-set to occur as needed. For example, they can be triggered when a user initiates a voice wake-up operation, when a user emits voice input on the local end, or when pre-defined trigger conditions are met, such as when a trigger time is reached. Wake-up trigger voice data consists of voice data emitted by meeting participants on their local end, used to trigger a wake-up call for the meeting and switch the meeting's audio status on the local end.
[0056] Specifically, during a meeting conducted via a conference client on a terminal, when the local end is in a non-speaking state awaiting voice wake-up, it indicates that the local end is not speaking but is listening to the voice broadcast from other clients. The terminal responds to the voice wake-up trigger event by acquiring the wake-up trigger voice data from the local end and then performs voice wake-up processing based on this data. In practical applications, when a client enters a meeting, it can be set by default to a non-speaking state awaiting voice wake-up on the local end to prevent voice leakage when various conference clients enter the meeting. Alternatively, users can actively configure the voice state, such as setting the meeting to a non-speaking state on the local end, indicating that the user is not currently speaking. The terminal can detect whether a voice wake-up trigger event has occurred on the local end. If a voice wake-up trigger event is detected, the terminal acquires the wake-up trigger voice data from the local end and performs voice wake-up processing based on this data.
[0057] Step 204: Perform wake-up word detection on the wake-up trigger voice data to obtain the wake-up word detection result.
[0058] The wake word is a keyword used to initiate speech. Wake words can be set according to actual needs. For example, wake words can include "start speaking," "local speaking enabled," "request to speak," etc. Wake word detection involves performing keyword detection on the voice data to obtain wake word detection results, thereby determining whether the voice data contains a wake word, i.e., whether a wake word has been matched.
[0059] Specifically, the terminal performs wake-up word detection on the wake-up trigger voice data from the local end. Specifically, it can perform speech recognition on the wake-up trigger voice data to obtain the speech recognition text corresponding to the wake-up trigger voice data, and match the speech recognition text with the pre-set wake-up word to determine whether the wake-up trigger voice data includes the wake-up word, thereby obtaining the wake-up word detection result.
[0060] Step 206: The wake-up trigger voice data is matched with the standard voiceprint data of the account with speaking permissions associated with the local terminal to obtain the voiceprint matching result.
[0061] Among them, a speaking permission account refers to an account with speaking privileges on the local end. Speaking permission accounts can be pre-configured according to actual needs. For example, when starting a meeting, permissions can be set for meeting participants, configuring speaking permissions for each participant. The participant with speaking privileges is the speaking permission account. There can be one or more speaking permission accounts. One speaking permission account indicates that only one speaking permission account is supported on the local end; multiple speaking permission accounts indicate that multiple speaking permission accounts are supported on the local end, meaning multiple speaking permission accounts can be triggered to speak via voice wake-up. Standard voiceprint data is the baseline voiceprint data corresponding to the speaking permission account, which can be pre-entered by the user based on the speaking permission account. Using standard voiceprint data, identity recognition can be performed on the voice data on the local end to determine the identity of the user sending the voice data. Voiceprint matching refers to the process of matching wake-up trigger voice data with standard voiceprint data to identify the user sending the wake-up trigger voice data. Based on the obtained voiceprint matching results, it can be determined whether the wake-up trigger voice data was sent by the user corresponding to the speaking permission account.
[0062] Specifically, the terminal can verify the identity of the user who issued the wake-up trigger voice data based on the wake-up trigger voice data. This can be achieved by the terminal obtaining the standard voiceprint data of the account associated with speaking permissions on the local machine. The speaking permission account has speaking permissions on the local machine and can be pre-configured according to actual needs, with its standard voiceprint data recorded for identification. The terminal then performs voiceprint matching between the wake-up trigger voice data and the standard voiceprint data. In practice, the standard voiceprint data can include standard voiceprint features extracted from standard voiceprint data. The terminal can extract voiceprint features from the wake-up trigger voice data and match these extracted features with the standard voiceprint features to obtain the voiceprint matching result.
[0063] Step 208: When the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match, the meeting is switched to the speaking state on the local end.
[0064] The target wake-up word is a pre-defined wake-up word that triggers a user's voice activation to speak. There can be one or more target wake-up words, which can be flexibly set according to actual needs. Local speaking status refers to the audio state where the conference supports speaking locally; that is, audio data collected locally can be transmitted to various conference clients, thus enabling local speaking.
[0065] Specifically, if the terminal determines that the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word and the voiceprint matching result is a match, it means that the voice data sent by the user corresponding to the speaking account with speaking authority on the local end triggered the voice wake-up. If the user needs to speak on the local end, the terminal will switch the meeting to the speaking state on the local end, thereby supporting the transmission of voice data collected on the local end in the meeting.
[0066] Step 210: When the local device is in speaking mode, collect the local device's speech data and send the speech data to the meeting client.
[0067] In this context, conference speaking voice data refers to the voice data emitted by users on their local devices that needs to be transmitted to various conference clients. In other words, conference speaking voice data is the user's speech data during the conference. The conference clients are those that join the conference and send the conference speaking voice data to their respective clients, thereby enabling the dissemination of user speech throughout the conference and achieving effective communication in the online meeting. Specifically, when the terminal switches to local speaking mode, the user can emit voice messages locally. The terminal collects the conference speaking voice data emitted by the local user and sends it to the conference clients, thus enabling the dissemination of the local speech content throughout the conference.
[0068] In a specific application, a user enters a meeting through a conference client on their terminal. Upon entering the meeting, the default setting is a non-speaking state, awaiting voice wake-up, meaning the user cannot speak during the meeting via the conference client. If the user needs to speak, they can send voice data, generating a voice wake-up trigger event. The terminal detects this event, indicating the user needs to speak, and then acquires the wake-up trigger voice data from the local terminal. The terminal performs wake-up word detection on the local wake-up trigger voice data, obtaining the wake-up word detection result. If the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word, the terminal further performs voiceprint recognition processing on the wake-up trigger voice data. Specifically, it matches the wake-up trigger voice data with the standard voiceprint data of the account associated with speaking permissions on the local terminal, obtaining a voiceprint matching result. If the voiceprint matching result matches, it indicates that the wake-up trigger voice data was sent by a user with speaking permissions, and the meeting is switched to a speaking state on the local terminal, thus supporting the user's speaking on the local terminal. The terminal collects the meeting speaking voice data sent by the user on the local terminal and sends the meeting speaking voice data to various conference clients, thereby enabling the dissemination of the local terminal's speaking content throughout the meeting. The content spoken by a user on the local device is only transmitted in the meeting after the wake-up triggers the voice data through wake word detection and voiceprint matching, thus switching the meeting to a local speaking state. This effectively prevents the leakage of content unrelated to the local device's speech into the meeting, thereby improving meeting security.
[0069] Furthermore, if the wake-up word detection result indicates that the wake-up trigger voice data does not include the target wake-up word, or the voiceprint matching result is inconsistent, it means that the voice data sent by the user on the local end is not used for voice wake-up, or the voice data sent on the local end is not sent by a user with speaking permissions. In this case, the terminal continues to maintain the non-speaking state of the meeting on the local end, thereby avoiding the leakage of content on the local end that is unrelated to the meeting's speaking content into the meeting and improving the security of the meeting.
[0070] In practical applications, for wake-up word detection and voiceprint matching processing of wake-up trigger voice data, wake-up word detection can be performed on the wake-up trigger voice data first. If the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word, then voiceprint matching can be performed based on the wake-up trigger voice data. Alternatively, voiceprint matching can be performed on the wake-up trigger voice data first. If the voiceprint matching result is a match, then wake-up word detection can be performed based on the wake-up trigger voice data. This two-level confirmation process for wake-up trigger voice data can ensure the accuracy of wake-up triggering while improving the processing efficiency of wake-up triggering.
[0071] In the aforementioned meeting interaction method, when the meeting is in a non-speaking state awaiting voice wake-up on the local end, wake-up trigger voice data on the local end is subjected to wake-up word detection and voiceprint matching. When the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word and the voiceprint matching result is a match, the meeting is switched to a speaking state on the local end, and the collected meeting speaking voice data is sent to the meeting client. During the meeting speaking control process, when the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word and the voiceprint matching result is a match, it ensures that the speaking on the local end is effectively confirmed, which can avoid voice leakage and thus improve the security of the meeting.
[0072] In one embodiment, wake-up word detection is performed on wake-up trigger voice data to obtain wake-up word detection results, including: extracting features from the wake-up trigger voice data to obtain initial audio features of the wake-up trigger voice data; performing differential mapping on the initial audio features to obtain recognized audio features; and performing wake-up word detection based on the recognized audio features to obtain wake-up word detection results.
[0073] The initial audio features are extracted from the wake-up-triggered speech data and are used to characterize the speech signal features of the wake-up-triggered speech data. For example, the initial audio features can be Mel Frequency Cepstral Coefficients (MFCC) features. MFCC features are constructed based on Mel cepstral coefficients, which can be applied in automatic speech and speaker recognition. Mel cepstral coefficients are a set of key coefficients used to construct the Mel cepstral spectrum. From a segment of speech data, a set of cepstral coefficients that sufficiently represents this speech data can be obtained. The Mel cepstral coefficients are the cepstral coefficients derived from this cepstral spectrum, i.e., the spectrum of the spectrum. Unlike ordinary cepstral spectra, the biggest feature of Mel cepstral spectra is that the frequency bands on the Mel cepstral spectrum are uniformly distributed on the Mel scale. That is to say, the frequency bands are closer to the nonlinear human auditory system than the generally seen linear cepstral representation. Differential mapping can be performed on the initial audio features through differential mapping at one level or more to obtain the recognized audio features. The recognized audio features obtained by mapping can specifically be K-dimensional feature vectors. The identified audio features are used for wake word detection. Further feature processing and classification can be performed on the identified audio features to achieve wake word detection and obtain wake word detection results.
[0074] Specifically, the terminal extracts features from the wake-up trigger voice data to obtain the initial audio features of the wake-up trigger voice data. Specifically, the terminal can extract Mel-spectrum features from the wake-up trigger voice data and obtain the initial audio features based on the obtained Mel-spectrum features. For example, the terminal can directly use the Mel-spectrum features as the initial audio features of the wake-up trigger voice data, or it can further process the Mel-spectrum features, such as further mapping processing, to obtain the initial audio features of the wake-up trigger voice data. The terminal performs differential mapping on the initial audio features, such as first-order and second-order differential processing, to obtain the recognized audio features, which can specifically be a K-dimensional feature vector. The terminal performs wake-up word detection based on the recognized audio features. For example, the terminal can input the recognized audio features into a pre-trained wake-up word detection model, so that the wake-up word detection model can detect the wake-up word based on the recognized audio features and output the wake-up word detection result. The wake-up word detection result indicates whether the wake-up trigger voice data includes the target wake-up word.
[0075] In this embodiment, after the terminal extracts the initial audio features from the wake-up trigger voice data, it performs differential mapping and performs wake-up word detection based on the obtained recognition audio features to obtain the wake-up word detection result. This enables the terminal to obtain highly expressive recognition audio features from the wake-up trigger voice data. By performing wake-up word detection based on the recognition audio features, the accuracy of wake-up word detection can be ensured.
[0076] In one embodiment, feature extraction of wake-up trigger speech data to obtain initial audio features of the wake-up trigger speech data includes: performing time-domain preprocessing on the wake-up trigger speech data to obtain intermediate speech data; performing frequency-domain transformation on the intermediate speech data to obtain the energy spectrum of the intermediate speech data; performing frequency transformation on the energy spectrum through a filter bank to obtain the power spectrum; and performing discrete transformation on the power spectrum to obtain the initial audio features of the wake-up trigger speech data.
[0077] Temporal preprocessing refers to preprocessing the wake-up trigger speech data in the time domain, including but not limited to various preprocessing techniques such as framing, pre-emphasis, and windowing. Intermediate speech data is obtained by performing temporal preprocessing on the wake-up trigger speech data. Frequency domain transformation refers to converting the intermediate speech data from the time domain to the frequency domain to obtain the energy spectrum. The filter bank can be a triangular bandpass filter bank constructed on the Mel frequency axis. By performing frequency transformation on the energy spectrum using the filter bank, the linear frequencies of the energy spectrum can be converted to Mel frequencies. Discrete transformation of the power spectrum, such as discrete cosine transform, can obtain the initial audio features of the wake-up trigger speech data. These initial audio features can be Mel cepstral features.
[0078] Specifically, the terminal performs temporal preprocessing on the wake-up trigger voice data, including framing, pre-emphasis, and windowing, to obtain intermediate voice data. The terminal then performs frequency domain transformation on the intermediate voice data, such as using a Fast Fourier Transform (FFT), to obtain the energy spectrum. The terminal acquires a triangular bandpass filter bank constructed on the Mel frequency axis and performs frequency transformation on the energy spectrum using the filter bank to obtain the power spectrum. Finally, the terminal performs discrete transformation on the power spectrum, such as a Discrete Cosine Transform, to obtain Mel cepstral features. Based on these Mel cepstral features, the initial audio features of the wake-up trigger voice data can be obtained.
[0079] In this embodiment, the terminal performs time-domain preprocessing, frequency-domain conversion, frequency conversion, and discrete transformation on the wake-up trigger voice data in sequence to obtain the initial audio features characterizing the wake-up trigger voice data. Based on the initial audio features, wake-up word detection processing can be performed to ensure the accuracy of wake-up word detection.
[0080] In one embodiment, performing differential mapping on initial audio features to obtain recognizable audio features includes: performing differential processing on the initial audio features at least once to obtain recognizable audio features.
[0081] Specifically, the terminal performs at least one differential processing on the initial audio features, such as performing first-order and second-order differential processing on the initial audio features to obtain the recognized audio features.
[0082] Furthermore, wake word detection is performed based on the identified audio features to obtain wake word detection results, including: obtaining a wake word detection model; the wake word detection model is obtained by training historical triggered speech data carrying wake word tags; wake word detection is performed on the identified audio features through the wake word detection model to obtain the wake word detection results output by the wake word detection model.
[0083] The wake word detection model is obtained by training on historical triggered speech data carrying wake word labels. This historical triggered speech data is historical data; by training on this data, the wake word detection model can be obtained. The wake word detection model can detect wake words on the input audio features and output the wake word detection result.
[0084] Specifically, the terminal acquires a pre-trained wake word detection model, which is trained using historical trigger speech data. The specific network structure of the wake word detection model can be configured according to actual needs, and may include, but is not limited to, various neural network architectures such as fully connected networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and attention mechanisms. The terminal uses the wake word detection model to detect wake words based on the identified audio features. Specifically, the terminal inputs the identified audio features into the wake word detection model, which then performs wake word detection based on the identified audio features and outputs the wake word detection result.
[0085] In this embodiment, the terminal performs at least one differential processing on the initial audio features to obtain recognition audio features suitable for input to the wake word detection model. The wake word detection model then performs wake word detection on the recognition audio features to obtain the wake word detection result output by the wake word detection model, which can ensure the accuracy and processing efficiency of wake word detection.
[0086] In one embodiment, wake word detection is performed on the identified audio features using a wake word detection model to obtain the wake word detection result output by the wake word detection model, including: sequentially processing the identified audio features through a feature processing layer structure of at least two levels in the wake word detection model to obtain intermediate audio features; classifying wake words based on the intermediate audio features through a classification layer structure in the wake word detection model to obtain a classification probability distribution; and obtaining the wake word detection result based on the classification probability distribution.
[0087] The wake word detection model comprises a multi-layered structure, capable of performing various processing operations on the input, including feature extraction, feature mapping, and classification. The feature processing layer performs feature processing on the input, including feature extraction and feature mapping; the classification layer classifies the input as wake words, obtaining a classification probability distribution. This classification probability distribution represents the probability distribution of wake words included in the wake-up trigger speech data, and the wake word detection result can be obtained based on this distribution.
[0088] Specifically, the terminal sequentially processes the recognized audio features through a feature processing layer structure of at least two levels in the wake word detection model to obtain intermediate audio features. For example, the terminal can extract features from the recognized audio features using convolutional layers in the wake word detection model, perform feature mapping on the extracted feature output through pooling layers, extract features again from the mapped output through convolutional layers, and fuse the extracted feature output through fully connected layers to obtain intermediate audio features. The terminal then uses the classification layer structure in the wake word detection model to classify wake words based on the intermediate audio features, obtaining a classification probability distribution. For example, the terminal can use a normalized exponential function in the wake word detection model to classify wake words based on the intermediate audio features and obtain a classification probability distribution. The terminal obtains the wake word detection result based on the classification probability distribution. Specifically, based on the classification probability distribution, it can determine whether the wake-up trigger voice data includes the target wake word, thus obtaining the wake word detection result.
[0089] In this embodiment, the terminal uses the feature processing layer structure and classification layer structure in the pre-trained wake word detection model to perform wake word detection based on the input recognition audio features, thereby obtaining the wake word detection result, which can ensure the accuracy and processing efficiency of wake word detection.
[0090] In one embodiment, such as Figure 3 As shown, the voiceprint matching process involves matching the wake-up trigger voice data with the standard voiceprint data of the account associated with speaking permissions on the local device to obtain the voiceprint matching result, including:
[0091] Step 302: Extract wake-up trigger voiceprint features from wake-up trigger voice data.
[0092] The wake-up trigger voice data consists of voice data emitted by meeting participants on their local devices. This data is used to trigger the meeting's wake-up and switch the meeting's audio status locally. The wake-up trigger voiceprint feature is the voiceprint feature extracted from the wake-up trigger voice data. A voiceprint is a sound wave spectrum carrying speech information displayed using electroacoustic instruments. Human voice has specificity and stability; like fingerprints, voiceprints serve an identification function. Voiceprint features are used to characterize a user's voiceprint.
[0093] Specifically, the terminal extracts voiceprint features from the wake-up trigger voice data to obtain wake-up trigger voiceprint features. These voiceprint features can be used to identify the user who sent the wake-up trigger voice data, thus determining whether the wake-up trigger voice data was sent by a user with speaking privileges. In practice, the terminal can extract the wake-up trigger voiceprint features from the wake-up trigger voice data using methods such as MFCC or deep feature extraction.
[0094] Step 304: Determine the standard voiceprint characteristics of the standard voiceprint data of the account with speaking permissions associated with the local terminal.
[0095] Among them, the speaking permission account refers to the account with speaking privileges on the local end, and the standard voiceprint data is the baseline voiceprint data corresponding to the speaking permission account, which can be pre-entered by the user based on the speaking permission account. Using the standard voiceprint data, identity recognition can be performed on the voice data on the local end to determine the identity of the user sending the voice data. The standard voiceprint feature is the voiceprint feature extracted from the standard voiceprint data.
[0096] Specifically, the terminal obtains the standard voiceprint data of the account associated with speaking permissions on the local end, and further obtains the standard voiceprint features of the standard voiceprint data. The standard voiceprint features can be extracted from the standard voiceprint data in advance. For example, after the user corresponding to the speaking permission account enters the standard voiceprint data, the voiceprint features can be extracted from the standard voiceprint data to obtain the standard voiceprint features. The standard voiceprint features are then associated with the corresponding speaking permission account. Based on this association, the terminal can query the corresponding standard voiceprint features according to the speaking permission account.
[0097] Step 306: Match the wake-up trigger voiceprint features with the standard voiceprint features to obtain the voiceprint matching result.
[0098] Specifically, the terminal performs voiceprint feature matching between the wake-up trigger voiceprint feature and the standard voiceprint feature to obtain the voiceprint matching result. Specifically, the terminal can determine the similarity between the wake-up trigger voiceprint feature and the standard voiceprint feature, and identify the identity of the user who sent the wake-up trigger voice data based on the obtained similarity, so as to determine whether the user who sent the wake-up trigger voice data has speaking permissions on the local end, that is, whether the account corresponding to the user belongs to the speaking permission account associated with the local end.
[0099] In this embodiment, the terminal extracts the wake-up trigger voiceprint features from the wake-up trigger voice data and performs voiceprint feature matching with the standard voiceprint features of the standard voiceprint data to achieve voiceprint matching processing of the wake-up trigger voice data. The voiceprint features can be used to accurately identify the identity of the user who sent the wake-up trigger voice data, ensuring the accuracy of the voiceprint matching results.
[0100] In one embodiment, in response to a voice wake-up trigger event, acquiring wake-up trigger voice data on the local end includes: collecting voice data on the local end; performing silence detection based on the voice features of the voice data collected on the local end to obtain a silence detection result; and when the silence detection result indicates that voice wake-up is triggered on the local end, acquiring the wake-up trigger voice data on the local end from the voice data collected on the local end.
[0101] The voice wake-up trigger event refers to the event that triggers wake-up via voice. Specifically, it can be generated when a user initiates a voice wake-up operation, or when voice input is detected from the local terminal. The local terminal refers to the client running on the local machine. Specifically, when listening to or speaking during a conference using a conference client, the local terminal can be the conference client running locally on the terminal. The terminal collects voice data locally, thereby capturing the voice input emitted by the user. The voice features are extracted from the voice data collected locally. Based on these voice features, silence detection can be performed to determine whether the user has emitted voice input locally. Wake-up trigger voice data consists of voice input emitted by conference participants on the local terminal, used to trigger a wake-up of the conference and switch the local voice status.
[0102] Specifically, the terminal collects voice data locally to detect whether users participating in the meeting are speaking locally. The terminal performs silence detection on the locally collected voice data. Specifically, the terminal can extract voice features from the voice data and use these features for silence detection to obtain a silence detection result. In practical applications, the terminal can compare the voice features of the collected voice data with silence detection feature thresholds to achieve silence detection processing specific to the local end. Voice features may include, but are not limited to, short-time zero-crossing rate, short-time energy, and fundamental frequency. The zero-crossing rate (ZCR) refers to the number of times the voice signal crosses zero in each frame, i.e., the number of times it changes from positive to negative or vice versa; short-time energy refers to the energy of the voice signal within a short period; and fundamental frequency refers to the frequency of the fundamental tone in a polyphony, which is the lowest frequency and strongest in a voice signal. When the silence detection result indicates that a voice wake-up is triggered locally, it means that the user participating in the meeting locally has decided to start speaking. The terminal then obtains the wake-up trigger voice data from the locally collected voice data. Specifically, the terminal can extract the wake-up trigger voice data issued by the user from the collected semantic data, and then use the wake-up trigger voice data to perform wake-up trigger processing for the meeting.
[0103] In this embodiment, the terminal performs silence detection on the voice features of the voice data collected locally. When the silence detection result indicates that voice wake-up is triggered locally, the wake-up trigger voice data of the local terminal is extracted from the collected voice data. The wake-up trigger voice data is used to perform wake-up trigger processing for the meeting, which can avoid false triggering of wake-up.
[0104] In one embodiment, when the meeting is in a non-speaking state awaiting voice wake-up on the local end, in response to a voice wake-up trigger event, the wake-up trigger voice data on the local end is obtained, including: in response to a trigger operation on the meeting entry, entering the meeting associated with the meeting entry; setting the meeting to be in a non-speaking state awaiting voice wake-up on the local end; and in response to a voice wake-up trigger event, obtaining the wake-up trigger voice data on the local end.
[0105] The meeting entry point refers to the point through which users access the meeting. Users can trigger an action by interacting with the meeting entry point displayed on their terminal. Meeting entry points can take various forms, such as links or button controls. Meeting entry points are associated with meetings; a single meeting can have multiple different entry points, allowing users to enter related meetings through these entry points.
[0106] Specifically, the terminal can display a meeting entry in the meeting client. Users can trigger operations based on this meeting entry, and the terminal responds by entering the meeting associated with that entry. The terminal sets the meeting to a non-speaking state awaiting voice wake-up on the local end, thus preventing irrelevant voice content sent locally upon entering the meeting from being broadcast into the meeting. In response to a voice wake-up trigger event, the terminal obtains the wake-up trigger voice data from the local end and performs voice wake-up processing based on the obtained wake-up trigger voice data.
[0107] In this embodiment, after entering the associated meeting through the meeting portal, the terminal sets the meeting to a non-speaking state where it is awaiting voice wake-up on the local end. At this time, the voice data on the local end will not be transmitted to the meeting. This can prevent voice leakage caused by voice content that is not related to the meeting when entering the meeting. This improves the security of the meeting.
[0108] In one embodiment, the conference interaction method further includes: when the duration for which no conference speaking voice data is collected on the local end meets the speaking end duration condition, switching the conference from the speaking state on the local end back to the non-speaking state on the local end.
[0109] The "end of speech" condition is used to determine whether a speech in a meeting has ended. Specifically, it can include a "end of speech" threshold. When the duration for which no local audio data is collected reaches the "end of speech" threshold, it is considered that the user on the local end has stopped speaking and their speech has ended. The local audio status of the meeting can be switched, specifically, the meeting can be switched from the speaking state on the local end to the non-speaking state on the local end. This prevents unrelated audio generated on the local end after the speech ends from continuing to propagate into the meeting and causing audio leakage.
[0110] Specifically, the terminal detects local speech. Upon determining that local speech has ended, the terminal can switch the local voice status. When the terminal does not collect local conference speech data—that is, when the terminal does not collect user-generated conference speech data locally—the terminal calculates the duration for which no conference speech data has been collected. The terminal queries a pre-set speech end duration condition and compares the calculated duration with the speech end duration condition to determine if the duration meets the condition. For example, the speech end duration condition may include a speech end duration threshold. The terminal compares the duration with the speech end duration threshold; when the duration reaches the threshold, it indicates that the user has not spoken for a certain period, and the local user's speech can be considered finished. The microphone can then be muted to prevent voice leakage after the meeting ends.
[0111] In practical applications, users can switch the audio status of the meeting locally by triggering audio status switching operations, including switching from speaking to non-speaking, or vice versa. When a user ends their speech, they can trigger an audio status switching operation, such as clicking the audio status switching control on the meeting. The terminal responds to this operation, switching the local audio status between speaking and non-speaking.
[0112] In this embodiment, when the duration for which the terminal does not collect local conference speaking voice data meets the speaking end duration condition, the local speaking is considered to have ended. The terminal then switches the conference from the speaking state on the local end back to the non-speaking state on the local end. This prevents unrelated voice data generated on the local end after the speaking ends from continuing to propagate into the conference and causing voice leakage, thus ensuring the security of voice data in the conference.
[0113] In one embodiment, the meeting interaction method further includes: entering the meeting and displaying the meeting information area associated with the meeting; displaying a meeting status indicator indicating a non-speaking state in the meeting information area; when voice data is detected on the local end, prompting a voice state switch by triggering voice data through a perceptible wake-up trigger prompt message; and when the meeting switches to a speaking state on the local end, switching the meeting status indicator to indicate a speaking state.
[0114] The meeting information area displays various information, including participant information, meeting topic, communication details, meeting minutes, and meeting settings. The meeting status indicator identifies the local audio status, such as whether the meeting is in a speaking or non-speaking state. Different audio states can be clearly indicated by different display content, allowing users to easily understand the meeting's current audio status. The wake-up trigger prompt is used to prompt users to switch audio states by triggering wake-up voice data. The form and content of the wake-up trigger prompt can be customized. It can be displayed visually on the terminal interface or broadcast as a voice message, prompting users to switch audio states by triggering wake-up voice data.
[0115] Specifically, when a terminal enters a meeting, it displays the meeting information area associated with the meeting. This area shows various meeting-related information. A meeting status indicator indicating a non-speaking state is displayed in the meeting information area, thus informing the user that the meeting is in a non-speaking state locally upon entry, and that local voice data will not be transmitted to the meeting. The display method and content of the meeting status indicator can be flexibly configured according to actual needs; for example, a disabled microphone icon can be displayed to indicate that the meeting is in a non-speaking state locally. The terminal detects voice data locally. When voice data is detected locally, it indicates that the user on the local end has emitted voice data and may need to speak in the meeting. The terminal then prompts the user to switch the voice state by triggering voice data through wake-up. For example, the terminal can broadcast a voice prompt to prompt the user on the local end to switch the voice state by triggering voice data through wake-up, thereby switching the meeting from a non-speaking state locally to a speaking state for speaking. When the meeting switches to local speaking mode, the terminal switches the meeting status indicator to indicate speaking mode, thus prompting the meeting to be in speaking mode locally. At this time, the voice data sent locally will be transmitted to the meeting.
[0116] Furthermore, when a terminal enters a meeting, a perceptible wake-up trigger prompt can be used to indicate that the voice state switch is triggered by wake-up-triggered voice data, thus prompting the user to switch voice states by activating wake-up-triggered voice data. During the meeting, the terminal can also periodically use perceptible wake-up trigger prompts to remind the user to switch voice states by activating wake-up-triggered voice data.
[0117] In this embodiment, the terminal uses the display status of the meeting status indicator to indicate the audio status of the meeting on the local end. This provides a prompt regarding the audio status on the local end, preventing audio leakage and improving meeting security. The terminal also prompts the user on the local end to switch audio status via voice wake-up through a perceptible wake-up trigger message. This allows the user to switch to the local speaking state and speak without needing to manually trigger the audio status switch, thus improving the efficiency of audio status switching during the meeting.
[0118] In one embodiment, the meeting interaction method further includes: displaying a meeting setting operation area in response to a meeting setting trigger operation triggered by the meeting; the meeting setting operation area includes a wake-up word setting item; and determining a target wake-up word based on the wake-up word setting operation in response to a wake-up word setting operation triggered by the wake-up word setting item.
[0119] The meeting settings trigger operation is used to activate settings specific to the meeting, such as triggering user-defined meeting settings controls. The meeting settings operation area is for configuring meeting details, including but not limited to settings for the meeting topic, time, location, participants, speaking permissions, and wake-up word. This area includes a wake-up word setting section, used to configure the target wake-up word for the meeting. The target wake-up word is a pre-set wake-up word that triggers the user's voice activation. By triggering the wake-up word setting operation, users can customize the target wake-up word for the meeting according to their needs.
[0120] Specifically, users can trigger meeting settings, such as by activating meeting settings controls. The terminal responds to these actions by displaying the meeting settings area, which includes at least the wake-up word setting item. This allows users to flexibly set the target wake-up word for the meeting. The terminal can also trigger wake-up word setting actions based on these actions, determining the target wake-up word set by the user. In practice, the wake-up word setting item allows users to select a wake-up word from a pre-selected list. Users can also edit the wake-up word, customizing it to personalize the meeting's target wake-up word. Furthermore, the meeting settings area also allows users to configure settings such as whether the camera is on, whether voice wake-up is enabled, and the content displayed for meeting participants.
[0121] In this embodiment, the user can trigger a wake-up word setting operation in the meeting settings operation area. The terminal responds to the wake-up word setting operation to determine the target wake-up word set by the user for the meeting, thereby performing voice wake-up processing through the target wake-up word, expanding the applicable scenarios for voice wake-up of meeting speeches.
[0122] In one embodiment, the meeting interaction method further includes: in response to a configuration operation for meeting participants on the local end, obtaining the local participant account on the local end; determining the speaking permission account associated with the local end based on the local participant account; and establishing the association between the standard voiceprint data of the local participant account and the speaking permission account.
[0123] The participant configuration operation refers to the process of configuring the participants attending the meeting locally. This operation allows you to set the accounts corresponding to users attending the meeting on your local meeting client. Each local participant account corresponds to one account. Standard voiceprint data can be pre-entered by the user corresponding to each local participant account and is used for identity verification. Speaking permission accounts refer to accounts with speaking privileges locally, which can be determined from the local participant accounts. For example, you can select the speaking permission accounts from the local participant accounts to grant speaking privileges.
[0124] Specifically, users can configure meeting participants locally. This configuration can be done when initiating the meeting, upon joining the meeting, or during the meeting. In practical applications, users can trigger participant configuration operations on the local meeting client on the terminal, such as triggering operations on participant configuration controls. The terminal responds to the user-triggered participant configuration operation and obtains the local participant accounts configured according to that operation. Each local participant account can be associated with one user, and the associated user can be determined based on the local participant account. Not every user participating in the meeting locally on the terminal needs to speak; therefore, the terminal can further determine which accounts require speaking permissions from the local participant accounts. In the specific implementation, users can further configure the speaking permissions of local participant accounts on the terminal, granting speaking permissions to participants who need to speak during the meeting, and designating them as the associated speaking permission accounts on the local terminal. In practical applications, by default, all local participants' accounts can be set to have speaking permissions associated with the local terminal. This means that all participants on the local terminal have speaking permissions in the meeting and can be activated by voice to speak.
[0125] After the terminal determines the speaking permission account associated with the local terminal, it establishes a relationship between the standard voiceprint data of the local participant account and the speaking permission account. Specifically, the terminal can determine the standard voiceprint data of the local participant account. This standard voiceprint data can be pre-entered by the user using their account credentials. When a user's account is configured as a local participant account, the terminal can obtain the corresponding standard voiceprint data. In practical applications, when a user logs into the meeting client with their account, they can enter standard voiceprint data in the meeting client, allowing for user identification based on the entered standard voiceprint data. The entered standard voiceprint data can be stored in a standard voiceprint database, from which the terminal can retrieve the standard voiceprint data of the local participant account. The terminal associates the standard voiceprint data of the local participant account with speaking permissions with the corresponding speaking permission account, establishing a relationship. This allows the terminal to use the standard voiceprint data of the local participant account as the standard voiceprint data of the speaking permission account, thus identifying the user who is speaking and determining whether they have speaking permissions.
[0126] In this embodiment, the user terminal can configure the users participating in the meeting locally, and determine the speaking permission account associated with the local terminal based on the local participant account, establish the standard voiceprint data of the local participant account and the association relationship between the speaking permission account, so as to support multiple people to participate in the meeting through the same terminal's meeting client, and support multiple people to speak in the meeting by voice wake-up.
[0127] In one embodiment, the meeting interaction method further includes: when it is determined that the account of the next participant to speak in the meeting belongs to the account with speaking permissions associated with the local end, switching the meeting from a non-speaking state to a speaking state on the local end; providing speaking prompts on the local end through perceptible speaking prompt information; and switching the meeting from a speaking state back to a non-speaking state on the local end when the account of the next participant to speak finishes speaking.
[0128] The "participant account" refers to the user account used to participate in the meeting. Each participant account corresponds to one user, allowing for user identification. The "next speaker account" refers to the participant account that will speak next after the current speaker finishes. Speaking prompts are used to notify the user of the next speaker account to speak on their local device. These prompts are presented in a perceptible manner, such as displaying them on the terminal interface or using voice prompts to remind the user to speak on their local device.
[0129] Specifically, the terminal determines the account of the next speaker in the meeting. This account can be determined based on the meeting schedule, such as the speaking order of all participants. Alternatively, the terminal can also analyze the content of the current speech. If the current speech includes the account of the next speaker, it can determine that account. For example, if the current speech is "My speech is finished. Next, please welcome Zhang San to speak," the terminal can identify Zhang San as the next speaker and determine the account by querying Zhang San's account. The terminal then matches the account of the next speaker with the associated speaking permission account on the local machine to determine if the account has speaking permission and thus whether it has the authority to speak in the meeting.
[0130] Once the terminal determines that the account of the next speaker belongs to the account associated with speaking permissions on the local terminal (i.e., the account of the next speaker has speaking privileges), the terminal switches the meeting from a non-speaking state to a speaking state on the local terminal, thus supporting the user of the next speaker's account to speak in the meeting locally. Furthermore, the terminal acquires pre-set, perceptible speaking prompts and displays them locally to remind the user of the next speaker's account to speak promptly. For example, the terminal can display text-based speaking prompts on the local interface to remind the user of the next speaker's account to speak promptly; the terminal can also play audio-based speaking prompts locally to remind the user of the next speaker's account to speak promptly. When the next speaker's account ends their speech, if the terminal detects that the speech content includes an expression indicating the end of the speech, or if the terminal does not collect audio data locally, determining that the speech has ended, the terminal switches the meeting from a speaking state back to a non-speaking state on the local terminal. This prevents irrelevant audio content from continuing to be disseminated in the meeting, thus avoiding audio leakage and ensuring meeting security.
[0131] Furthermore, if the account of the next speaker does not belong to the account associated with speaking permissions on the local end (i.e., it does not have speaking permissions on the local end), speaking permission configuration can be triggered for the next speaker's account, thereby granting speaking permissions to the next speaker's account. In specific applications, when the terminal determines that the account of the next speaker in the meeting belongs to the account associated with speaking permissions on the local end, the terminal can also provide a speaking prompt on the local end only through perceptible speaking prompt information, prompting the user to activate the system via voice to switch to speaking mode on the local end, thereby preventing the leakage of voice messages unrelated to the meeting.
[0132] In this embodiment, when the terminal determines that the next participant's account has speaking permissions on the local end, the terminal automatically switches the meeting from a non-speaking state to a speaking state on the local end. It also provides a noticeable speaking prompt to the user on the local end, further eliminating the need for wake-up triggering and improving the efficiency of meeting speech processing. When the next participant's account finishes speaking, the terminal automatically switches the meeting back from a speaking state to a non-speaking state on the local end, thus preventing voice leakage after the meeting ends and improving meeting security.
[0133] This application also provides an application scenario in which the above-described meeting interaction method is applied. Specifically, the meeting interaction method is applied in this scenario as follows:
[0134] The rapid development of the internet has transformed people's lives and work habits. Web conferencing has broken geographical limitations, allowing users to access meetings anytime, anywhere, meeting their needs for cloud-based office work, cloud-based teaching, and online collaboration. However, this convenient communication also makes it easy for users to forget to mute their microphones, leading to privacy breaches. Furthermore, in some scenarios, users are not physically present and cannot effectively switch speaking modes. Current conferencing applications offer various speaking settings. The host can mute the entire meeting, and users can unmute themselves. Users can control their speaking status via mouse clicks or keyboard input, and can also set whether their voice is enabled upon joining the meeting. However, while these settings collectively determine the user's final voice status, the ease with which microphones are muted upon joining, after speaking, or even accidental activation of the meeting access button on mobile devices frequently results in unintentional speech during meetings. This is especially problematic when discussing private matters unrelated to the meeting, leading to privacy breaches and embarrassing situations for participants, resulting in a poor user experience.
[0135] Based on this, the meeting interaction method provided in this embodiment, based on voiceprint recognition and wake-up word voice activation, ensures that the user's voice content during the meeting has been confirmed by the user's own voice, preventing the spread of voice content unrelated to the meeting, avoiding privacy leaks caused by accidental activation, and reducing unexpected situations when users participate in important meetings. At the same time, through voice control, users do not need to switch meeting applications to perform application operations, improving the efficiency of switching voice states. When not near the device, users can also start speaking, improving the convenience of speaking in meetings, improving the efficiency of users freely switching speaking when using mobile devices, and enhancing the user experience.
[0136] Specifically, such as Figure 4 As shown in the diagram, the interface for controlling voice status using a wake-up word in a network audio / video conferencing application is illustrated. The lower left corner displays the current voice status as "awaiting voice wake-up," and a slash is added to the microphone to indicate that it is not currently activated. If a user speaks without using a wake-up word during the meeting, they are prompted whether they need to say the wake-up word to activate their voice. To avoid accidental wake-ups, a prompt tone such as "Speaking on" can be emitted through the speaker, or the relevant status can be displayed on the meeting interface. The meeting application interface displays information about the meeting participants, including their names, identities, and avatars. It also displays microphones marked as "voice wake-up," indicating that the microphone is not currently activated on the local device. A floating text prompt in the meeting application interface prompts the user to speak by saying the wake-up word. Figure 5As shown, this is a schematic diagram of the interface after a user activates voice communication. In the meeting application interface, the microphone in the lower left corner is displayed as activated, indicating that the microphone is enabled locally and the user can speak during the meeting. Furthermore, if the user does not speak for a certain period, it automatically switches to a mute state, and the user can restart voice communication using a wake-up word. Further, as... Figure 6 As shown, the conference interaction method provided in this embodiment also supports user configuration to enable the voice wake-up function, and optionally supports custom wake-up words. For custom wake-up words with pre-defined symbol requirements, users can use the custom wake-up words to start speaking. In the conference settings area, users can configure various parameters of the conference, including enabling the voice wake-up function and customizing the wake-up word. Voice wake-up (keyword spotting, KWS) refers to the real-time detection of specific segments of a speaker in a continuous speech stream; the wake-up word is the keyword that wakes the device from sleep mode and triggers a specified response when the user issues the voice command.
[0137] Specifically, the meeting interaction method provided in this embodiment, such as Figure 7As shown, the process includes steps 702 to 712, wherein: Step 702, detecting whether the user has enabled the voice wake-up function; the terminal can detect whether the user has enabled the voice wake-up function for the current member. The meeting application provides a setting to enable the voice wake-up function, which the user can check. After enabling, the meeting application determines whether to turn on the microphone based on the user's voice input. Step 704, detecting device voice input; when the user has enabled the voice wake-up function for the meeting, the terminal detects the device's voice input, i.e., determines whether a user is speaking on the local end. Step 706, whether a wake-up word is hit; when voice input is detected, wake-up word detection is performed on the collected voice data to determine whether a wake-up word is hit, i.e., whether the voice data includes a wake-up word. Step 708, voiceprint recognition to determine if the user has speaking privileges; if the voice data includes a wake-up word, the terminal further performs voiceprint recognition based on the voice data to determine whether the user who sent the voice data is a user with speaking privileges. Step 720: Initiate voice communication; if voiceprint recognition passes, meaning the user has speaking privileges, the terminal initiates voice communication. Specifically, the microphone is activated to collect voice data, which is then sent to the conference client. Step 712: Mute microphone if no speech is received within the timeout period; if the user remains silent for a period exceeding the timeout threshold, the user's speech is considered complete, the microphone is muted, and the local audio is no longer transmitted to the conference. Furthermore, the terminal can maintain its voice state even if the user has not activated voice wake-up, the wake-up word is not detected, or voiceprint recognition fails. Voiceprint recognition is a technology that extracts the speaker's voice characteristics and speech content information to automatically verify the speaker's identity.
[0138] Furthermore, regarding wake-word command recognition and processing, since individual users are mostly silent during multi-person conferences, the silence detection module on the conference equipment can first determine whether the user has started speaking. Wake-word detection is not activated when the user is not speaking. Silence detection is performed by extracting audio signal feature values from the user's speech and using pre-set thresholds. Audio signal features may include short-time zero-crossing rate, short-time energy, and fundamental frequency. Since most speakers in current conferences communicate through headphones, noise issues are generally not significant; therefore, energy-based silence detection is more effective. When the user begins speaking, the silence detection module is activated, and the wake-word detection module starts.
[0139] There are two main types of methods used in wake word detection systems. One is the traditional keyword / filler model based on Hidden Markov Models (HMMs). This model splits keywords into HMM states and represents non-keywords as filler states. A decoding network is built for each wake word, and dynamic programming algorithms, such as the Viterbi algorithm, are used to find the optimal path and make a decision, thus achieving wake word detection. With the development of deep learning, extracting audio features and feeding them into a multi-layer neural network structure to predict wake word probabilities allows for better feature recognition of wake words, offering advantages such as low computational cost, fast response, and low latency.
[0140] Among these, audio features typically employ Mel-frequency cepstral (MFCC) features under short-time spectral density. The MFCC feature extraction process, such as... Figure 8 As shown, the process includes steps 802 to 816, wherein: Step 802, obtaining the speech time-domain signal; the terminal obtains the speech time-domain signal, that is, the wake-up trigger speech data collected locally; Step 804, preprocessing such as framing, pre-emphasis, and windowing; the terminal performs preprocessing such as framing, pre-emphasis, and windowing on the speech time-domain signal; Step 806, Fast Fourier Transform; the terminal performs Fast Fourier Transform on the preprocessed output; Step 808, obtaining the energy spectrum; the terminal obtains the energy spectrum through Fast Fourier Transform; Step 810, filtering with a Mel filter bank; the terminal filters the energy spectrum through a Mel filter bank; Step 812, obtaining the logarithmic spectrum; the terminal obtains the filtered logarithmic spectrum, specifically the Log logarithmic spectrum; Step 814, Discrete Cosine Transform; the terminal performs Discrete Cosine Transform on the logarithmic spectrum; Step 816, obtaining Mel cepstral features; the terminal obtains the Discrete Cosine Transform of the Discrete Cosine Transform.
[0141] Furthermore, the input speech signal is processed by framing, pre-emphasis, and windowing to facilitate speech analysis. Then, the speech signal's spectrum is obtained through Fast Fourier Transform (FFT), and the energy spectrum is obtained by squaring the spectrum. M triangular bandpass filter banks are constructed on the Mel frequency axis to convert the obtained linear frequencies into Mel frequencies. The logarithm of the output of each filter is taken to obtain the logarithmic power spectrum of the corresponding frequency band, and a Discrete Cosine Transform (DCT) is performed to obtain L MFCC coefficients. Based on these L MFCC coefficients, Mel cepstral features are obtained. Specifically, after framing the original signal, many frames are obtained. An FFT (Fast Fourier Transform) is performed on each frame. The FFT converts the time-domain signal to the frequency-domain signal. Stacking the frequency-domain signals (spectral graphs) of each frame in time yields the spectrogram.
[0142] Furthermore, by extracting Mel-Cepstral features (MFCC features) and then performing first- and second-order differencing, the speech signal can be transformed into a K-dimensional feature vector for input to the neural network model. The neural network model is then constructed, with the MFCC feature vector as input. The model can utilize various neural network architectures, including fully connected networks, convolutional neural networks, recurrent neural networks, and attention mechanisms. Taking a wake-word structure composed of convolutional and linear networks as an example, the model structure is as follows: Figure 9 As shown, from top to bottom, there is one convolutional layer (64*20*8), one pooling layer (64*2*2), one convolutional layer (64*10*4), and four fully connected layers. The output state probability is estimated by the normalized exponential function (sotfmax). After the model judges, it can be determined whether the current speech signal contains a wake word, and the wake word recognition result is obtained.
[0143] Furthermore, for the provided custom wake word function, custom wake words can be recognized through model retraining. To reduce false positives, user-defined wake words are judged and required to be longer than a certain length, such as four syllables or more, to reduce the likelihood of them appearing in regular speech. In addition, there are certain requirements for the syllables within the wake word, using words with clear, voiceless syllables, reducing the use of zero initials and reduplicated words, and avoiding conflicts with words commonly used in meetings, thereby improving the wake word accuracy.
[0144] For voiceprint recognition processing, to avoid wake-up words triggered by voices other than the user's own, voiceprint recognition is further introduced to determine if the voice originates from the current user, i.e., whether it comes from a user participating in the meeting through the terminal. For example... Figure 10As shown, by recording user voice data, specifically the voice data of the wake-up word pronunciation, and performing feature extraction and voiceprint registration, a voiceprint model for the user is obtained. Specifically, voiceprint features can be extracted using methods such as MFCC extraction or deep feature extraction. Then, a voiceprint model is generated using voiceprint models, including but not limited to GMM-UBM (Gaussian Mixture Model-Universal Background Model), GMM / Ivector, DNN (Deep Neural Networks) / Ivector, and GSV (Gaussian Super Vector). During verification, test features are generated in the same way. The test features are compared with the model to obtain a final score. A score greater than a certain threshold indicates that the voice matches the current meeting user; otherwise, the user is not awakened to speak. Specifically, features are extracted from the voice data collected locally. After voiceprint verification, the voiceprint is matched with the voiceprint model to obtain a similarity score. Based on the similarity score, it is determined whether the user uttering the voice data is the current meeting participant.
[0145] The meeting interaction method provided in this embodiment is applied to online audio and video conferencing. It identifies whether the user's voice contains a specified wake-up word and matches the user's voiceprint to activate the speaking state. The method then reminds the user to speak via a prompt tone or interface display. If no one speaks for a certain period, the microphone is automatically muted. This effectively prevents accidental voice leakage caused by users accidentally activating their microphones, ensuring that the user's speech has been confirmed by the user, reducing embarrassment caused by accidental remarks during important meetings, and improving the user experience. Furthermore, voice control of the speaking state allows users to freely switch speaking modes when they are not near the device, enhancing the convenience of online meetings.
[0146] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0147] Based on the same inventive concept, this application also provides a conference interaction device for implementing the conference interaction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more conference interaction device embodiments provided below can be found in the limitations of the conference interaction method described above, and will not be repeated here.
[0148] In one embodiment, such as Figure 11 As shown, a conference interaction device 1100 is provided, including: a wake-up trigger voice acquisition module 1102, a wake-up word detection module 1104, a voiceprint matching module 1106, a voice state switching module 1108, and a speech processing module 1110, wherein:
[0149] The wake-up trigger voice acquisition module 1102 is used to acquire wake-up trigger voice data from the local end when the meeting is in a non-speaking state awaiting voice wake-up on the local end in response to a voice wake-up trigger event.
[0150] The wake word detection module 1104 is used to detect wake words in wake-up trigger voice data and obtain wake word detection results.
[0151] The voiceprint matching module 1106 is used to match the wake-up trigger voice data with the standard voiceprint data of the speaking permission account associated with the local terminal to obtain the voiceprint matching result.
[0152] The voice state switching module 1108 is used to switch the conference to the local speaking state when the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word and the voiceprint matching result is a match.
[0153] The speech processing module 1110 is used to collect the speech data of the local device when the speaker is in a speaking state during the meeting, and send the speech data of the local device to the meeting client.
[0154] In one embodiment, the wake-up word detection module 1104 includes a feature extraction module, a differential mapping module, and a recognition audio feature processing module; wherein: the feature extraction module is used to extract features from the wake-up trigger voice data to obtain the initial audio features of the wake-up trigger voice data; the differential mapping module is used to perform differential mapping on the initial audio features to obtain recognition audio features; and the recognition audio feature processing module is used to perform wake-up word detection based on the recognition audio features to obtain the wake-up word detection result.
[0155] In one embodiment, the feature extraction module includes a preprocessing module, a frequency domain conversion module, a frequency conversion module, and a discrete transformation module; wherein: the preprocessing module is used to perform time-domain preprocessing on the wake-up trigger speech data to obtain intermediate speech data; the frequency domain conversion module is used to perform frequency domain conversion on the intermediate speech data to obtain the energy spectrum of the intermediate speech data; the frequency conversion module is used to perform frequency conversion on the energy spectrum through a filter bank to obtain the power spectrum; and the discrete transformation module is used to perform discrete transformation on the power spectrum to obtain the initial audio features of the wake-up trigger speech data.
[0156] In one embodiment, the differential mapping module is further configured to perform differential processing on the initial audio features at least once to obtain the recognized audio features; the recognized audio feature processing module is further configured to obtain a wake word detection model; the wake word detection model is obtained by training historical triggered speech data carrying wake word tags; the wake word detection model is used to perform wake word detection on the recognized audio features to obtain the wake word detection result output by the wake word detection model.
[0157] In one embodiment, the audio feature processing module is further configured to sequentially process the identified audio features through a feature processing layer structure of at least two levels in the wake word detection model to obtain intermediate audio features; classify wake words based on the intermediate audio features through a classification layer structure in the wake word detection model to obtain a classification probability distribution; and obtain a wake word detection result based on the classification probability distribution.
[0158] In one embodiment, the voiceprint matching module 1106 is further configured to extract wake-up trigger voiceprint features from wake-up trigger voice data; determine the standard voiceprint features of the standard voiceprint data of the local terminal associated speaking permission account; and perform voiceprint feature matching between the wake-up trigger voiceprint features and the standard voiceprint features to obtain a voiceprint matching result.
[0159] In one embodiment, the wake-up trigger voice acquisition module 1102 is further configured to collect voice data on the local end; perform silence detection based on the voice features of the voice data collected on the local end to obtain a silence detection result; and when the silence detection result indicates that voice wake-up is triggered on the local end, obtain the wake-up trigger voice data of the local end from the voice data collected on the local end.
[0160] In one embodiment, the wake-up trigger voice acquisition module 1102 is further configured to enter the meeting associated with the meeting entry in response to a trigger operation on the meeting entry; set the meeting to be in a non-speaking state awaiting voice wake-up on the local end; and acquire the wake-up trigger voice data on the local end in response to a voice wake-up trigger event.
[0161] In one embodiment, the system further includes an end-switching module, which switches the conference from a speaking state on the local end back to a non-speaking state on the local end when the duration for which no local conference speaking voice data has been collected meets the speaking end duration condition.
[0162] In one embodiment, the system further includes an information area display module, a status indicator display module, a prompt wake-up module, and a status indicator switching module; wherein: the information area display module is used to enter the meeting and display the meeting information area associated with the meeting; the status indicator display module is used to display a meeting status indicator indicating a non-speaking state in the meeting information area; the prompt wake-up module is used to prompt a voice state switch by triggering voice data through a perceptible wake-up trigger prompt message when voice data is detected on the local end; and the status indicator switching module is used to switch the meeting status indicator to indicate the speaking state when the meeting switches to a speaking state on the local end.
[0163] In one embodiment, the system further includes a trigger setting module and a wake-up word determination module; wherein: the trigger setting module is used to display a meeting setting operation area in response to a meeting setting trigger operation triggered by the meeting; the meeting setting operation area includes a wake-up word setting item; the wake-up word determination module is used to determine a target wake-up word based on the wake-up word setting operation triggered by the wake-up word setting item.
[0164] In one embodiment, the system further includes a participant configuration module, a speaking permission account determination module, and a voiceprint data association module; wherein: the participant configuration module is used to obtain the local participant accounts in response to the participant configuration operation on the local end of the meeting; the speaking permission account determination module is used to determine the speaking permission account associated with the local end based on the local participant accounts; and the voiceprint data association module is used to establish the association relationship between the standard voiceprint data of the local participant accounts and the speaking permission accounts.
[0165] In one embodiment, the system further includes a voice state switching module and a speaking prompt module; wherein: the voice state switching module is used to switch the meeting from a non-speaking state to a speaking state on the local end when it is determined that the account of the next speaker in the meeting belongs to the speaking permission account associated with the local end; the speaking prompt module is used to provide speaking prompts on the local end through perceptible speaking prompt information; the voice state switching module is also used to switch the meeting from a speaking state back to a non-speaking state on the local end when the account of the next speaker finishes speaking.
[0166] Each module in the aforementioned conference interaction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0167] In one embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows: Figure 12 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores conference data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a conference interaction method.
[0168] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0169] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0170] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0171] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0173] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0174] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0175] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A meeting interaction method, characterized in that, The method includes: In response to a trigger operation on the meeting entry point, enter the meeting associated with the meeting entry point, set the meeting to a non-speaking state on the local end that is awaiting voice wake-up, and in response to a voice wake-up trigger event, obtain the wake-up trigger voice data on the local end. The non-speaking state is a state in which the local end does not support speaking in the meeting. The wake-up trigger voice data is subjected to wake-up word detection to obtain wake-up word detection results; The wake-up trigger voice data is matched with the standard voiceprint data of the speaking permission account associated with the local terminal to obtain the voiceprint matching result. When the wake word detection result indicates that the wake-up trigger voice data includes the target wake word and the voiceprint matching result is a match, the meeting is switched to the speaking state on the local end. The speaking state is a state that supports the local end to speak for the meeting. When the local terminal is in speaking mode during the meeting, the local terminal's speaking voice data is collected and sent to the meeting client. The speaking voice data is the speaking data sent from the local terminal to the meeting. When it is determined, based on the meeting schedule information, that the account of the next speaker in the meeting belongs to the account with speaking permissions associated with the local terminal, the meeting is switched from the non-speaking state to the speaking state on the local terminal; a speaking prompt is displayed on the local terminal through perceptible speaking prompt information to prompt the account of the next speaker to speak directly on the local terminal; when the account of the next speaker finishes speaking, the meeting is switched back from the speaking state to the non-speaking state on the local terminal. If the account of the next participant to speak is not an account associated with speaking permissions on the local device, speaking permission configuration is triggered for the account of the next participant to speak, thereby granting speaking permission to the account of the next participant to speak. If the account of the next participant to speak is an account associated with speaking permissions on the local device, a speaking prompt is made on the local device through perceptible speaking prompt information, prompting the account of the next participant to speak to be woken up by voice and switch to speaking state on the local device.
2. The method according to claim 1, characterized in that, The step of detecting a wake-up word in the wake-up trigger voice data to obtain a wake-up word detection result includes: Feature extraction is performed on the wake-up trigger voice data to obtain the initial audio features of the wake-up trigger voice data; The initial audio features are differentially mapped to obtain the recognized audio features; Based on the identified audio features, wake word detection is performed to obtain wake word detection results.
3. The method according to claim 2, characterized in that, The step of extracting features from the wake-up trigger voice data to obtain the initial audio features of the wake-up trigger voice data includes: The wake-up trigger voice data is preprocessed in the time domain to obtain intermediate voice data; The intermediate speech data is frequency domain converted to obtain the energy spectrum of the intermediate speech data; The power spectrum is obtained by frequency conversion of the energy spectrum using a filter bank. The power spectrum is discretized to obtain the initial audio features of the wake-up trigger voice data.
4. The method according to claim 2, characterized in that, The step of performing differential mapping on the initial audio features to obtain the recognized audio features includes: The initial audio features are subjected to at least one differential processing to obtain the recognition audio features; The wake word detection based on the identified audio features, to obtain the wake word detection result, includes: A wake-up word detection model is obtained; the wake-up word detection model is obtained by training historical triggered speech data carrying wake-up word tags. The wake word detection model is used to detect the wake word in the recognized audio features, and the wake word detection result output by the wake word detection model is obtained.
5. The method according to claim 4, characterized in that, The step of detecting wake words in the recognized audio features using the wake word detection model to obtain the wake word detection result output by the wake word detection model includes: The wake word detection model uses at least two levels of feature processing layers to sequentially process the recognized audio features to obtain intermediate audio features. By using the classification layer structure in the wake word detection model, wake words are classified based on the intermediate audio features to obtain the classification probability distribution; The wake word detection results are obtained based on the classification probability distribution.
6. The method according to claim 1, characterized in that, The step of matching the wake-up trigger voice data with the standard voiceprint data of the speaking permission account associated with the local terminal to obtain the voiceprint matching result includes: Extract wake-up trigger voiceprint features from the wake-up trigger voice data; Determine the standard voiceprint features of the standard voiceprint data of the account with speaking permissions associated with the local terminal; The wake-up trigger voiceprint feature is matched with the standard voiceprint feature to obtain the voiceprint matching result.
7. The method according to claim 1, characterized in that, The step of acquiring the wake-up trigger voice data from the local terminal in response to the voice wake-up trigger event includes: Voice data is collected at the local terminal; Silence detection is performed based on the speech features of the speech data collected at the local terminal to obtain silence detection results; When the silence detection result indicates that voice wake-up is triggered on the local end, the wake-up trigger voice data of the local end is obtained from the voice data collected on the local end.
8. The method according to claim 1, characterized in that, The method further includes: When the duration for which no conference speaking voice data is collected on the local end meets the speaking end duration condition, the conference is switched from the speaking state on the local end back to the non-speaking state on the local end.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Upon entering the meeting, the meeting information area associated with the meeting is displayed; The meeting information area displays a meeting status indicator representing the non-speaking state; When voice data is detected on the local end, a perceptible wake-up trigger prompt message is sent to indicate that the voice state is switched by triggering the wake-up trigger voice data; When the meeting switches to speaking mode on the local end, the meeting status indicator is switched to indicate the speaking mode.
10. The method according to any one of claims 1 to 8, characterized in that, The method further includes: In response to a meeting settings trigger operation triggered by the meeting, the meeting settings operation area of the meeting is displayed; the meeting settings operation area includes a wake-word setting item; In response to a wake-word setting operation triggered by the wake-word setting item, the target wake-word is determined according to the wake-word setting operation.
11. The method according to any one of claims 1 to 8, characterized in that, The method further includes: In response to a configuration operation for participants in the meeting on the local end, obtain the local participant accounts in the meeting on the local end; The speaking permission account associated with the local terminal is determined based on the local participant account; Establish the association between the standard voiceprint data of the local participant accounts and the accounts with speaking privileges.
12. A conference interaction device, characterized in that, The device includes: The wake-up trigger voice acquisition module is used to respond to the trigger operation for the meeting entry, enter the meeting associated with the meeting entry, set the meeting to a non-speaking state on the local end that is waiting for voice wake-up, and in response to the voice wake-up trigger event, acquire the wake-up trigger voice data on the local end. The non-speaking state is a state in which the local end does not support speaking for the meeting. The wake-up word detection module is used to detect wake-up words in the wake-up trigger voice data and obtain wake-up word detection results. The voiceprint matching module is used to match the wake-up trigger voice data with the standard voiceprint data of the speaking permission account associated with the local terminal to obtain the voiceprint matching result. The voice state switching module is used to switch the conference to a speaking state on the local end when the wake-up word detection result indicates that the wake-up trigger voice data includes the target wake-up word and the voiceprint matching result is a match. The speaking state is a state that supports the local end to speak for the conference. The speech processing module is used to collect the speech data of the local terminal when the meeting is in a speaking state, and send the speech data of the local terminal to the meeting client. The speech data of the local terminal is the speech data sent to the meeting. The voice state switching module is used to switch the meeting from the non-speaking state to the speaking state on the local terminal when it is determined from the meeting arrangement information of the meeting that the account of the next speaker in the meeting belongs to the speaking permission account associated with the local terminal. The speaking prompt module is used to provide speaking prompts on the local terminal through perceptible speaking prompt information, so as to prompt the next participating member account to speak directly on the local terminal; The voice state switching module is also used to switch the meeting from the speaking state back to the non-speaking state on the local end when the next speaking member account finishes speaking. When the account of the next speaker does not belong to the account with speaking permissions associated with the local terminal, the speaker configuration module is triggered to configure speaking permissions for the account of the next speaker, thereby granting speaking permissions to the account of the next speaker. When the account of the next speaker belongs to the account with speaking permissions associated with the local terminal, the speaking prompt module provides a speaking prompt on the local terminal through perceptible speaking prompt information, prompting the account of the next speaker to be woken up by voice and switch to the speaking state on the local terminal.
13. The apparatus according to claim 12, characterized in that, The wake-up word detection module is also used to extract features from the wake-up trigger voice data to obtain the initial audio features of the wake-up trigger voice data; and to perform differential mapping on the initial audio features to obtain the recognition audio features. Based on the identified audio features, wake word detection is performed to obtain wake word detection results.
14. The apparatus according to claim 13, characterized in that, The wake-up word detection module is further configured to perform time-domain preprocessing on the wake-up trigger speech data to obtain intermediate speech data; perform frequency-domain conversion on the intermediate speech data to obtain the energy spectrum of the intermediate speech data; perform frequency conversion on the energy spectrum through a filter bank to obtain the power spectrum; and perform discrete transformation on the power spectrum to obtain the initial audio features of the wake-up trigger speech data.
15. The apparatus according to claim 13, characterized in that, The wake word detection module is further configured to perform at least one differential processing on the initial audio features to obtain recognized audio features; obtain a wake word detection model; the wake word detection model is obtained by training historical triggered speech data carrying wake word tags; and perform wake word detection on the recognized audio features through the wake word detection model to obtain the wake word detection result output by the wake word detection model.
16. The apparatus according to claim 15, characterized in that, The wake word detection module is also used to sequentially process the recognized audio features through at least two levels of feature processing layer structure in the wake word detection model to obtain intermediate audio features; The wake word detection model uses a classification layer structure to classify wake words based on the intermediate audio features, thereby obtaining a classification probability distribution; and then obtains the wake word detection result based on the classification probability distribution.
17. The apparatus according to claim 12, characterized in that, The voiceprint matching module is also used to extract wake-up trigger voiceprint features from the wake-up trigger voice data; determine the standard voiceprint features of the standard voiceprint data of the speaking permission account associated with the local terminal; and perform voiceprint feature matching between the wake-up trigger voiceprint features and the standard voiceprint features to obtain a voiceprint matching result.
18. The apparatus according to claim 12, characterized in that, The wake-up trigger voice acquisition module is also used to collect voice data on the local end; perform silence detection based on the voice features of the voice data collected on the local end, and obtain a silence detection result; when the silence detection result indicates that voice wake-up is triggered on the local end, the wake-up trigger voice data of the local end is obtained from the voice data collected on the local end.
19. The apparatus according to claim 12, characterized in that, The device further includes: The end-switching module is used to switch the conference from the speaking state on the local end back to the non-speaking state on the local end when the duration of the absence of conference speaking voice data on the local end meets the speaking end duration condition.
20. The apparatus according to any one of claims 12 to 19, characterized in that, The device further includes: An information area display module is used to enter the meeting and display the meeting information area associated with the meeting; A status indicator display module is used to display a meeting status indicator indicating the non-speaking state in the meeting information area; The wake-up prompt module is used to prompt a voice state switch triggered by the wake-up trigger voice data when voice data is detected on the local terminal; The status identifier switching module is used to switch the meeting status identifier to indicate the speaking status when the meeting is switched to the speaking status on the local end.
21. The apparatus according to any one of claims 12 to 19, characterized in that, The device further includes: A trigger module is configured to display the meeting settings operation area in response to a meeting settings trigger operation triggered by the meeting; the meeting settings operation area includes a wake-word setting item; The wake word determination module is used to determine the target wake word in response to a wake word setting operation triggered by the wake word setting item.
22. The apparatus according to any one of claims 12 to 19, characterized in that, The participant configuration module is also used to obtain the local participant account in the meeting on the local end in response to the participant configuration operation for the meeting on the local end; The device further includes: The speaking permission account determination module is used to determine the speaking permission account associated with the local terminal based on the local participant account; The voiceprint data association module is used to establish the association between the standard voiceprint data of the local participant account and the account with speaking privileges.
23. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Characteristic extraction method of geographical name speech signals
CN106782499A
Conference audio control method, system and device and computer readable storage medium
CN110300001A
Microphone control method, electronic device and computer readable storage medium
CN111429914A
Network conference control method and device, electronic equipment and storage medium
CN112885350A